As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.
Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.
Closes#5805 although their (ML?) write-up was incorrect.
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.
Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
- stlxp does monitor check before alignment check, use loads for all
for consistency
- Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
ARMv9.0-a hardware
- Adds product name to CPUID
- Oryon-3 being CPU PartID 2 isn't a mistake.
- No distinction between Oryon-1 and Oryon-2, both are partid 1.
cpuinfo:
```
processor : 0
BogoMIPS : 38.40
Features : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part : 0x002
CPU revision : 1
```
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.
Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the fit
script.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
Previously reported as some old CPU without AVX and SSE4 and other
things.
Start advertising as something more modern that actually shipped with
AVX2 and other features. Should help some modern libraries that do bad
family and model checks rather that CPUID features checks.
Also removes the silly `(ES)` tag from CPU-Z.
Also moves generation in to a constexpr function that can actually range
check these 4-bit and 8-bit values.
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.
Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.
Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.
Fixes#4026
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.
Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.
This lets us keep the Linux specific code in the frontend.
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.
Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.
This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.
Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.
Just check the feature bit directly in the HostRunner frontend for these
bits.
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Moves it to the hypervisor leafs.
Before:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
```
After:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-X1C
model name : Cortex-X1C
model name : Cortex-X1C
model name : Cortex-X1C
```
Now the FEX string is in the hypervisor functions as a leaf, so if some
utility wants the FEX version they can query that directly
Ex:
```bash
$ ./Bin/FEXInterpreter get_cpuid_fex
Maximum 4000_0001h sub-leaf: 2
We are running under FEX on host: 2
FEX version string is: 'FEX-2404-113-g820494d'
```
Instead of only enabling enhanced rep movs if software TSO is disabled,
Enable it if software tso is disabled OR memcpysettso is disabled. This
is because now we hit the fast path when memcpysettso is disabled alone
but global TSO is disabled.
Retested Hades and performance was fine in this configuration.
Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
Need #3348 merged first.
As I was casually thinking, this code made me realize that it was quite
branch heavy and could likely be optimized to logic.
The previous code generated some fairly nasty branch heavy code. This
can be optimized to be branchless and take roughly five instructions
per flag. Using a bitfield for each feature would turn each calculation
in to 3-4 instructions but that seems overkill.
Very minor thing.