- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.
Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the fit
script.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
Previously reported as some old CPU without AVX and SSE4 and other
things.
Start advertising as something more modern that actually shipped with
AVX2 and other features. Should help some modern libraries that do bad
family and model checks rather that CPUID features checks.
Also removes the silly `(ES)` tag from CPU-Z.
Also moves generation in to a constexpr function that can actually range
check these 4-bit and 8-bit values.
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.
Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.
Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.
Fixes#4026
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.
Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.
This lets us keep the Linux specific code in the frontend.
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.
Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.
This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.
Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.
Just check the feature bit directly in the HostRunner frontend for these
bits.
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Moves it to the hypervisor leafs.
Before:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-A78C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
model name : FEX-2404-101-gf9effcb Cortex-X1C
```
After:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-A78C
model name : Cortex-X1C
model name : Cortex-X1C
model name : Cortex-X1C
model name : Cortex-X1C
```
Now the FEX string is in the hypervisor functions as a leaf, so if some
utility wants the FEX version they can query that directly
Ex:
```bash
$ ./Bin/FEXInterpreter get_cpuid_fex
Maximum 4000_0001h sub-leaf: 2
We are running under FEX on host: 2
FEX version string is: 'FEX-2404-113-g820494d'
```
Instead of only enabling enhanced rep movs if software TSO is disabled,
Enable it if software tso is disabled OR memcpysettso is disabled. This
is because now we hit the fast path when memcpysettso is disabled alone
but global TSO is disabled.
Retested Hades and performance was fine in this configuration.
Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
Need #3348 merged first.
As I was casually thinking, this code made me realize that it was quite
branch heavy and could likely be optimized to logic.
The previous code generated some fairly nasty branch heavy code. This
can be optimized to be branchless and take roughly five instructions
per flag. Using a bitfield for each feature would turn each calculation
in to 3-4 instructions but that seems overkill.
Very minor thing.
No need to wait for initialization on for this anymore.
Ever since Init was refactored to do basically no work, this hasn't been
necessary.
CPUID does need to still be initialized after HostFeatures though, so
need to ensure correct member ordering there.
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.
Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
Hades and the vcruntime hits this very hard in memmove.
`86.56% [JIT] tid 458574 [.] JIT_0x18000c375_0x7fffc94790c8`
```asm
0x00007fffc94790f8: ldaprb w3, [x2]
0x00007fffc94790fc: stlrb w3, [x1]
0x00007fffc9479100: add x1, x1, #0x1
0x00007fffc9479104: add x2, x2, #0x1
0x00007fffc9479108: sub x0, x0, #0x1
0x00007fffc947910c: cbnz x0, 0x7fffc94790f8
```
This performance is terrible because Cortex's LRCPC performance is bottom-tier.
Work around the performance issue by forcing things to do larger moves with vector moves instead.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>