Commit Graph
60 Commits
Author SHA1 Message Date
Justin Becker 89a13cd5fd AVX-VNNI 2026-09-10 15:56:03 -07:00
Ryan Houdek 631b8c3f58 FEXCore/HostFeatures: Adds HostFeatures hashing support
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.

Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
2026-08-24 20:50:35 -07:00
Ryan Houdek 9365e6240b CPUID: Adds a few new bits
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.

Closes #5805 although their (ML?) write-up was incorrect.
2026-08-05 14:20:56 -07:00
LC c5eddd922d FEXCore: Resolve missing prototype warnings
Makes sure we mark everything internally linked as necessary, or make
declarations visible to their implementation.
2026-07-17 02:48:45 -04:00
LC 849d60253c CPUID: Fix MIDR walking in SetupHostHybridFlag() 2026-07-10 05:46:16 -04:00
Ryan Houdek 9d18ecc5cb FEXCore: Fixes a crash with multiblock if ProcessorID IR op is encountered
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.

Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
2026-07-08 17:41:03 -07:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek 8bfa6b817f CPUID: Add option for hiding hybrid Big.Little CPUs
Required for Denuvo?
2026-02-02 12:59:44 -08:00
Ryan Houdek 6b33613bb0 CPUID: Adds AmpereOneC identifier 2026-01-12 20:04:56 -08:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Ryan Houdek 94edbc3436 CPUID: Stop accidentally exposing the HTT bit
We don't support this.
2025-11-11 11:20:26 -08:00
Ryan Houdek 5eab1e559a CPUID: Fixes APICID for processor count calculation.
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.

Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
2025-11-11 11:20:26 -08:00
Ryan Houdek 75a0bc79be FEX: Update CPUID and detect script for newer qemu
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the  fit
script.
2025-11-01 12:46:13 -07:00
Ryan Houdek 3baa598b9e HostFeatures: Adds flag for SSE4a 2025-10-04 02:51:28 -07:00
Ryan Houdek 327b62ea45 CPUID: Adds new CPU product names that exist 2025-09-10 11:51:04 -07:00
Billy Laws 1e21416ccb Disable 3DNow by default on WOW64 FEX 2025-08-06 22:30:38 +01:00
Ryan Houdek 84704d1cc2 CPUID: Update documentation comments
Additional reserved bits have set uses now.
Additionally set the cpuid bit for bus-lock-detect, because FEX
definitely detects bus-locks.
2025-07-30 15:13:08 -07:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Ryan Houdek 0dff992fab FEXCore: Implement xsaveopt
It's the same as xsave because we don't track a hidden "xinuse" hardware
mask. So this is a trivial implementation.
2025-07-14 17:31:47 -07:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek 9a92f6f743 FEXCore: Remove CPUID option for SHA
Instead use the HostFeatures option that exists which is controlled by
the {enable,disable}crypto option.
2025-03-21 19:43:38 -07:00
Ryan Houdek a668492fb7 CPUID: Remove duplicated ARM Neoverse-N2
This was declared twice in the list.
2025-01-07 15:32:38 -08:00
Ryan Houdek 048e967546 FEXCore: Adds support for CPU Index through TPIDRRO 2024-10-25 15:07:57 -07:00
Ryan Houdek 09879bd962 CPUID: Update to something a little more modern
Previously reported as some old CPU without AVX and SSE4 and other
things.
Start advertising as something more modern that actually shipped with
AVX2 and other features. Should help some modern libraries that do bad
family and model checks rather that CPUID features checks.

Also removes the silly `(ES)` tag from CPU-Z.

Also moves generation in to a constexpr function that can actually range
check these 4-bit and 8-bit values.
2024-09-16 01:42:35 -07:00
Ryan Houdek fd4f6b8020 FEXCore: Dynamically scale TSC
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.

Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.

Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.

Fixes #4026
2024-09-03 13:41:31 -07:00
James Calligeros f588304b12 CPUID: add missing Apple core part numbers
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.

Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
2024-09-01 12:14:39 +10:00
Ryan Houdek 0ec724cf1f HostFeatures: Moves MIDR querying to the frontend
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.

This lets us keep the Linux specific code in the frontend.
2024-08-23 00:53:53 -07:00
Ryan Houdek 9c8438f264 FEXCore: Splits up atomic enablement checks
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.

Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.

This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.

Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
2024-08-20 14:34:34 -07:00
Ryan Houdek a1f55f0b0b HostFeatures: Removes feature flags always supported by FEX
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.

Just check the feature bit directly in the HostRunner frontend for these
bits.
2024-08-08 19:05:55 -07:00
Ryan Houdek 0653b346e0 CPUID: Adds a few missing CPU names for new CPU cores
These should be making their way to the market sooner rather than later
so make sure we have the descriptor text for them.
2024-07-07 02:40:19 -07:00
Ryan Houdek dad47b7bda CPUID: Oops, forgot to enable AVX2 2024-06-26 17:43:56 -07:00
Ryan Houdek b5e696b3cb CPUID: Implement support for XCR0 when AVX is enabled
This enables AVX, AVX2, FMA3 for the entire CPUID!

```bash
$ FEX_HOSTFEATURES=enableavx,enableavx2 ./Bin/FEXInterpreter /usr/bin/cat /proc/cpuinfo
processor       : 0
vendor_id       : GenuineIntel
cpu family      : 6
model           : 23
model name      : Cortex-A78AE
stepping        : 0
microcode       : 0x0
cpu MHz         : 3000
cache size      : 512 KB
physical id     : 0
siblings        : 12
core id         : 0
cpu cores       : 12
apicid          : 0
initial apicid  : 0
fpu             : yes
fpu_exception   : yes
cpuid level     : 22
wp              : yes
flags           : fpu vme tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht tm syscall nx mmxext fxsr_opt rdtscp lm 3dnow 3dnowext constant_tsc art rep_good nopl xtoplogy nonstop_tsc cpuid tsc_known_freq pni pclmulqdq dtes64 monitor tm2 ssse3 fma cx16 sse4_1 sse4_2 movbe popcnt aes xsave avx hypervisor lahf_lm cmp_legacy extapic abm 3dnowprefetc
h tce fsgsbase bmi1 avx2 smep bmi2 erms invpcid adx clflushopt clwb sha_ni clzero arat vpclmulqdq rdpid fsrm
bugs            :
bogomips        : 8000.0
TLB size        : 2560 4K pages
clflush size    : 64
cache_alignment  : 64
address sizes   : 40 bits physical, 48 bits virtual
power management:
```

Notice avx, avx2, and fma
2024-06-26 14:56:01 -07:00
Ryan Houdek 94fd100fc7 Merge pull request #3719 from lioncash/f16c
OpcodeDispatcher: Handle F16C operations
2024-06-26 12:12:13 -07:00
Lioncache b9ff36b5d9 CPUID: Signify F16C support if AVX is available
On Aarch64 hardware, if we have SVE2 available (which we use in the AVX implementation),
then we can also enable F16C support.
2024-06-26 15:05:03 -04:00
Ryan Houdek 45c27b2965 CPUID: Enable support for FMA3 when AVX is enabled 2024-06-25 11:24:53 -07:00
Ryan Houdek a8255aa475 CPUID: Expose support for VPCLMULQDQ
Wasn't exposed before since we couldn't unit test the SVE256
implementation.
2024-06-25 10:03:33 -04:00
Ryan Houdek e614340c0c CPUID: Update labeling on some reserved bits
These aren't reserved and I was confused that they were missing.
2024-06-21 05:34:44 -07:00
Ryan Houdek 643bc10d52 CPUID: Expose VAES if supported 2024-06-19 05:51:47 -07:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Ryan Houdek 5f0427c253 CPUID: Adds Qualcomm Oryon product name
From https://github.com/llvm/llvm-project/pull/91022

Easy enough
2024-05-03 20:16:46 -07:00
Ryan Houdek faa494c288 Merge pull request #3605 from Sonicadvance1/move_fex_versionstring_cpuid
CPUID: Removes FEX version string from CPU model name
2024-05-02 11:20:49 -07:00
Ryan Houdek 6228226c08 CPUID: Fix inverted RDTSCP check
This was inverted and always enabling the RDTSCP cpuid bit for wine.
Thus always disabling it elsewhere.
2024-05-01 18:31:41 -07:00
Ryan Houdek 31341bb7c2 CPUID: Removes FEX version string from CPU model name
Moves it to the hypervisor leafs.

Before:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
```

After:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
```

Now the FEX string is in the hypervisor functions as a leaf, so if some
utility wants the FEX version they can query that directly

Ex:
```bash
$ ./Bin/FEXInterpreter get_cpuid_fex
Maximum 4000_0001h sub-leaf: 2
We are running under FEX on host: 2
FEX version string is: 'FEX-2404-113-g820494d'
```
2024-05-01 16:27:13 -07:00
Ryan Houdek 84d5b3ee59 CPUID: Enable enhanced rep movs in more situations
Instead of only enabling enhanced rep movs if software TSO is disabled,
Enable it if software tso is disabled OR memcpysettso is disabled. This
is because now we hit the fast path when memcpysettso is disabled alone
but global TSO is disabled.

Retested Hades and performance was fine in this configuration.
2024-04-21 18:50:17 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek f79991a9d8 OpcodeDispatcher: Implement rdpid
Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
2024-03-14 20:07:58 -07:00
Ryan Houdek b902b8edab Implement small TSC scaling
Games engines are expecting >1Ghz cycle counters. Scale them to work
around the issue.

Resolves the excessive busy waiting in Unreal Engine 5 games.
2024-02-20 12:05:44 -08:00
Ryan Houdek 5d37d5db1a FEXCore: Optimize HostFeatures and CPUID feature calculation
Need #3348 merged first.

As I was casually thinking, this code made me realize that it was quite
branch heavy and could likely be optimized to logic.

The previous code generated some fairly nasty branch heavy code. This
can be optimized to be branchless and take roughly five instructions
per flag. Using a bitfield for each feature would turn each calculation
in to 3-4 instructions but that seems overkill.

Very minor thing.
2023-12-25 04:58:15 -08:00