Commit Graph
3076 Commits
Author SHA1 Message Date
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
Ryan Houdek 2e6a7f869c OpcodeDispatcher: Fixes nop encoded prefetch instruction
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.

Fixes `Devil May Cry 4`
2026-04-06 11:04:46 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
LC ae3fa6a836 Merge pull request #5403 from Sonicadvance1/127
IR: Adds support for printing strings
2026-03-31 16:23:42 -04:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Tony Wasserka fd6cea4698 CodeCache: Fix glibc debug mode assertion
If begin == end, the first vector::erase() call would invalidate the begin
iterator.
2026-03-22 10:11:53 +01:00
Lioncache 928a932a43 MemoryOps: Collapse duplicate add/sub in Memset
We can just use ilog2 to deduplicate this a bit.
2026-03-19 21:34:03 -04:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Ryan Houdek 68480f6e43 Merge pull request #5356 from lioncash/mops
MemoryOps: Drop MOPS handling into place for MemSet/MemCpy
2026-03-19 18:04:14 -07:00
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek 957c1fc420 Merge pull request #5357 from Sonicadvance1/106
Config: Enable Dynamic L1 and Disabled L2 caches by default
2026-03-11 14:12:38 -07:00
Ryan Houdek 86acfb35aa Config: Enable Dynamic L1 and Disabled L2 caches by default
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.

Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
2026-03-11 13:42:32 -07:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
Lioncache c8f3772753 X87Tables: Handle aliases for FSTP 2026-02-26 19:08:48 -05:00
Lioncache b046942a28 X87Tables: Handle aliases for FXCH 2026-02-26 17:44:12 -05:00
Ryan Houdek 4c49036b8a Merge pull request #5332 from lioncash/fcomp
X87Tables: Add handling for FCOMP DE D0 aliases
2026-02-26 13:49:31 -08:00
Lioncache 6909d49683 X87Tables: Add handling for FCOMP DE D0 aliases
Handles the single remaining alias for FCOMP.
2026-02-26 16:14:55 -05:00
Ryan Houdek 32471953ed Merge pull request #5226 from neobrain/fix_jit_consistency
JIT: Ensure code cache consistency by zero-initializing padding bytes
2026-02-26 13:13:10 -08:00