Commit Graph
13231 Commits
Author SHA1 Message Date
Ryan Houdek 2229c04d4d LookupCache: Fixes assert
These two asserts could never fail, Add assert to the base allocation
instead.
2025-11-01 15:11:44 -07:00
Ryan Houdek 7610243b0c Merge pull request #5010 from neobrain/fix_asahi_regression
Switch back to jemalloc to fix regression in muvm-based setups
2025-11-01 13:34:36 -07:00
LC e11349b577 Merge pull request #5013 from Sonicadvance1/drm_v6.17
IoctlEmulation: Update to v6.17
2025-11-01 16:17:20 -04:00
Ryan Houdek 9b425697cb IoctlEmulation: Update to v6.17
Nova isn't handled yet because the API is in flux, but it's in v6.17 so
track it.
2025-11-01 13:04:52 -07:00
Ryan Houdek 42596ff91e External/drm-headers: Update to v6.17 2025-11-01 13:04:05 -07:00
LC b3c2ff47f3 Merge pull request #5012 from Sonicadvance1/qemu_apple
FEX: Update CPUID and detect script for newer qemu
2025-11-01 15:56:53 -04:00
Ryan Houdek 75a0bc79be FEX: Update CPUID and detect script for newer qemu
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the  fit
script.
2025-11-01 12:46:13 -07:00
Tony Wasserka 5a002ad08d Revert "Merge pull request #4969 from Sonicadvance1/rpmalloc"
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.

The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
2025-10-30 15:22:17 +01:00
Tony Wasserka fbac6f86d1 Merge pull request #5008 from neobrain/fix_removed_option
FEXConfig: Fix crash caused by no longer recognized option
2025-10-29 21:10:46 +01:00
Tony Wasserka ca18bf2a3d FEXConfig: Fix crash caused by no longer recognized option 2025-10-29 20:45:38 +01:00
Ryan Houdek 199effdff7 Merge pull request #5007 from bylaws/fasterefdfdgf
JIT: Restore behaviour of emitting interrupt checks at every block entry
2025-10-28 17:46:18 -07:00
Billy Laws 8212f4b7fb JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:34:56 +00:00
Ryan Houdek e1a45a2720 Merge pull request #4969 from Sonicadvance1/rpmalloc
Switch over to rpmalloc instead of jemalloc.
2025-10-28 17:25:25 -07:00
Ryan Houdek bd7edd8651 Merge pull request #5005 from bylaws/oodsakj
Profiler: Fix missing include
2025-10-28 17:25:02 -07:00
Billy Laws 4e1d10a46f Profiler: Fix missing include 2025-10-28 23:52:57 +00:00
LC 46d019fe02 Merge pull request #5001 from Sonicadvance1/non_repeating_strings
FEXCore: Have non-repeat strings operations listen to non-tso config
2025-10-27 22:41:37 -04:00
LC d56f689e15 Merge pull request #5002 from Sonicadvance1/fix_typo_in_the_long_long
LinuxSyscalls/Threads: Fixes typo in long jump handler
2025-10-27 22:40:03 -04:00
Ryan Houdek 90702b4102 LinuxSyscalls/Threads: Fixes typo in long jump handler
PR #4892 already found this, but since that isn't merged, make sure this
typo is fixed at least.
2025-10-27 14:05:11 -07:00
Ryan Houdek 0cf105b64a Merge pull request #4993 from Sonicadvance1/moar_stats
FEXCore: Adds some more per-thread stats.
2025-10-27 13:46:12 -07:00
Ryan Houdek 5c74d9458c FEXCore: Have non-repeat strings operations listen to non-tso config
This was missed before, where the non-repeating strings instructions
were still using TSO even when the memcpy/set config option was
disabled. Make sure it listens to the config option and disable TSO in
those instances.

Noticed this while profiling Dishonored, and WINE's `sse2_memmove`
function was showing up as a high amount of CPU time. This is due to
them using non-repeating string operations on the header and tail of
their memmove to align to 16-byte.

With this fixed, it causes the game to go from ~62FPS to ~67FPS,
becoming bottlenecked by x87 emulation instead of memmove. Doing about
23 million soft-float operations per second, because it needs full
precision to remove some flickering artifacts.
2025-10-27 13:05:54 -07:00
Ryan Houdek 75391bf834 External: Remove jemalloc (jemalloc_glibc still exists) 2025-10-27 12:05:08 -07:00
Ryan Houdek 63304a1d88 Windows: rpmalloc 2025-10-27 12:05:08 -07:00
Ryan Houdek 1ab79bd72e FEXCore: Adds some more per-thread stats.
- Cache miss counts
  - Useful for determining if L2 cache or dynamic cache could help
- Cache read/write lock contention times
  - Useful to see if threads are blocking each other on contention
  - Read lock is the case where a read-lock is beneficial, even if we
    currently use a write lock.
- JIT count
  - Useful to see if any new JIT blocks are generating

On top of #4951 because it fiddles with the cache stuff.
2025-10-27 11:25:59 -07:00
Ryan Houdek 985bdf2b6c Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2025-10-27 11:23:13 -07:00
Ryan Houdek b57ea83aea External: Add rpmalloc 2025-10-27 11:23:13 -07:00
Tony Wasserka d716e22476 Merge pull request #4997 from Sonicadvance1/delete_bad_flags
FEXCore: Remove ABILocalFlags hack
2025-10-27 09:33:35 +01:00
Ryan Houdek bbb8e1ccab FEXCore: Remove ABILocalFlags hack
With our flags being optimized, this does even less than when it was
introduced. It's a hack, people are tinkering with it thinking it'll do
something. Get rid of it.
2025-10-24 17:34:57 -07:00
Ryan Houdek 96f20779c6 Merge pull request #4951 from Sonicadvance1/dynamically_delicious
LookupCache: Adds an option to dynamically scale L1 cache
2025-10-24 14:01:18 -07:00
Ryan Houdek 908313e378 Merge pull request #4996 from cjacek/stlxr-xzr
Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation
2025-10-24 13:57:22 -07:00
Jacek Caban 12e5c60633 Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation 2025-10-24 22:38:42 +02:00
Ryan Houdek 11946ffc4c InstcountCI: Update 2025-10-24 11:11:54 -07:00
Ryan Houdek f44cd9c545 LookupCache: Adds an option to dynamically scale L1 cache
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.

Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)

The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.

The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.

Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.

Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.

Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.

These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.

Based on #4940 as a base which needs to be merged first.

Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
       JIT Time: 0.486630 ms/second (0.00 percent)
    Signal Time: 0.065880 ms/second (0.00 percent)
     SIGBUS Cnt: 38 (38.160780 per second)
        SMC Cnt: 0
  Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
    JIT resident:             95 mB
    OpDispatcher resident:    38 mB
    Frontend resident:        8 mB
    CPUBackend resident:      624 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 7 mB
    ThreadStates resident:    460 kB
    Unaccounted resident:     217 mB
```
2025-10-24 11:11:54 -07:00
Tony Wasserka 9c5ccb13de Merge pull request #4994 from lioncash/cast
SpinWaitLock, etc: Make use of std::atomic_ref over reinterpret_cast
2025-10-23 17:43:36 +02:00
Lioncache e9bcfd4784 Thread: Make use of std::atomic_ref over cast 2025-10-22 11:07:02 -04:00
Lioncache fd1e8d4566 Arm64: Make use of std::atomic_ref over cast 2025-10-22 11:01:50 -04:00
Lioncache 68dcce0739 SpinWaitLock: Make use of std::atomic_ref over cast
Has a more well-defined way of applying atomic operations to values.
2025-10-22 10:32:49 -04:00
LC 3c554cd787 Merge pull request #4992 from Sonicadvance1/Remove_the_paranoia
FEXCore: Remove Paranoid TSO mode.
2025-10-22 10:05:48 -04:00
LC 9e5f2269d9 Merge pull request #4990 from Sonicadvance1/gettls
wow64/arm64ec: Call GetTLS less frequently
2025-10-21 23:53:36 -04:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
LC eda8ca5449 Merge pull request #4984 from Sonicadvance1/shm_guaranteed_or_your_money_back
SHMStats: Add a 16-byte alignment guarantee
2025-10-21 13:46:53 -04:00
Tony Wasserka 35dd8972e2 Merge pull request #4991 from Sonicadvance1/4096
Removes some hardcoded 4096 constants
2025-10-21 18:28:06 +02:00
Ryan Houdek 643dd56b74 FEX: Removes sone hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:07 -07:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek f4eaab6977 Merge pull request #4971 from Sonicadvance1/fix_fixington 2025-10-21 06:16:43 -07:00
Ryan Houdek 900c114831 Revert "TestHarnessRunner: Avoid frontend SMC handling"
This reverts commit 2556acb82d.
2025-10-20 16:46:38 -07:00
Ryan Houdek 7937b7e52d unittests: Fixes mixture of code and data in the same page
Test behaviour themselves not changed at all, just data moved or
aligned.

For tests that aren't explicitly testing out SMC behaviour, we were
accidentally relying on some aggressive SMC tracking by mixing data and
code in the same page. To fix this just align the test's data to the
next page boundary which means FEX's SMC tracking won't get triggered
since it is no longer living in the same page.

This has been a thorn for a while, so just get rid of it. We obviously
still have ASM tests that still exist that /do/ rely on SMC, and those
are still expected to work.
2025-10-20 16:46:36 -07:00
Ryan Houdek b40e707771 wow64/arm64ec: Call GetTLS less frequently
If called back-to-back, the compiler can't optimize the object creation
resulting in multiple indirections. Save the creation and pass it
around, allowing the compiler to merge loadstores, and remove redundant
loads.

NFC
2025-10-20 15:45:31 -07:00
Ryan Houdek edde5c8516 Merge pull request #4986 from Sonicadvance1/disable_trace_profiler_default
FEX: Disable trace profiler by default
2025-10-20 12:29:39 -07:00
Ryan Houdek 197facb845 Merge pull request #4980 from neobrain/fix_infer_mapping_base
LinuxSyscalls: Fix base address inference for ELF binaries
2025-10-20 10:26:20 -07:00
Ryan Houdek ca697d0d5d FEX: Disable trace profiler by default
Use a config option to turn it on.
2025-10-20 10:25:17 -07:00