Commit Graph
2844 Commits
Author SHA1 Message Date
Billy Laws 9fe5eb1979 JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:35:46 +00:00
Ryan Houdek f414c92963 Code view 2025-10-28 23:53:15 +00:00
Ryan Houdek b7c7789a01 FEXCore/JIT: Supports restarting JIT in case of encoding failure
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.

We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.

Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.

This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).

But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
2025-10-28 23:53:15 +00:00
Ryan Houdek f653c5e0c0 FEXCore/JIT: Ignore local encoding limit checks
These are guaranteed not to hit encoding distance limits, so we can
ignore the returns.
2025-10-28 23:53:15 +00:00
Ryan Houdek 65fff73959 FEXCore/Dispatcher: Check encoding errors 2025-10-28 23:53:15 +00:00
Billy Laws 47619063c2 LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-10-28 23:53:12 +00:00
Billy Laws cb7076cbab FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-10-28 23:53:12 +00:00
Billy Laws 8dde79826e LookupCache: Drop unused state frame argument for delinker cbs 2025-10-28 23:53:12 +00:00
LC 46d019fe02 Merge pull request #5001 from Sonicadvance1/non_repeating_strings
FEXCore: Have non-repeat strings operations listen to non-tso config
2025-10-27 22:41:37 -04:00
Ryan Houdek 5c74d9458c FEXCore: Have non-repeat strings operations listen to non-tso config
This was missed before, where the non-repeating strings instructions
were still using TSO even when the memcpy/set config option was
disabled. Make sure it listens to the config option and disable TSO in
those instances.

Noticed this while profiling Dishonored, and WINE's `sse2_memmove`
function was showing up as a high amount of CPU time. This is due to
them using non-repeating string operations on the header and tail of
their memmove to align to 16-byte.

With this fixed, it causes the game to go from ~62FPS to ~67FPS,
becoming bottlenecked by x87 emulation instead of memmove. Doing about
23 million soft-float operations per second, because it needs full
precision to remove some flickering artifacts.
2025-10-27 13:05:54 -07:00
Ryan Houdek 1ab79bd72e FEXCore: Adds some more per-thread stats.
- Cache miss counts
  - Useful for determining if L2 cache or dynamic cache could help
- Cache read/write lock contention times
  - Useful to see if threads are blocking each other on contention
  - Read lock is the case where a read-lock is beneficial, even if we
    currently use a write lock.
- JIT count
  - Useful to see if any new JIT blocks are generating

On top of #4951 because it fiddles with the cache stuff.
2025-10-27 11:25:59 -07:00
Ryan Houdek bbb8e1ccab FEXCore: Remove ABILocalFlags hack
With our flags being optimized, this does even less than when it was
introduced. It's a hack, people are tinkering with it thinking it'll do
something. Get rid of it.
2025-10-24 17:34:57 -07:00
Ryan Houdek f44cd9c545 LookupCache: Adds an option to dynamically scale L1 cache
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.

Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)

The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.

The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.

Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.

Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.

Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.

These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.

Based on #4940 as a base which needs to be merged first.

Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
       JIT Time: 0.486630 ms/second (0.00 percent)
    Signal Time: 0.065880 ms/second (0.00 percent)
     SIGBUS Cnt: 38 (38.160780 per second)
        SMC Cnt: 0
  Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
    JIT resident:             95 mB
    OpDispatcher resident:    38 mB
    Frontend resident:        8 mB
    CPUBackend resident:      624 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 7 mB
    ThreadStates resident:    460 kB
    Unaccounted resident:     217 mB
```
2025-10-24 11:11:54 -07:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek edde5c8516 Merge pull request #4986 from Sonicadvance1/disable_trace_profiler_default
FEX: Disable trace profiler by default
2025-10-20 12:29:39 -07:00
Ryan Houdek ca697d0d5d FEX: Disable trace profiler by default
Use a config option to turn it on.
2025-10-20 10:25:17 -07:00
Ryan Houdek 854a741ea4 OpcodeDispatcher: Fix #4982
Forgot to move the OpcodeDispatcher
2025-10-17 14:01:47 -07:00
Ryan Houdek ed1952a79a Merge pull request #4983 from Sonicadvance1/detect_partial_decode
Frontend: Detect partial decoded instructions
2025-10-17 13:19:27 -07:00
Ryan Houdek 39e8f5122f Merge pull request #4982 from Sonicadvance1/fix_fex_conflict
OpcodeDispatcher: Move FEX reserved instruction
2025-10-17 11:23:09 -07:00
Ryan Houdek eb41cb2261 Frontend: Detect partial decoded instructions
Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.

Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.

Right now just change the log so we can determine if we need to support
this edge case.
2025-10-16 13:14:26 -07:00
Ryan Houdek d654f55c3c OpcodeDispatcher: Move FEX reserved instruction
This now conflicts with an SMX instruction, so move it over to another
bytecode that is unlikely to be used.
2025-10-16 11:20:58 -07:00
Paulo Matos eb1689e79f instcountci: Revert Fix quiet and signalling nan propagation 2025-10-16 09:11:01 +02:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Ryan Houdek 40a29ca9a7 FEXCore: Adds option to disable L2 cache lookups
This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.

L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.

One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.

Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!

Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.

before:
```
Total FEX Anon memory resident: 1132 mB
    JIT resident:             60 mB
    OpDispatcher resident:    97 mB
    Frontend resident:        37 mB
    CPUBackend resident:      500 kB
    Lookup cache resident:    629 mB
    Lookup L1 cache resident: 108 mB
    ThreadStates resident:    436 kB
```

after:
```
Total FEX Anon memory resident: 438 mB
    JIT resident:             62 mB
    OpDispatcher resident:    56 mB
    Frontend resident:        22 mB
    CPUBackend resident:      496 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 109 mB
    ThreadStates resident:    436 kB
```
2025-10-15 13:22:19 -07:00
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Ryan Houdek 474f2dc267 FEX: Name remaining allocations as "Misc"
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.

With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
        Misc resident:        54 MiB
    JEMalloc resident:        208 MiB
```

So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
2025-10-10 17:34:03 -07:00
Ryan Houdek 5f390c16be OpcodeDispatcher: Fixes Scalar FMA size calculation
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.

Correct the implementation to use the correct register size for FMA.
2025-10-09 15:18:26 -07:00
Lioncache c6e60ff3f5 Core: Add missing std::move in AddForceTSOInformation()
All callsites move the instructions into the function, but we weren't
further passing the rvalue-reference to merge().
2025-10-09 12:36:44 -04:00
Lioncache f5d450b95c OpcodeDispatcher: Amend wonky bitwise AND usage in LoadMemPairAutoTSO/_StoreMemPairAutoTSO 2025-10-09 11:39:17 -04:00
Lioncache b98d5f30c0 RegisterAllocationPass: Ensure relevant members are initialized
Ensures that they have deterministic values on construction
2025-10-09 04:17:21 -04:00
LC f8ff46f3e3 Merge pull request #4952 from Sonicadvance1/naming_block_links
FEXCore/fexl: Support a named monotonic_buffer_resource
2025-10-08 23:04:41 -04:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Ryan Houdek 282f091e85 FEXCore/fexl: Support a named monotonic_buffer_resource
Lets us track our memory usage of our block links.
2025-10-08 16:43:28 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Ryan Houdek 9626a64340 Merge pull request #4944 from lioncash/type
Addressing: Shave 8 bytes off AddressMode
2025-10-06 14:01:09 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Lioncache b7117b86ea X86Tables: Remove unused LateInitCopyTable() 2025-10-06 15:04:36 -04:00
Lioncache 72cb29d36b OpcodeDispatcher: Remove unimplemented function prototypes
Just cleans out the interface a little.
2025-10-06 13:08:07 -04:00
Lioncache a16d4ff1f3 OpcodeDispatcher: Deduplicate in LEAOp/SMSWOp
We can shorten a few lines here by just storing the op addr value
to a local variable.
2025-10-06 11:13:00 -04:00
Lioncache 305b1ecf2d OpcodeDispatcher: Default alignment parameters for store helpers
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.

Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
2025-10-06 11:04:12 -04:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
Ryan Houdek 45978474f3 Allocator: Name FEX's VMA regions for allocation
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.

Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.

This tracking will be the first step towards seeing where our memory
usage is going.
2025-10-05 19:30:59 -07:00
Ryan Houdek 06ff0a45a7 Merge pull request #4935 from lioncash/move
IR: Remove Swap1/Swap2 ops
2025-10-05 11:57:35 -07:00
Ryan Houdek c948d532a7 Merge pull request #4934 from lioncash/select
OpcodeDispatcher: Move off implicit _Select
2025-10-05 11:54:59 -07:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Lioncache c84801adb3 RedundantFlagCalc: Avoid vector copy in OptimizeParity()
Previously this was making a copy of the vector, when we only
need to read from it.
2025-10-05 12:17:45 -04:00