Commit Graph
19 Commits
Author SHA1 Message Date
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Ryan Houdek 438501e49c Merge pull request #4998 from Sonicadvance1/i_like_my_writes_quick_and_monitored
LookupCache: Convert mutex to new WritePriorityMutex
2025-11-04 16:02:44 -08:00
Ryan Houdek c034e99aaf LookupCache: Convert mutex to new WritePriorityMutex
Changes the single highly-contended lock in `FindBlock` to be a
read-lock.
2025-11-04 15:48:52 -08:00
Ryan Houdek 2229c04d4d LookupCache: Fixes assert
These two asserts could never fail, Add assert to the base allocation
instead.
2025-11-01 15:11:44 -07:00
Ryan Houdek f44cd9c545 LookupCache: Adds an option to dynamically scale L1 cache
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.

Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)

The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.

The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.

Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.

Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.

Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.

These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.

Based on #4940 as a base which needs to be merged first.

Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
       JIT Time: 0.486630 ms/second (0.00 percent)
    Signal Time: 0.065880 ms/second (0.00 percent)
     SIGBUS Cnt: 38 (38.160780 per second)
        SMC Cnt: 0
  Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
    JIT resident:             95 mB
    OpDispatcher resident:    38 mB
    Frontend resident:        8 mB
    CPUBackend resident:      624 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 7 mB
    ThreadStates resident:    460 kB
    Unaccounted resident:     217 mB
```
2025-10-24 11:11:54 -07:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Ryan Houdek 282f091e85 FEXCore/fexl: Support a named monotonic_buffer_resource
Lets us track our memory usage of our block links.
2025-10-08 16:43:28 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Tony Wasserka 4078840ef1 Core: Reduce JIT time by sharing CodeBuffers between threads
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
  be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
  in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
  the modifying thread. Old versions of the CodeBuffer persist as read-only
  data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
  decrease the reference count and eventually trigger deallocation of the
  old version
2025-06-01 22:44:49 +02:00
Tony Wasserka 6681d7dcf9 LookupCache: Split L3 cache into a dedicated interface
This data isn't really a cache, since the JIT is directly responsible of
writing its contents. Instead it be considered the source to populate the
L1/L2 caches from.

Furthermore, splitting off this data allows it to be shared across threads
in the future without affecting L1/L2 caches.
2025-06-01 22:42:55 +02:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Billy Laws 0135e2d78c LookupCache: Use emulated overcommit if necessary
Allocates uncommited memory and calls the frontend callbacks so it can
commit pages on faults as necessary.
2024-09-10 16:08:45 +00:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek e4613477b1 FEXCore/Interface/Core: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00