Commit Graph
155 Commits
Author SHA1 Message Date
Ryan Houdek 168f4b1e6b Context: Add support for single-step RIP ranges
Useful when debugging a range.
2026-07-08 12:21:16 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Ryan Houdek 739e85032b FEXCore: Add support for developer single stepping, read/write watching. 2026-04-16 14:13:19 -07:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Billy Laws 064a48e965 CodeCache: Make LoadData Thread argument an optional pointer
Windows doesn't have access to Thread when loading the main image and
ntdll.
2025-12-23 23:45:07 +00:00
Tony Wasserka 8fcb84112b CodeCache: Implement runtime cache validation 2025-12-22 17:52:43 +01:00
Tony Wasserka 70eff81f19 CodeCache: Implement cache loading 2025-12-15 16:12:43 +01:00
Tony Wasserka 994613260c CodeCache: Support reverse application of relocations
This allows code to be serialized consistently across runs.
2025-12-02 22:27:53 +01:00
Tony Wasserka 33e06058c6 JIT: Move ApplyRelocations to CodeCache 2025-12-02 18:38:59 +01:00
Tony Wasserka 90cb76312c LinuxSyscalls: Implement code map writing for future code caching 2025-11-20 19:13:18 +01:00
Tony Wasserka b34b711161 CodeCache: Add interfaces to describe and generate code maps
Code maps describe per-binary metadata used to generate caches. Currently,
this includes compiled block offsets and loaded shared libraries.
2025-11-20 19:13:18 +01:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Billy Laws efd95efb40 FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-11-20 00:38:03 +00:00
Billy Laws 99ad7ea45c LookupCache: Drop unused state frame argument for delinker cbs 2025-11-20 00:38:03 +00:00
Ryan Houdek 42d0324304 FEX: Moves FEX thunk callback function generation to the frontend
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.

Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.

Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
2025-11-18 14:15:04 -08:00
Ryan Houdek 43d9384b1c FEXCore/JIT: Add a pool allocator that understands a guard page
The size asked for has its final page guarded. It's up to the code
asking for allocations to ensure it never uses the final page if
necessary.
2025-11-10 11:54:25 -08:00
Ryan Houdek bbb8e1ccab FEXCore: Remove ABILocalFlags hack
With our flags being optimized, this does even less than when it was
introduced. It's a hack, people are tinkering with it thinking it'll do
something. Get rid of it.
2025-10-24 17:34:57 -07:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Paulo Matos e7a47a647c Fix quiet and signalling nan propagation
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.

Fix for __builtin_issignaling() test of SPEC2017 classify test.
2025-09-23 13:48:45 +02:00
Lioncache e0f27bb855 Interface/Context: Remove unnecessary headers/forward declarations 2025-09-12 04:13:26 -04:00
Tony Wasserka b5b8ff01d5 CodeCache: Move LibraryJITNaming and GDBSymbols checks to Core 2025-09-11 17:03:50 +02:00
Tony Wasserka 70a1d92d9f CodeCache: Drop ComputeCodeMapId member function 2025-09-11 17:03:50 +02:00
Tony Wasserka 0749477eb9 CodeCache: Introduce revamped interfaces 2025-09-11 17:03:50 +02:00
Tony Wasserka db601d333b Core: Rename and move AOTIR.cpp and AOTIR.h
The new names better reflect the contents after recent/upcoming API changes.
2025-09-11 10:49:07 +02:00
Tony Wasserka fee9e91c4f Core: Drop support for TSO auto migration 2025-09-10 15:36:23 +02:00
Lioncache aa29201e00 Dispatcher: Reduce header dependencies
Gets rid of some unnecessary dependencies and also resolves some indirect dependencies.
2025-09-02 12:10:05 -04:00
Tony Wasserka 7b1db40d4a Config: Remove legacy code caching interfaces 2025-09-02 12:06:57 +02:00
Tony Wasserka dcaa90a855 CodeCache: Remove legacy interfaces 2025-09-02 12:06:52 +02:00
Billy Laws 8b14bd4e87 FEXCore: Fix broken RemoveCustomIREntrypoint
Unused, but would have crashed prior due to providing a nullptr thread.
2025-08-06 22:39:17 +01:00
Billy Laws e049596252 FEXCore: Use MonoBackpatcherWrite for XCHG ops in the mono backpatcher block 2025-08-06 22:39:17 +01:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Billy Laws da3b7f5a41 FEXCore: Add an option to enable future mono-specific hacks 2025-08-06 22:39:17 +01:00
Billy Laws dcd2794ff5 FEXCore: Make ThreadRemoveCodeEntryFromJit invalidate over all threads
Avoids redundant CompileBlock hits and generally easier to reason about.
2025-08-06 22:39:17 +01:00
Billy Laws 0c6fcd1678 FEXCore: Add function to recover the current block entrypoint 2025-08-06 22:39:17 +01:00
Ryan Houdek a6bb9739d4 OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2025-07-29 12:02:37 -07:00
Ryan Houdek a402c308ad FEXCore: Replace CustomIREntry tuple with struct 2025-07-25 12:33:25 -07:00
Billy Laws 44107757a3 JIT: Rewrite block linking to support direct ExitFunction calls
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.

To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset

If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.

Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.

(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.

(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).

(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).

(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).

(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).

(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
2025-07-24 14:53:09 +01:00
Billy Laws 68270ad425 LookupCache: Return whether Erase removed any cache entries 2025-07-24 14:53:09 +01:00
Billy Laws fd58f17dbe FEXCore: Track block executable ranges prior to adding cache entries
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.

This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
2025-07-24 14:52:54 +01:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Tony Wasserka 4bbaef58e9 JIT: Re-enable parallel compilation by compiling to a temporary buffer 2025-06-01 22:45:50 +02:00
Tony Wasserka 95791a985a Core: Minimize the time CodeBufferWriteMutex is held 2025-06-01 22:45:50 +02:00
Tony Wasserka 0dfefe9730 Core: Re-check LookupCache before running compiler backend
This further reduces lock contention by skipping the backend phase in case
another thread raced the active one for the same block.
2025-06-01 22:45:49 +02:00