Commit Graph
407 Commits
Author SHA1 Message Date
Ryan Houdek f414c92963 Code view 2025-10-28 23:53:15 +00:00
Ryan Houdek 150bf7b30c FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-10-28 23:53:15 +00:00
Billy Laws 47619063c2 LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-10-28 23:53:12 +00:00
Ryan Houdek 1ab79bd72e FEXCore: Adds some more per-thread stats.
- Cache miss counts
  - Useful for determining if L2 cache or dynamic cache could help
- Cache read/write lock contention times
  - Useful to see if threads are blocking each other on contention
  - Read lock is the case where a read-lock is beneficial, even if we
    currently use a write lock.
- JIT count
  - Useful to see if any new JIT blocks are generating

On top of #4951 because it fiddles with the cache stuff.
2025-10-27 11:25:59 -07:00
Ryan Houdek f44cd9c545 LookupCache: Adds an option to dynamically scale L1 cache
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.

Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)

The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.

The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.

Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.

Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.

Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.

These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.

Based on #4940 as a base which needs to be merged first.

Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
       JIT Time: 0.486630 ms/second (0.00 percent)
    Signal Time: 0.065880 ms/second (0.00 percent)
     SIGBUS Cnt: 38 (38.160780 per second)
        SMC Cnt: 0
  Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
    JIT resident:             95 mB
    OpDispatcher resident:    38 mB
    Frontend resident:        8 mB
    CPUBackend resident:      624 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 7 mB
    ThreadStates resident:    460 kB
    Unaccounted resident:     217 mB
```
2025-10-24 11:11:54 -07:00
LC 3c554cd787 Merge pull request #4992 from Sonicadvance1/Remove_the_paranoia
FEXCore: Remove Paranoid TSO mode.
2025-10-22 10:05:48 -04:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
LC eda8ca5449 Merge pull request #4984 from Sonicadvance1/shm_guaranteed_or_your_money_back
SHMStats: Add a 16-byte alignment guarantee
2025-10-21 13:46:53 -04:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek 6ee77eb1a1 SHMStats: Adds ThreadStats size to header
Reused the padding area so the header format doesn't change. If it is
non-zero then it should be used by the tool.
2025-10-16 14:07:29 -07:00
Ryan Houdek 46fc45b952 SHMStats: Add a 16-byte alignment guarantee
We want to take advantage of 16-byte single-copy atomicity. Which I am
relying on, but didn't codify it the first time.

Additionally add comments to explain that new members should be added to
the end to allow tools time to gain support gradually. This will allow
me to add new members without fully breaking mangohud, they'll just not
display the new information until support is added.

We're not guaranteeing backwards compatibility, just an attempt not to
constantly churn the format unless necessary. This way if we do break
compatibility, the tool will have an upper bound on supported versions
before needing to rewrite code.
2025-10-16 13:58:31 -07:00
Jacek Caban f36ac0498d Arm64: Emulate LDAXR/STLXR instructions in non-JIT ARM64EC code 2025-10-10 00:14:48 +02:00
Jacek Caban 5fd3852fd3 ARM64EC: Emulate unaligned atomic access in non-JIT EC code 2025-10-10 00:13:56 +02:00
LC f8ff46f3e3 Merge pull request #4952 from Sonicadvance1/naming_block_links
FEXCore/fexl: Support a named monotonic_buffer_resource
2025-10-08 23:04:41 -04:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek 282f091e85 FEXCore/fexl: Support a named monotonic_buffer_resource
Lets us track our memory usage of our block links.
2025-10-08 16:43:28 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Lioncache d89c119d9d EnumUtils: Remove fmt include
Forgot to remove this in the previous PR.
2025-10-07 12:54:32 -04:00
Ryan Houdek 802f3bed47 Merge pull request #4947 from lioncash/format
EnumUtils: Further simplify enum passthrough formatting
2025-10-07 06:47:56 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Lioncache d94b9fae94 SourceCodeResolver: Pass string_view by value for GenerateMap()
Generally this should be passed as a value type unless there's a good
reason not to.
2025-10-07 01:27:15 -04:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
Ryan Houdek 3baa598b9e HostFeatures: Adds flag for SSE4a 2025-10-04 02:51:28 -07:00
Lioncache a0c2ce0068 EnumUtils: Add define for default passthrough formatting
Handles a normal case where printing an enum type as an integral value
is still desirable.

Mainly just a way to reduce boilerplate.
2025-10-03 14:21:19 -04:00
Lioncache 569d7297f0 fextl/string: Correct <filesystem> include to <functional>
This was unintentionally putting all the filesystem utilities into headers implicitly.
2025-09-29 23:54:08 -04:00
Ryan Houdek 61719115e5 Merge pull request #4817 from neobrain/refactor_code_cache_new_interfaces
CodeCache: Introduce new interfaces
2025-09-11 12:59:33 -07:00
Tony Wasserka 70a1d92d9f CodeCache: Drop ComputeCodeMapId member function 2025-09-11 17:03:50 +02:00
Tony Wasserka 0749477eb9 CodeCache: Introduce revamped interfaces 2025-09-11 17:03:50 +02:00
Tony Wasserka 48787ab460 Context: Drop unneeded header includes 2025-09-11 10:49:07 +02:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Tony Wasserka fee9e91c4f Core: Drop support for TSO auto migration 2025-09-10 15:36:23 +02:00
Ryan Houdek 1121f2a1fb Merge pull request #4853 from lioncash/allocator
FEXCore/Allocator: Remove unused headers
2025-09-08 20:25:25 -07:00
Lioncache a7989eb79f FEXCore/Allocator: Remove unused headers
Reveals some more indirect inclusions.
2025-09-08 22:53:09 -04:00
Lioncache b2407352a9 IR: Remove unnecessary forward declarations/headers
Pares the base IR header down to a modest size and also reveals a few indirect
reliances on the allocator facilities.
2025-09-08 22:03:25 -04:00
Ryan Houdek 09793d295f Merge pull request #4849 from lioncash/config
Config: Remove unused Context.h include
2025-09-08 16:59:38 -07:00
Lioncache f69821d7db Config: Remove unused Context.h include
Removes quite a heavy include from the config system and specifies any
indirect inclusion that were relied on because of it.
2025-09-08 15:23:30 -04:00
Lioncache 6b2363d3f6 InternalThreadState: Remove unused headers
Removes some includes in the core header and resolves indirect includes elsewhere.
2025-09-08 14:13:01 -04:00
Ryan Houdek e76af56c86 Merge pull request #4845 from lioncash/thunk
Core/Thunks: Remove unused IR header/forward declarations
2025-09-08 10:03:34 -07:00
Ryan Houdek b7173a3ae2 Merge pull request #4844 from lioncash/builtin
General: std::alignment_of -> alignof
2025-09-08 10:03:24 -07:00
Lioncache 9b9ab16b55 Common: Remove duplicated StringUtil functionality
We already have an equivalent header within FEXCore that's header only,
so we can adapt it to conform for both cases, allowing for removal of
one of them.
2025-09-08 09:44:25 -04:00
Lioncache d36c387c1f Core/Thunks: Remove unused IR/forward declarations
Moves the IR include into one of the more specific headers, which avoids dumping the IR header into any core bits that use the interface.

Also uncovered a missing header guard.
2025-09-06 22:28:27 -04:00
Lioncache 7e61b172b2 General: std::alignment_of -> alignof
We can just use the compiler built-in in these cases. std::alignment_of
mainly has utility in metaprogramming.
2025-09-06 21:23:10 -04:00
Ryan Houdek ea8ae2bf2c Merge pull request #4839 from lioncash/utils
Common Utils: Remove unused/unnecessary inclusions from headers
2025-09-06 15:50:00 -07:00
Lioncache 2613762ac2 InternalThreadState: Make bool operator explicit
We definitely don't want the boolean null test to be able to be implicitly converted
(e.g. to an int or whatever else).

For example a non-explicit bool operator allows for silly things like:

NonMovableUniquePtr<...> ptr;
// ...
auto k = 5 + ptr;

to build without issue, which we should really force the user to be explicit about if it's *really* a desired behavior.
2025-09-05 16:29:19 -04:00
Lioncache 62f16a3c4b FEXServerClient: Remove unnecessary includes from header
We can reduce includes, such as logging by specifying a concrete size for logging levels,
allowing the enum to be forward declared. We can also move FillHeader into the cpp file,
allowing the syscalls header to be removed.
2025-09-05 13:45:31 -04:00
Lioncache dd735464a2 Threads: Remove unnecessary includes
We have quite a few unnecessary includes here left over from code movement,
so we can clean those out.

Also fix an indirect include in the header.
2025-09-03 11:00:57 -04:00
Tony Wasserka 7b1db40d4a Config: Remove legacy code caching interfaces 2025-09-02 12:06:57 +02:00
Tony Wasserka dcaa90a855 CodeCache: Remove legacy interfaces 2025-09-02 12:06:52 +02:00
Lioncache fde99a8dc1 Dispatcher: Centralize SignalDelegator config creation
Since all of the information comes from the Dispatcher, we can have the dispatcher
provide that information. This way we can also eliminate a bunch of now-redundant
public interface members and simplify the config setup within InitCoreImpl().

Conveniently, this also allows making all members of the dispatcher non-public.
2025-08-29 11:19:07 -04:00
Lioncache 2e692a1146 SignalDelegator: Move struct outside of SignalDelegator
The name itself is already qualified with SignalDelegator, so this can reasonably be outside the class itself.
This also allows for forward declarations of the config struct
2025-08-29 10:23:03 -04:00