Commit Graph
490 Commits
Author SHA1 Message Date
Ryan Houdek 41241d7500 Wow64: Spin loop on atomic with WFE
Instead of burning roughly a million watts, put this spinloop on a WFE.
This tends to occur on a crash during shutdown that isn't fully able to
be avoided. The least we can do is not consume all the power in the
world.
2026-07-07 16:09:00 -07:00
Egor Lazarchuk 1eb5abb8db Allocator: rename DetermineVASize to GetHostVABits
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
2026-07-05 12:46:41 +01:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00
Ryan Houdek 7d1c625e32 FEXCore: Allow InterruptFaultPage to be significantly further away
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.

Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
2026-05-18 15:57:25 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
Tony Wasserka c6d2ce043f Merge pull request #5388 from Sonicadvance1/123
Config: Finish wiring up Regex app overrides
2026-04-02 11:00:37 +02:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
badumbatish 81d4e8fe9d Initial implementation for regex engine
Add support for question mark and plus mark in regex, supply testing for star

Added more characters to the regex alphabets, add more test case

Added support for regex matching of configs, awaiting reviews

Rename variable to CamelCase

Addresses PR reviews

Remove unnecessary features and test cases

Rewrite to naive regex with dp

Addresses PR reviews

Build fixes

Code Review
2026-04-01 10:45:59 -07:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
crueter 33a01f2773 Core: Use long commit hash instead of abbreviating
Fixes #5230

Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.

Note: I have no idea how fmt will handle that format string with an
array of uint8_t.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-26 10:08:41 -05:00
Tony Wasserka 35e4d452c0 Config: Remove unused function from the old code cache skeleton implementation 2026-02-23 17:49:15 +01:00
Ryan Houdek 19ff098db5 FEXCore: Ensure that CpuStateFrame is 64-byte aligned
This is our cacheline and CLZero line size. Need to ensure it is
aligned.
2026-02-21 15:11:47 -08:00
Ryan Houdek b130d14de4 CoreState: Add some more alignment checks
Just to ensure we don't hit performance penalties of some CPUs.
We already weren't, but make sure.
2026-02-20 12:57:21 -08:00
Tony Wasserka 44a6bfd874 Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
CodeCache.h has no dependency on SourcecodeResolver.h.

This reverts commit f2bbc0eccd.
2026-02-16 10:29:58 +01:00
Ryan Houdek 98617a4ba5 Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2026-02-10 08:28:51 -08:00
LC a5ceb2a61d Merge pull request #5270 from Sonicadvance1/69
FEXCore: Move two functions to the frontend
2026-01-28 16:01:35 -05:00
Paulo Matos eb6c050cf0 FEXCore: Fix initial PF flag value
Initialize pf_raw to 1 instead of 0 so that the reconstructed Parity
Flag matches x86 reset state (PF=0).
2026-01-27 11:25:08 +01:00
Ryan Houdek a8bee07d8f FEXCore: Move two functions to the frontend
Only used in the frontend and is OS specific.
2026-01-26 20:11:36 -08:00
crueter 884b96c300 [externals] use ankerl::unordered_dense over tsl
ankerl::unordered_dense is faster on average and has less memory usage
than tsl::robin_map. It is pretty significantly faster than std but
we'll keep that as is for now.

Obviously, this will need a lot of testing.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-15 14:07:08 -05:00
Ryan Houdek 1caa9d5294 FEX: Move SBRK handling to the frontend
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.

There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.

NFC
2026-01-12 13:18:01 -08:00
Ryan Houdek d7c2b8f513 FEXCore: Cleanup pointers structure
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.

NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
2026-01-07 12:55:34 -08:00
Ryan Houdek 817a927e31 FEXCore: Fixes circular dependency with thunk callback
Fixes crash in thunks that use callbacks, introduced in #5148.

The dispatcher would call the syscallhandler to get the VDSO thunk
callback. But due to reordering initialization, the VDSO thunk would
have not been loaded at that point. This would cause thunks that use
callbacks to crash with a nullptr exception.

Instead, defer the thunk callback pointer loading until the thread
starts executing, and load the pointer in to our thread state's pointer
struct instead.

Didn't get caught in my initial test sweep since I didn't run a Wine
game with thunks.
2026-01-07 12:03:19 -08:00
Billy Laws 6b583ee697 Relocations: Switch to robin_map to improve lookup perf 2025-12-31 15:12:05 +00:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Billy Laws f2bbc0eccd CodeCache: Use defaulted dtor for ExecutableFileInfo 2025-12-28 00:30:47 +00:00