Commit Graph
161 Commits
Author SHA1 Message Date
Ryan Houdek 2bcbfe8747 Windows: rpmalloc 2026-02-10 08:28:51 -08:00
Ryan Houdek 98617a4ba5 Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2026-02-10 08:28:51 -08:00
Ryan Houdek 45376f0dab Merge pull request #5282 from Sonicadvance1/78
FEXCore/Allocator: Disable additional VMA name attempts on failure
2026-02-09 11:43:13 -08:00
Simon Scherer 085d80792f Fix Buffer Overflow in LoadFileImpl 2026-02-05 18:42:32 +01:00
Ryan Houdek 5270983dc4 FEXCore/Allocator: Disable additional VMA name attempts on failure
Reduces spam log in strace on platforms that don't support it.
2026-02-03 19:46:00 -08:00
Ryan Houdek 0cf3108af7 FEXCore/Allocator: Slightly more verbose logs
Somehow managed to hit this in a broken setup. Add some more logs.
2026-02-03 19:43:14 -08:00
Ryan Houdek a8bee07d8f FEXCore: Move two functions to the frontend
Only used in the frontend and is OS specific.
2026-01-26 20:11:36 -08:00
Ryan Houdek 1caa9d5294 FEX: Move SBRK handling to the frontend
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.

There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.

NFC
2026-01-12 13:18:01 -08:00
Ryan Houdek 9fa8148cc6 Merge pull request #5153 from Sonicadvance1/30
WritePriorityMutex: Add some more documentation
2025-12-31 09:45:54 -08:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Ryan Houdek ba5fa35f09 WritePriorityMutex: Add some more documentation
Just my brain spinning as I try and determine what is causing some
hanging. Seems to be WINE specific so might not even be in FEX code.

Good to have some more documentation so when I read this again I don't
need to make some more logic deductions.
2025-12-29 11:04:50 -08:00
Billy Laws 2edee2855c WritePriorityMutex: Fix rare case of dropped read waiter wakes
The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.

The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).

Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
2025-12-27 22:52:24 +00:00
Ryan Houdek 2a5c1684de AllocatorHooks: Add missing header 2025-12-07 10:08:59 -08:00
Ryan Houdek f423b110a8 Utils/WritePriorityMutex: Support being forkable
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.

This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
2025-11-24 11:50:20 -08:00
Ryan Houdek 4dc1dd2511 FEXCore/Utils/SpinWaitLock: Adds WaitPred for waiting on a predicate
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.

We will need a predicated version for the next commit.
2025-11-20 14:23:01 -08:00
Ryan Houdek bddc2f227d FEXCore/Win32: Move WritePriorityMutex away from SRWLock
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.

Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.

While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.

Dark Souls Remastered before:
```
  $RDLck Time: 4.531100 ms/second (0.04 percent)
  $WRLck Time: 2.122560 ms/second (0.02 percent)
```

after:
```
  $RDLck Time: 1.441620 ms/second (0.01 percent)
  $WRLck Time: 0.963720 ms/second (0.01 percent)
```
2025-11-10 17:43:47 -08:00
Ryan Houdek 7ad7f181d7 FEXCore/LongJump: Add a way to manually load from a longjump
The frontends will need this when loading a longjump buffer in to a
context.
2025-11-07 15:33:47 -08:00
Ryan Houdek 3ccdf6508e Merge pull request #5024 from neobrain/feature_jit_encoder_recovery
FEXCore/JIT: Add support for recovering from branch encoding failures
2025-11-05 09:27:27 -08:00
Tony Wasserka 2bb64ad1c6 FEXCore/Allocator: Require caller to move unique_ptr into release workaround
This further isolates the workaround to the implementation by highlighting
at the call-site that ownership is moved away.
2025-11-05 11:29:38 +01:00
Tony Wasserka 79a685c15e FEXCore: Rename LongJump to UncheckedLongJump
This better reflects the difference to std::longjmp.
2025-11-05 10:11:30 +01:00
Ryan Houdek 8db3670ecc FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-11-05 09:47:02 +01:00
Ryan Houdek baee367532 FEXCore/Allocator: Move memory leak to a unified location 2025-11-04 16:06:05 -08:00
Ryan Houdek 8da4e72d87 FEXCore/Allocator: Fixes bug where MAP_FIXED could overallocate
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.

Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
2025-11-04 16:06:05 -08:00
Ryan Houdek b4a84a2317 Allocator: Fixes false OOM issue in allocator
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.

Adds a unittest to track this.
2025-11-04 16:06:04 -08:00
Ryan Houdek 43d6347212 FlexBitSet: Add TestAndSet helper 2025-11-04 16:06:04 -08:00
Ryan Houdek 438501e49c Merge pull request #4998 from Sonicadvance1/i_like_my_writes_quick_and_monitored
LookupCache: Convert mutex to new WritePriorityMutex
2025-11-04 16:02:44 -08:00
Ryan Houdek d2d0ca2de9 FEXCore: Implement a write-priority mutex
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.

The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.

- Windows: Uses an SRWLock instead.
  - Only way for WINE to provide us a futex fallback that priorities
    write-priority without stampeding.
2025-11-04 15:48:52 -08:00
Ryan Houdek 8430a2f7e6 FEXCore/Allocator: Fixes FlexBitSet
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.

Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.

Additionally some minorly incorrect assertions.

Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.

Fixes a spurious crash in Ender Magnolia.
2025-11-02 18:37:00 -08:00
Ryan Houdek 07cff27fa2 SpinWaitLock: Adds one-shot WFE helper 2025-11-01 13:58:22 -07:00
Ryan Houdek 2cf86998bc SpinWaitLock: Fix missing pragma 2025-11-01 13:58:22 -07:00
Tony Wasserka 5a002ad08d Revert "Merge pull request #4969 from Sonicadvance1/rpmalloc"
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.

The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
2025-10-30 15:22:17 +01:00
Ryan Houdek e1a45a2720 Merge pull request #4969 from Sonicadvance1/rpmalloc
Switch over to rpmalloc instead of jemalloc.
2025-10-28 17:25:25 -07:00
Billy Laws 4e1d10a46f Profiler: Fix missing include 2025-10-28 23:52:57 +00:00
Ryan Houdek 63304a1d88 Windows: rpmalloc 2025-10-27 12:05:08 -07:00
Ryan Houdek 985bdf2b6c Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2025-10-27 11:23:13 -07:00
Jacek Caban 12e5c60633 Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation 2025-10-24 22:38:42 +02:00
Lioncache fd1e8d4566 Arm64: Make use of std::atomic_ref over cast 2025-10-22 11:01:50 -04:00
Lioncache 68dcce0739 SpinWaitLock: Make use of std::atomic_ref over cast
Has a more well-defined way of applying atomic operations to values.
2025-10-22 10:32:49 -04:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Ryan Houdek ca697d0d5d FEX: Disable trace profiler by default
Use a config option to turn it on.
2025-10-20 10:25:17 -07:00
Ryan Houdek 474f2dc267 FEX: Name remaining allocations as "Misc"
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.

With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
        Misc resident:        54 MiB
    JEMalloc resident:        208 MiB
```

So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
2025-10-10 17:34:03 -07:00
Jacek Caban f36ac0498d Arm64: Emulate LDAXR/STLXR instructions in non-JIT ARM64EC code 2025-10-10 00:14:48 +02:00
Jacek Caban 18bdd9b665 Arm64: Factor out DoCAS 2025-10-10 00:13:57 +02:00
Jacek Caban 5fd3852fd3 ARM64EC: Emulate unaligned atomic access in non-JIT EC code 2025-10-10 00:13:56 +02:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
Ryan Houdek 45978474f3 Allocator: Name FEX's VMA regions for allocation
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.

Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.

This tracking will be the first step towards seeing where our memory
usage is going.
2025-10-05 19:30:59 -07:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Ryan Houdek 1121f2a1fb Merge pull request #4853 from lioncash/allocator
FEXCore/Allocator: Remove unused headers
2025-09-08 20:25:25 -07:00