This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.
This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.
We will need a predicated version for the next commit.
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.
Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.
While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.
Dark Souls Remastered before:
```
$RDLck Time: 4.531100 ms/second (0.04 percent)
$WRLck Time: 2.122560 ms/second (0.02 percent)
```
after:
```
$RDLck Time: 1.441620 ms/second (0.01 percent)
$WRLck Time: 0.963720 ms/second (0.01 percent)
```
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.
Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.
Adds a unittest to track this.
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.
The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.
- Windows: Uses an SRWLock instead.
- Only way for WINE to provide us a futex fallback that priorities
write-priority without stampeding.
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.
Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.
Additionally some minorly incorrect assertions.
Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.
Fixes a spurious crash in Ender Magnolia.
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.
The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.
With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
Misc resident: 54 MiB
JEMalloc resident: 208 MiB
```
So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.
Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.
This tracking will be the first step towards seeing where our memory
usage is going.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.