This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.
In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.
There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word
These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.
Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
those many contiguous bits within a single atomic word
- Allocation can fail if there are cross-word contiguous bits of the
size available
- Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
- Searches for contiguous bits of free space between multiple atomic
words.
- If found, will attempt to allocate tracking which bits were allocated
- If allocation fails, unwind bits already acquired and continue
scanning
Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
- If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
acquire semantics for scanning can technically overlook some
allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.
Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.
The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.
This used to be used for the intrusively allocated `LiveVMARegion` but
that is all handled internally to the object now, making this
unnecessary. It was always receiving zero and doing nothing so just
remove it.
Because our bitset type is uint64_t, then that means Memory + ManagedSize
is more like: Memory + (ManagedSize * 8), which is way larger of a base
than we need.
Since this was a reference, this would end up overwriting Region[0] with
whatever the smallest region was instead of just being a running pointer
to what happened to be the current smallest region.
We can switch over to a pointer to avoid obliterating the first memory
region.
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
Commit abf9724 ("Allocator: Fix and optimize VA range detection")
removed assignment to the `HostVASize` global thus making each call to
`DetermineVASize` redo all the work with potential to produce incorrect
results. Set the global again to fix this.
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.
This will speed up coredumps.
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:
```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
88 | struct prctl_mm_map {
| ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
134 | struct prctl_mm_map {
| ^
1 error generated.
```
prefer <sys/prctl.h> and do not include <linux/prctl.h>
fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.
Instead compare it to the appname that now gets passed over to it for
matching.
This allows us to do the following:
```
{
"Config": {
"ProfileStats": "1",
"X87ReducedPrecision": "1",
"TSOEnabled": "1",
"VectorTSOEnabled": "0",
"MemcpySetTSOEnabled": "0",
"HalfBarrierTSOEnabled":"1",
"MaxInst": "500",
"Multiblock": "1"
},
"AppOverrides" : {
"setup*" : {
"Comment": [
"292030 - The Witcher 3: Wild Hunt"
],
"X87ReducedPrecision": "0"
}
}
}
```
Based on #121 which needs to get merged first.
Code Review
Code Review: Class deletion
Add support for question mark and plus mark in regex, supply testing for star
Added more characters to the regex alphabets, add more test case
Added support for regex matching of configs, awaiting reviews
Rename variable to CamelCase
Addresses PR reviews
Remove unnecessary features and test cases
Rewrite to naive regex with dp
Addresses PR reviews
Build fixes
Code Review
Disables THP on some key locations that are fairly sparse
- rpmalloc
- This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
- These get in the hundreds of megabytes, while not being sparse they
trend towards only using a handful of pages and ballooning to 2MB
per thread is quite heavy.
- Lookup cache
- L1 specifically gets hit here which adds a decent chunk of overhead
due to sparsity.
Win32 for all of these also aren't handled, but that will need to be a
followup.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.
There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.
NFC
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Just my brain spinning as I try and determine what is causing some
hanging. Seems to be WINE specific so might not even be in FEX code.
Good to have some more documentation so when I read this again I don't
need to make some more logic deductions.
The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.
The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).
Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.
This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.
We will need a predicated version for the next commit.
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.
Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.
While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.
Dark Souls Remastered before:
```
$RDLck Time: 4.531100 ms/second (0.04 percent)
$WRLck Time: 2.122560 ms/second (0.02 percent)
```
after:
```
$RDLck Time: 1.441620 ms/second (0.01 percent)
$WRLck Time: 0.963720 ms/second (0.01 percent)
```