`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
Commit abf9724 ("Allocator: Fix and optimize VA range detection")
removed assignment to the `HostVASize` global thus making each call to
`DetermineVASize` redo all the work with potential to produce incorrect
results. Set the global again to fix this.
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.
This will speed up coredumps.
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:
```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
88 | struct prctl_mm_map {
| ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
134 | struct prctl_mm_map {
| ^
1 error generated.
```
prefer <sys/prctl.h> and do not include <linux/prctl.h>
fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.
Instead compare it to the appname that now gets passed over to it for
matching.
This allows us to do the following:
```
{
"Config": {
"ProfileStats": "1",
"X87ReducedPrecision": "1",
"TSOEnabled": "1",
"VectorTSOEnabled": "0",
"MemcpySetTSOEnabled": "0",
"HalfBarrierTSOEnabled":"1",
"MaxInst": "500",
"Multiblock": "1"
},
"AppOverrides" : {
"setup*" : {
"Comment": [
"292030 - The Witcher 3: Wild Hunt"
],
"X87ReducedPrecision": "0"
}
}
}
```
Based on #121 which needs to get merged first.
Code Review
Code Review: Class deletion
Add support for question mark and plus mark in regex, supply testing for star
Added more characters to the regex alphabets, add more test case
Added support for regex matching of configs, awaiting reviews
Rename variable to CamelCase
Addresses PR reviews
Remove unnecessary features and test cases
Rewrite to naive regex with dp
Addresses PR reviews
Build fixes
Code Review
Disables THP on some key locations that are fairly sparse
- rpmalloc
- This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
- These get in the hundreds of megabytes, while not being sparse they
trend towards only using a handful of pages and ballooning to 2MB
per thread is quite heavy.
- Lookup cache
- L1 specifically gets hit here which adds a decent chunk of overhead
due to sparsity.
Win32 for all of these also aren't handled, but that will need to be a
followup.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.
There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.
NFC
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Just my brain spinning as I try and determine what is causing some
hanging. Seems to be WINE specific so might not even be in FEX code.
Good to have some more documentation so when I read this again I don't
need to make some more logic deductions.
The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.
The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).
Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.
This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.
We will need a predicated version for the next commit.
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.
Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.
While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.
Dark Souls Remastered before:
```
$RDLck Time: 4.531100 ms/second (0.04 percent)
$WRLck Time: 2.122560 ms/second (0.02 percent)
```
after:
```
$RDLck Time: 1.441620 ms/second (0.01 percent)
$WRLck Time: 0.963720 ms/second (0.01 percent)
```
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.
Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.
Adds a unittest to track this.
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.
The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.
- Windows: Uses an SRWLock instead.
- Only way for WINE to provide us a futex fallback that priorities
write-priority without stampeding.
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.
Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.
Additionally some minorly incorrect assertions.
Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.
Fixes a spurious crash in Ender Magnolia.
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.
The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.