- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Just my brain spinning as I try and determine what is causing some
hanging. Seems to be WINE specific so might not even be in FEX code.
Good to have some more documentation so when I read this again I don't
need to make some more logic deductions.
The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.
The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).
Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.
This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.
Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.
While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.
Dark Souls Remastered before:
```
$RDLck Time: 4.531100 ms/second (0.04 percent)
$WRLck Time: 2.122560 ms/second (0.02 percent)
```
after:
```
$RDLck Time: 1.441620 ms/second (0.01 percent)
$WRLck Time: 0.963720 ms/second (0.01 percent)
```
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.
The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.
- Windows: Uses an SRWLock instead.
- Only way for WINE to provide us a futex fallback that priorities
write-priority without stampeding.