mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-06 17:00:19 +02:00
Turns out I was reading six year old code for Wine's implementation for SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation. It's still write-priority but it's actually significantly slower than I was expecting due to futex queue usage and some other implementation details. Instead of using Wine's implementation, use win32's Wait/Wake on address functionality and reuse all our other mechanism for implementing this futex. This grants us our regular low-overhead codepath that I tested on Linux, while the fallback is the only "slow" path. This also allows us to still support a pseudo `WAIT_BITSET` code-path that reduces stampeding even on Win32. The reader side just waits on the upper-half of the futex (the writer bits) and the `WaitOnAddress` means only the exact match address will be woken. We also get the regular reader<->writer hand-offs working. While this path still uses the futex queue, the majority of the time our mutexes get acquired in the WFE loop already, so it's a significant win. Dark Souls Remastered before: ``` $RDLck Time: 4.531100 ms/second (0.04 percent) $WRLck Time: 2.122560 ms/second (0.02 percent) ``` after: ``` $RDLck Time: 1.441620 ms/second (0.01 percent) $WRLck Time: 0.963720 ms/second (0.01 percent) ```