The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.
The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).
Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
- Some CMake LSPs have aneurysms when you put the end parenthesis on a
different line. Annoying? Yes, but this is all we can really do about
it for now.
- `set`, `option`, and `message` should not have spaces before their
opening parenthesis.
Signed-off-by: crueter <crueter@eden-emu.dev>
- `AddObject` doesn't need an additional Type parameter since it's
already assumed to be an object library
- The static and shared libraries don't need any explicit compile
options, as the object already handled this.
- They also don't need to be linked to FEXCore_Base. Object libraries
already handle that for us since the symbols already get pulled in
anyways.
- The object library doesn't need an output name. Only the user-facing
libraries do
Signed-off-by: crueter <crueter@eden-emu.dev>
CMake has had the `MINGW` builtin to describe MinGW targets since at
least version 3.2, so it can safely be used. This variable is also set
for the MSYS2 environments, so CLANGARM64 also correctly sets `MINGW`.
Note that this depends on https://github.com/FEX-Emu/jemalloc/pull/11.
Signed-off-by: crueter <crueter@eden-emu.dev>
Recent CPUs do nop fusion with the following instruction, this gives the
CPU the best chance to do fusion with something that actually does work.
Very trivial, doesn't do this for the more complex handling below these
as counting the number of moves before nop emitting is messy.
Instead of just a trivial pad being on or off, support a tri-state
on/off/auto where on will always pad, off will never pad, and auto will
pad only if code caching is enabled.
Further augment this by allowing a byte-width to be passed in, which can
be used with pointers to force a 48-bit VA width to only ever pad to
three instructions, reducing the common worst-case situation from 4
instructions to 3. This works because we're not going to expose a VA
width larger than 47-bit to the guest.
Fixes the handful of use-cases that explicitly chose their NOP padding,
and a bug in Arm64Relocations.cpp where it was incorrectly asking to not
receive padding even though it requires it.
When relocations are loaded, all immediates read from memory are checked
against the relocation map and transformed into an appropriately
sign-extended entrypoint-relative variant of the specific operands
addressing mode. As almost every case of an unhandled relocation will
lead to a later, likely harder to debug, crash at runtime just bail out
early if any such cases are encountered. Note that while this
handles/detects all cases of relocated immediates, if relocations were
applied to instructions themselves (occurs in some malware variants)
these would be missed without any errors reported.
When compiling code at runtime there is no harm to including jumps to
different sections within a multiblock, when enforcing as such would
introduce a lookup cost for every decode invocation (or some caching).
However when compiling offline as each cache blob is tied to a specific
library these boundaries should be enforced.
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
When loading code caches, these constants get patched up for the new guest
address. The new value may be larger than the original, so the padding bytes
ensure the maximum of 16 bytes of encoding space is always available.