The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
This would become an issue when multiple threads are contending with the SIGBUS handler on the same code.
We were failing to mask the VR, OPC, Rm, and Option bits, resulting in a
comparison below always resulting in a false result if another thread
managed to backpatch.
This was just unlikely to be seen on LRCPC2 supporting hardware and
since we fixed `LDSTUNSCALED_MASK` before, this wasn't really getting
seen.
When multiple threads simultaneously SIGBUS on the same address, one of them
will perform the backpatching while the other will detect the backpatched
instruction sequence and hence report the SIGBUS as "handled".
This typo broke the instruction detection logic: The second thread would
assume the source of the SIGBUS was unrelated to TSO emulation and hence
report the signal as unhandled (generally triggering program abortion).
In practice, this problem did not manifest as FEX does not currently share
CodeBuffers between threads.
The immediate offset masking was at the completely wrong offset when I
wrote these handlers. No idea how I managed to mess those up so badly.
Should fix at least some of the issues with #4216
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.
I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.
One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.
This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
When #3899 refactored some of this code, it had split an if-else chain
in to two independent if-else chains. Turns out there was a minor
dependency between them. Add back the few checks necessary to ensure
that FEAT_LRCPC falls through before hitting the FEAT_LSE checks.
This is necessary because their instructions share an encoding so the
masks between the two instruction classes overlap slightly.
Instead of clearing a hardcoded 16 bytes, adjust for the actual number
of instructions modified. The implementation will still only clear a
single cacheline so it doesn't change behaviour.
This used to exist in the FEXCore header since the unaligned handler was
done in the frontend. Once it got moved in to FEXCore it had stayed
there. Move it over now.
In the case of a visibility tear when one thread is backpatching while
another is executing. The executing thread can /potentially/ see the
writing of instructions depending on coherency rules or filling of
cachelines.
By ensuring the DMB instructions are backpatched over the NOP
instructions first, this ensures correct atomic visibility even on tear.
When code buffers are shared between threads, FEX needs to be careful
around backpatching its code buffers, since one thread might have
backpatched the code that another thread was also planning on
backpatching.
To handle this case, when the handler fails to find a backpatchable
instruction, check if it was already backpatched. This can be determined
by atomically reading the instructions back and seeing if they have
turned in to the non-atomic variants.
In most cases we can just return saying that it has been handled, in the
case of a store we need to back the PC up 4 bytes to ensure the DMB is
executed before the non-atomic store.
These handlers don't do any code backpatching so locking the spinlock
futex isn't necessary. Move them before the lock to make them a bit more
efficient once code buffers get shared.
A feature of FEX's JIT is that when an unaligned atomic load/store
operation occurs, the instructions will be backpatched in to a barrier
plus a non-atomic memory instruction. This is the half-barrier technique
that still ensures correct visibility of loadstores in an unaligned
context.
The problem with this approach is that the dmb instructions are HEAVY,
because they effectively stop the world until all memory operations in
flight are visible. But it is a necessary evil since unaligned atomics
aren't a thing on ARM processors. FEAT_LSE only gives you unaligned
atomics inside of a 16-byte granularity, which doesn't match x86
behaviour of cacheline size (effectively always 64B).
This adds a new TSO option to disable the half-barrier on unaligned
atomic and instead only convert it to a regular loadstore instruction,
ommiting the half-barrier. This gives more insight in to how well a
CPU's LRCPC implementation is by not stalling on DMB instructions when
possible.
Originally implemented as a test to see if this makes Sonic Adventure 2
run full speed with TSO enabled (but all available TSO options disabled)
on NVIDIA Orin. Unfortunately this basically makes the code no longer
stall on dmb instructions and instead just showing how bad the LRCPC
implementation is, since the stalls show up on `ldapur` instructions
instead.
Tested Sonic Adventure 2 on X13s and it ran at 60FPS there without the
hack anyway.
This is no longer necessary to be part of the public API. Moves the
header internally.
Needed to pass through `IsAddressInCodeBuffer` from CPUBackend through
the Context object, but otherwise no functional change.
Currently no functional change but public API breaks should come early.
The thread state object will be used for looking up thread specific
codebuffers in the future when we support MDWE with code mirrors.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>