Commit Graph
38 Commits
Author SHA1 Message Date
Ryan Houdek 538a9624ec ArchHelpers/Arm64: Fixes zero register usage
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
2026-06-20 19:56:18 -07:00
Ryan Houdek c4c69ca8de ArchHelpers/Arm64: Support CAS/CASP in non-JIT SIGBUS handler
Proton was using this
2026-06-20 19:55:39 -07:00
fixedcat 6216f22cb9 Arm64: Fix byte-size handling in unaligned STLXR emulation 2026-05-19 18:53:44 +08:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Jacek Caban 12e5c60633 Arm64: Fix XZR register handling in ARM64EC unaligned STLXR emulation 2025-10-24 22:38:42 +02:00
Lioncache fd1e8d4566 Arm64: Make use of std::atomic_ref over cast 2025-10-22 11:01:50 -04:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Jacek Caban f36ac0498d Arm64: Emulate LDAXR/STLXR instructions in non-JIT ARM64EC code 2025-10-10 00:14:48 +02:00
Jacek Caban 18bdd9b665 Arm64: Factor out DoCAS 2025-10-10 00:13:57 +02:00
Jacek Caban 5fd3852fd3 ARM64EC: Emulate unaligned atomic access in non-JIT EC code 2025-10-10 00:13:56 +02:00
Ryan Houdek 16552d1194 Arm64: Remove pair usage in 128-bit loader
NFC
2025-07-28 15:58:11 -07:00
Ryan Houdek b0c61e2b69 ArchHelpers: Remove pair usage in unaligned handler
It's being treated like optional, where a value means it has been
handled, and no value means it hasn't been handled. Stop using pair in
this case.
2025-07-25 13:01:18 -07:00
Lioncache edbfe45a20 Arm64: Mark several functions as internally linked
These aren't directly used outside of the translation unit.
2025-04-16 09:16:44 -04:00
Ryan Houdek f9b369c550 Telemetry: Removes unnecessary indirection
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.

Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).

This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
2025-03-29 15:10:37 -07:00
Ryan Houdek 41e9309a36 ArchHelpers/Arm64: Fix loadstore mask
This would become an issue when multiple threads are contending with the SIGBUS handler on the same code.

We were failing to mask the VR, OPC, Rm, and Option bits, resulting in a
comparison below always resulting in a false result if another thread
managed to backpatch.

This was just unlikely to be seen on LRCPC2 supporting hardware and
since we fixed `LDSTUNSCALED_MASK` before, this wasn't really getting
seen.
2025-02-13 01:10:03 -08:00
Tony Wasserka 51d355da30 Arm64: Fix bitmask used to match load/store instructions
When multiple threads simultaneously SIGBUS on the same address, one of them
will perform the backpatching while the other will detect the backpatched
instruction sequence and hence report the SIGBUS as "handled".

This typo broke the instruction detection logic: The second thread would
assume the source of the SIGBUS was unrelated to TSO emulation and hence
report the signal as unhandled (generally triggering program abortion).

In practice, this problem did not manifest as FEX does not currently share
CodeBuffers between threads.
2025-01-30 14:56:17 +01:00
Ryan Houdek 2019f8138e ArchHelpers/Arm64: Fixes LDAPUR and STLUR backpatching
The immediate offset masking was at the completely wrong offset when I
wrote these handlers. No idea how I managed to mess those up so badly.

Should fix at least some of the issues with #4216
2024-12-19 17:29:45 -08:00
Ryan Houdek 54d332935e Arm64: Implement support for strict in-process split-locks
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.

I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.

One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.

This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
2024-08-23 16:18:19 -07:00
Ryan Houdek be9c5678c0 Arm64: Fixes SIGBUS handler for FEAT_LRCPC
When #3899 refactored some of this code, it had split an if-else chain
in to two independent if-else chains. Turns out there was a minor
dependency between them. Add back the few checks necessary to ensure
that FEAT_LRCPC falls through before hitting the FEAT_LSE checks.

This is necessary because their instructions share an encoding so the
masks between the two instruction classes overlap slightly.
2024-08-21 18:01:55 -07:00
Ryan Houdek a82fcdecd7 PR review comments 2024-08-16 10:21:18 -07:00
Ryan Houdek 27acbe305d ArchHelpers: Adjust ClearICache for its usage
Instead of clearing a hardcoded 16 bytes, adjust for the actual number
of instructions modified. The implementation will still only clear a
single cacheline so it doesn't change behaviour.
2024-08-16 10:21:18 -07:00
Ryan Houdek cd0739a534 Arm64Helpers: Moves instruction definitions to implementation
This used to exist in the FEXCore header since the unaligned handler was
done in the frontend. Once it got moved in to FEXCore it had stayed
there. Move it over now.
2024-08-16 10:21:18 -07:00
Ryan Houdek f1055d0713 Arm64: On backpatch ensure DMB instructions get patched in first
In the case of a visibility tear when one thread is backpatching while
another is executing. The executing thread can /potentially/ see the
writing of instructions depending on coherency rules or filling of
cachelines.

By ensuring the DMB instructions are backpatched over the NOP
instructions first, this ensures correct atomic visibility even on tear.
2024-08-16 10:21:18 -07:00
Ryan Houdek 7875b20594 Arm64: Handle backpatching in a thread-safe manner
When code buffers are shared between threads, FEX needs to be careful
around backpatching its code buffers, since one thread might have
backpatched the code that another thread was also planning on
backpatching.

To handle this case, when the handler fails to find a backpatchable
instruction, check if it was already backpatched. This can be determined
by atomically reading the instructions back and seeing if they have
turned in to the non-atomic variants.

In most cases we can just return saying that it has been handled, in the
case of a store we need to back the PC up 4 bytes to ensure the DMB is
executed before the non-atomic store.
2024-08-16 10:21:18 -07:00
Ryan Houdek 25679cd319 Arm64: Moves some unaligned atomic handling before mutex lock
These handlers don't do any code backpatching so locking the spinlock
futex isn't necessary. Move them before the lock to make them a bit more
efficient once code buffers get shared.
2024-08-16 10:21:18 -07:00
Ryan Houdek fa45ec32db Arm64: Early exit in unaligned handler if not on arm64
NFC, just removes a level of indent
2024-08-16 10:21:18 -07:00
Ryan Houdek 6463054fa3 Arm64: Adds another TSO hack to disable half-barrier TSO
A feature of FEX's JIT is that when an unaligned atomic load/store
operation occurs, the instructions will be backpatched in to a barrier
plus a non-atomic memory instruction. This is the half-barrier technique
that still ensures correct visibility of loadstores in an unaligned
context.

The problem with this approach is that the dmb instructions are HEAVY,
because they effectively stop the world until all memory operations in
flight are visible. But it is a necessary evil since unaligned atomics
aren't a thing on ARM processors. FEAT_LSE only gives you unaligned
atomics inside of a 16-byte granularity, which doesn't match x86
behaviour of cacheline size (effectively always 64B).

This adds a new TSO option to disable the half-barrier on unaligned
atomic and instead only convert it to a regular loadstore instruction,
ommiting the half-barrier. This gives more insight in to how well a
CPU's LRCPC implementation is by not stalling on DMB instructions when
possible.

Originally implemented as a test to see if this makes Sonic Adventure 2
run full speed with TSO enabled (but all available TSO options disabled)
on NVIDIA Orin. Unfortunately this basically makes the code no longer
stall on dmb instructions and instead just showing how bad the LRCPC
implementation is, since the stalls show up on `ldapur` instructions
instead.

Tested Sonic Adventure 2 on X13s and it ran at 60FPS there without the
hack anyway.
2024-04-24 13:09:00 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek f46e88ebdb FEXCore: Moves CPUBackend definition internal
This is no longer necessary to be part of the public API. Moves the
header internally.

Needed to pass through `IsAddressInCodeBuffer` from CPUBackend through
the Context object, but otherwise no functional change.
2024-03-29 02:27:29 -07:00
Paulo Matos e4560ed0c8 Code cleanup - mainly dead store removal; NFC
scan-build found a few dead stores that can be easily cleaned-up
2024-01-31 08:35:55 +00:00
Ryan Houdek ab6c00bbcf FEXCore/Utils: Rename FutexSpinWait to SpinWaitLock 2024-01-17 10:19:38 -08:00
Ryan Houdek e18453cb57 Jitarm64: Implements spin-loop futex for JIT blocks
This will ensure that multiple concurrent SIGBUS handlers in the same
code block doesn't modify the same code.
2024-01-17 10:19:38 -08:00
Ryan Houdek 39f49782da Arm64: Move ParanoidTSO checks up out of the non-paranoid code bath 2024-01-17 10:19:38 -08:00
Ryan Houdek 38ad3f0e05 FEXCore: Pass thread object to HandleUnalignedAccess
Currently no functional change but public API breaks should come early.

The thread state object will be used for looking up thread specific
codebuffers in the future when we support MDWE with code mirrors.
2023-12-21 01:55:25 -08:00
Ryan Houdek 9d3d33fa27 FEXCore/Utils: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00