Commit Graph
537 Commits
Author SHA1 Message Date
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Ryan Houdek a499ad6404 IR: Implement two new IR ops
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
2025-10-04 02:43:39 -07:00
Lioncache 48c6acbc87 IR: Move RegisterAllocationPass forward decl to JITClass
This isn't actually used anywhere in the IR header, so we can
move it to where it's actually used.

Now the IR interface header doesn't have anything related to the
independent passes in it.
2025-10-04 00:08:58 -04:00
Ryan Houdek 012c2ba851 IR: Fixes vector 64-bit binops
We were only supporting operating size of 256-bit and 128-bit. These
also support 64-bit which wasn't wired up.

SSE4a will want to use 64-bit.
2025-10-03 13:04:09 -07:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Lioncache 863d1e0007 IR: Convert FenceType to an enum class
Same thing minus an extra struct lingering around.
2025-10-01 23:13:08 -04:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Lioncache 84325d6b5c Arm64Emitter: Cull unnecessary includes
Also fixes an indirect include.
2025-09-30 09:32:25 -04:00
Lioncache 6b61093037 JITClass: Mark functions as static where applicable
These don't depend on any class state.
2025-09-29 23:42:43 -04:00
Ryan Houdek 7f8fffbb24 FEXCore: Implement support for AVX gathers with overflow
Falls down the emulated path so we can zero extend the address
calculation. No way for SVE gathers to use 32-bit addressing.
2025-09-29 16:06:02 -07:00
Lioncache e0f27bb855 Interface/Context: Remove unnecessary headers/forward declarations 2025-09-12 04:13:26 -04:00
Ryan Houdek 61719115e5 Merge pull request #4817 from neobrain/refactor_code_cache_new_interfaces
CodeCache: Introduce new interfaces
2025-09-11 12:59:33 -07:00
Ryan Houdek 0eedd55dfd Merge pull request #4866 from neobrain/refactor_reformat
Update code formatting
2025-09-11 10:13:42 -07:00
Ryan Houdek 4a9170e981 Merge pull request #4837 from bylaws/x87opt
X87: Better cache intermediate results in the X87 slowpath
2025-09-11 10:12:36 -07:00
Tony Wasserka 0749477eb9 CodeCache: Introduce revamped interfaces 2025-09-11 17:03:50 +02:00
Tony Wasserka db601d333b Core: Rename and move AOTIR.cpp and AOTIR.h
The new names better reflect the contents after recent/upcoming API changes.
2025-09-11 10:49:07 +02:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Billy Laws d2714f3338 JIT: Check for suspend interrupts at back-edges
Avoids most cases where an infinite loop could lead to suspend interrups
never getting checked.
2025-09-09 22:35:52 +01:00
Billy Laws b4e556e63a IR: Introduce operation to form a scaled address from STATE 2025-09-09 21:29:12 +01:00
Lioncache a7989eb79f FEXCore/Allocator: Remove unused headers
Reveals some more indirect inclusions.
2025-09-08 22:53:09 -04:00
Paulo Matos 14256e7a90 Fix mapping for VS and VC condition codes
This doesn't seem to be generated in-code therefore it's not fixing any
existing bug, but fixes the mapping for future uses of the condition code.
2025-09-08 11:41:15 +02:00
Ryan Houdek 6ae94581bb Merge pull request #4821 from bylaws/monof
Frontend: Fix tailcall handling when mono hacks are enabled
2025-09-06 16:51:08 -07:00
Ryan Houdek 343a529e63 Merge pull request #4824 from bylaws/tf2
OpDispatcher: Force a recheck of TF after POPF
2025-09-03 15:24:40 -07:00
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Billy Laws d72e121fd4 AtomicOps: Reimplement AtomicNeg using 8.1 CAS atomics 2025-09-02 22:23:14 +01:00
Billy Laws 2c883d7cdc OpDispatcher: Force a recheck of TF after POPF
The dispatcher/block linker will handle this, but if the instruction
following a POPF flag doesn't otherwise trigger one of those the
interrupt would be missed.
2025-09-02 22:21:42 +01:00
Tony Wasserka 7b1db40d4a Config: Remove legacy code caching interfaces 2025-09-02 12:06:57 +02:00
Tony Wasserka dcaa90a855 CodeCache: Remove legacy interfaces 2025-09-02 12:06:52 +02:00
Tony Wasserka 51e64c69f3 LogManager: Unconditionally evaluate assertion conditions
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
2025-08-25 10:36:01 +02:00
Ryan Houdek c9a33c638a Dispatcher: Most minor of optimizations for f64 x87
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
2025-08-17 19:50:21 -07:00
Alyssa Rosenzweig 2a00f23459 JIT: replace JumpTargets map with vector
Hashmaps are super expensive and there's no reason not to use a vector - we
already have compact block IDs so we don't benefit from the sparseness. Huge win
for very little effort.

Spotted when profiling FEX. CondJump() in the JIT was almost 4% of our time (?!)
and all because of map slowness. Easy fix.

Difference at 95.0% confidence
	-0.0196494 +/- 0.00194956
	-3.92827% +/- 0.389753%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Alyssa Rosenzweig c58ad8b593 IR: add AndShift op
will use it for next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:09:20 -04:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Billy Laws 3497870a45 JIT: Guard lookupcache locks with the code invalidation mutex
Avoids issues with forking, as the code invalidation mutex is fork-safe.
2025-07-24 14:53:09 +01:00
Billy Laws 44107757a3 JIT: Rewrite block linking to support direct ExitFunction calls
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.

To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset

If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.

Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.

(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.

(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).

(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).

(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).

(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).

(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
2025-07-24 14:53:09 +01:00
Billy Laws 45ba1af388 BranchOps: Use the call-ret stack to optimise indirect ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws ba9884a26a JIT: Emit entrypoint code for call return target blocks
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
2025-07-24 14:53:09 +01:00
Billy Laws d67b1645c9 FEXCore: Hold a frontend allocation for the call-ret stack
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
2025-07-24 14:53:09 +01:00
Billy Laws 963a8c2f08 FEXCore: Save and restore the call/ret SP from CPUState
For simplicity in cases like signal handling, always load it in
fill and store in spill, even the SP is stored in a callee save
register.
2025-07-24 14:53:09 +01:00
Ryan Houdek 7e1ee5bb07 FEXCore: Reintroduce support for CSSC
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.

Been a while since I last looked at this, added a new instcountci file
to show the improvement.
2025-07-17 15:09:27 -07:00
Paulo Matos 5267cde60e Whole-tree reformat with clang-format-19 2025-07-17 08:10:00 +02:00
Ryan Houdek 7edc8417b6 JIT: Add more padding
Go to a whole page of additional padding, #4670 adds some more size to a
block and overran the padding causing a unittest to fail.
2025-07-16 15:28:05 -07:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Lioncache 678e03b470 VectorOps: Correct benign VFAddV op cast
This was using the non-float VAddV variant, but had the same behavior,
since the fields were named the same. So this is just a correctness fix.
2025-07-15 11:41:28 -04:00