Commit Graph
420 Commits
Author SHA1 Message Date
Ryan Houdek 217bbf423b FEXCore: Switch constant emission to default to NoPad
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.

The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
2025-12-29 11:45:51 -08:00
Ryan Houdek a1a30cd9a6 IR: Remove default argument for padding in Constant op
All use cases now pass a pad type in to this.
2025-12-29 11:29:01 -08:00
Ryan Houdek cb432548bf IR: Allow passing padding and MaxBytes through Constant IR 2025-12-29 11:04:34 -08:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Ryan Houdek a499ad6404 IR: Implement two new IR ops
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
2025-10-04 02:43:39 -07:00
Ryan Houdek 012c2ba851 IR: Fixes vector 64-bit binops
We were only supporting operating size of 256-bit and 128-bit. These
also support 64-bit which wasn't wired up.

SSE4a will want to use 64-bit.
2025-10-03 13:04:09 -07:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Lioncache e37996873a IRDumper: Add missing FenceType entry 2025-10-02 02:55:34 -04:00
Lioncache d65c54ba21 IR: Remove TypeDefinition struct
These aren't used at all.
2025-10-02 00:50:51 -04:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Lioncache 863d1e0007 IR: Convert FenceType to an enum class
Same thing minus an extra struct lingering around.
2025-10-01 23:13:08 -04:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Ryan Houdek 7f8fffbb24 FEXCore: Implement support for AVX gathers with overflow
Falls down the emulated path so we can zero extend the address
calculation. No way for SVE gathers to use 32-bit addressing.
2025-09-29 16:06:02 -07:00
Billy Laws b4e556e63a IR: Introduce operation to form a scaled address from STATE 2025-09-09 21:29:12 +01:00
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Alyssa Rosenzweig 92b66dbf17 IR: describe immediate inlining in json
This drops some register class validation since InlineConstants don't have a
register class.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Alyssa Rosenzweig c58ad8b593 IR: add AndShift op
will use it for next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:09:20 -04:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
Alyssa Rosenzweig 263279d5dd IR: index blocks
this will let us avoid a costly hashmap in DCE.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 7eaf5ae9e0 IR: drop FillRegister original source
this is now unused, and it's problematic with future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig 8532593d91 IR: specify DestSize for RMWHandle
seems to just have been an oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:58:05 -04:00
Alyssa Rosenzweig 55bd16e2c8 IR: make FillRegister sizes explicit
instead of hacking around it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:57:51 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Ryan Houdek 9b0bb29d78 IR: Implement support for sha256h{2,}
I keep carrying this patch around. Not yet wired up to the instruction
implementation yet, but I don't want to forget about it.
2025-04-01 17:53:48 -07:00
Alyssa Rosenzweig 1f08f8df0d IR: allow VUShrNI with bitshift=0
encodes to Xtn, we need this to narrow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Alyssa Rosenzweig 0038a0b19c IR: plumb Vector_FToISized op
this exposes the frint* opcodes in a new ir op

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Ryan Houdek 4ab6c2252d IR: More validation for shifts 2025-03-27 04:36:41 -07:00
Ryan Houdek 4ebc307744 IR: Ensure BFE ops can't try to extract element larger than size 2025-03-27 04:26:40 -07:00
Alyssa Rosenzweig 694e674fe6 IR: tie VExtr
needed for sve-256 move reduction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:22:55 -04:00
Alyssa Rosenzweig 7ebc0f32b8 IR: tie VInsGPR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00
Alyssa Rosenzweig 68cacc2fc3 IR: tie VFMin/VFMax
this was missed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00