Commit Graph
400 Commits
Author SHA1 Message Date
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Alyssa Rosenzweig 92b66dbf17 IR: describe immediate inlining in json
This drops some register class validation since InlineConstants don't have a
register class.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Alyssa Rosenzweig c58ad8b593 IR: add AndShift op
will use it for next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:09:20 -04:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
Alyssa Rosenzweig 263279d5dd IR: index blocks
this will let us avoid a costly hashmap in DCE.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 7eaf5ae9e0 IR: drop FillRegister original source
this is now unused, and it's problematic with future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig 8532593d91 IR: specify DestSize for RMWHandle
seems to just have been an oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:58:05 -04:00
Alyssa Rosenzweig 55bd16e2c8 IR: make FillRegister sizes explicit
instead of hacking around it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:57:51 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Ryan Houdek 9b0bb29d78 IR: Implement support for sha256h{2,}
I keep carrying this patch around. Not yet wired up to the instruction
implementation yet, but I don't want to forget about it.
2025-04-01 17:53:48 -07:00
Alyssa Rosenzweig 1f08f8df0d IR: allow VUShrNI with bitshift=0
encodes to Xtn, we need this to narrow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Alyssa Rosenzweig 0038a0b19c IR: plumb Vector_FToISized op
this exposes the frint* opcodes in a new ir op

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Ryan Houdek 4ab6c2252d IR: More validation for shifts 2025-03-27 04:36:41 -07:00
Ryan Houdek 4ebc307744 IR: Ensure BFE ops can't try to extract element larger than size 2025-03-27 04:26:40 -07:00
Alyssa Rosenzweig 694e674fe6 IR: tie VExtr
needed for sve-256 move reduction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:22:55 -04:00
Alyssa Rosenzweig 7ebc0f32b8 IR: tie VInsGPR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00
Alyssa Rosenzweig 68cacc2fc3 IR: tie VFMin/VFMax
this was missed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 15:16:25 -04:00
Alyssa Rosenzweig 43eb597044 IR: tie VBSL source
this was missed before, noticed while experimenting with round robin RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-25 14:14:21 -04:00
Paulo Matos e18a661b50 Refactor OP_STORESTACKMEM case in x87 Stack Opt Pass 2025-03-18 18:34:06 +01:00
Ryan Houdek 9f2f10a65f IR: Add new 128-bit vector move operation
This allows us to slightly optimize some x87 behaviour.
2025-03-12 22:21:28 -07:00
Ryan Houdek a37d6a3841 JIT: Optimize CAS
Hey kid, want to see a sick trick?

Finally optimal codegen for 64-bit cmpxchg.
2025-03-07 14:24:45 -08:00
Ryan Houdek 66841ce35c IR: Remove some static analysis warnings
NFC
2025-03-05 12:30:27 -08:00
Paulo Matos b5241e0f60 Improve reciprocal estimate and tests
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.

* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.

Fixes #4319.
2025-02-27 15:07:14 +01:00
Ryan Houdek de6931b1f5 OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
```diff
     "sha1rnds4 xmm0, xmm1, 10b": {
-      "ExpectedInstructionCount": 55,
+      "ExpectedInstructionCount": 10,
```

So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
2025-02-22 04:57:58 -08:00
Ryan Houdek cdf6a16efc JIT: Implement ARM VSha1SU1 IR operation 2025-02-19 11:33:38 -08:00
Ryan Houdek bbcd4c168c JIT: Implement support for VSha256U1 operation 2025-02-18 18:03:31 -08:00
Paulo Matos 44c65c35c8 Revert "Enable RA of SVE Predicate Registers"
This reverts commit fcbf0de05a.

The initial user of this code has been re-implemented in  b148cc6c.
This is not needed any longer so we're removing it.
2025-01-29 11:56:19 +01:00
Paulo Matos 0bccb1ece5 Ensure predicate cache is reset when control flow leaves block
Whenever the control float leaves the block, it might clobber the
predicate register so we reset the cache whenever that happens.

Fixes #4264
2025-01-27 20:12:43 +01:00
Paulo Matos 2d53867668 x87 fst/fld optimization for different addrmodes
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.

Discussed in #4252.
2025-01-14 16:20:33 +01:00
Ryan Houdek 1ecfa3253d IR: Change convention from number of elements to elementsize
The IR stores elementsize, where the json was wanting number of
elements. While the IR Emitter function declaration always wanted
element size. This was causing us to do a little dance from ElementSize
-> Number of elements -> ElementSize. Just pass the ElementSize directly
instead of this bogus little dance.
2025-01-03 11:01:03 -08:00
Paulo Matos 1d3ce30e50 Generate SVE for 80bit load/stores when possible
Fixes #4166.
2024-12-06 10:15:29 +01:00
Paulo Matos fcbf0de05a Enable RA of SVE Predicate Registers 2024-12-02 18:35:31 +01:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
Ryan Houdek 82f936cb6d IR: Converts base IR operations to store OpSize sizes
NFC

Finally converts the IR operations themselves to store the OpSize for
the IR operation size and element sizes.

This also finally, FINALLY, converts that remaining `_Constant` helper
to stop using a size field that is specified in bits rather than bytes
like all the other IR op handlers. That thing was so confusing and now
it's gone.
2024-10-28 21:26:59 -07:00
Ryan Houdek 6cca007817 IR: Fix some missing OpSize conversions
Missed these in the previous PR.
2024-10-28 16:25:36 -07:00
Ryan Houdek 65ddae1b71 IR: Change F80VBSLStack to use IR::OpSize 2024-10-28 02:25:17 -07:00
Ryan Houdek 063f524084 IR: Change F80CVTToInt to use IR::OpSize 2024-10-28 02:24:25 -07:00