Commit Graph
663 Commits
Author SHA1 Message Date
Ryan Houdek d2c92808f5 FEXCore: Split out CodeBuffer management to its own file
NFC

- Renames CodeBufferManager to SharedCodeBufferManager to be more
  explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
  about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
  of the CPUBackend code

Makes it easier to parse ownership and lifetime semantics of these
buffers.
2026-07-20 18:09:29 -07:00
LC 2464633431 Merge pull request #5776 from Sonicadvance1/190
JIT: Remove JIT detection string
2026-07-20 21:07:23 -04:00
Ryan Houdek fe1ac1bc1d JIT: Remove JIT detection string
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
2026-07-20 17:44:15 -07:00
Ryan Houdek 9edd27b214 JIT: Rename temporary CPU buffer allocator
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.

NFC
2026-07-20 17:32:03 -07:00
Ryan Houdek 921ce59054 Merge pull request #5750 from lioncash/pred
VectorOps: Make use of unpredicated shifts
2026-07-15 09:15:04 -07:00
Ryan Houdek a7627ba39a Merge pull request #5751 from lioncash/sq
VectorOps: Add trivial case handling in VSQXTN2
2026-07-15 09:06:59 -07:00
Ryan Houdek 0f2463dda2 Merge pull request #5748 from lioncash/invariant
RegisterAllocationPass: Ensure pair reg invariant
2026-07-15 09:04:38 -07:00
LC ba8b0afe7a VectorOps: Use unpredicated shifts where applicable for 256-bit scalar shifts
Lets us trim some output
2026-07-15 09:10:57 -04:00
LC 98d45a6a9b VectorOps: Make use of unpredicated immediate shifts
Same behavior, just without introducing a predicate register dependency.
2026-07-15 08:00:58 -04:00
LC 6bc808cb06 VectorOps: Add trivial case handling in VSQXTN2
Lets us generate much more optimal code in the event the destination and
lower source are the same.
2026-07-15 07:49:50 -04:00
LC d6fb60d512 RegisterAllocationPass: Ensure pair reg invariant
Allows us to actually catch if this requirement ever gets broken in
the future.
2026-07-15 06:46:24 -04:00
LC ef35474f88 VectorOps: Simplify 256-bit VAddV
Didn't read the manual close enough on the first read award.
2026-07-15 04:43:47 -04:00
LC 9e8e87bbb2 AtomicOps: Avoid constrained unpredictable case in TelemetrySetValue()
STLXR cannot use the same register as both the status register and the
value register, otherwise it's architecturally unpredictable
behavior.

Only applies to hardware without FEAT_LSE, so this only meaningfully
affects hardware using the v8.0 spec, since FEAT_LSE becomes mandatory
in v8.1 and newer.
2026-07-13 19:08:17 -04:00
Ryan Houdek 28cdae4687 Merge pull request #5737 from lioncash/bsl
VectorOps: Simplify SVE 256-bit VOrn with BSL2N
2026-07-13 10:05:08 -07:00
LC 842e22915c VectorOps: Simplify SVE 256-bit VOrn with BSL2N
Lets us shave off an instruction and also avoid using a temporary
register in some cases. We can also tweak our worst case that requires a
predicate to eliminate the temporary as well.

We can also expand our cmpps cases, so that we can reflect the
BSL2N usages in instcountci.
2026-07-13 12:27:54 -04:00
LC 24720b67da VectorOps: Make SVE shift==0 case symmetric with ASIMD
Ensures that we have consistent behavior.
2026-07-13 10:26:42 -04:00
LC 9d3c388664 Crypto: Make use of XAR in SHA1NEXTE when available
Lets us shave an instruction off on hardware that supports XAR.

Closes #5730
2026-07-12 16:02:12 -04:00
Ryan Houdek 821dfe5b98 Merge pull request #5708 from lioncash/mrs
MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
2026-07-11 10:56:27 -07:00
LC 613e9ef701 MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
Previously this had the potential to not clear rounding bits properly
depending on incoming FPCR state.
2026-07-11 10:24:07 -04:00
LC 5080c6ffc5 MemoryOps: Avoid double application of base offset in {Load,Store}ContextIndexed unaligned case 2026-07-11 09:56:56 -04:00
LC 8b0f07b2dc Syscalls: Remove unnecessary usages of namespace FEXCore::IR
These aren't necessary.
2026-07-10 21:02:12 -04:00
Ryan Houdek c8dd9eefa6 Merge pull request #5698 from lioncash/dead
ConversionOps: Remove redundant code in Vector_FToS
2026-07-10 10:35:06 -07:00
LC 150b25f29e ConversionOps: Remove redundant code in Vector_FToS
These are already defined in an outer scope.
2026-07-10 09:16:35 -04:00
LC dd838d4ad3 MemoryOps: Make use of current working reg for cache operations
TMP1 technically isn't initialized properly here until after the first
iteration.
2026-07-10 07:13:17 -04:00
Ryan Houdek 5f2d19c7aa Merge pull request #5685 from lioncash/faddv 2026-07-10 03:33:43 -07:00
Ryan Houdek 95e7c866cf Merge pull request #5682 from lioncash/pid 2026-07-10 03:20:23 -07:00
LC ede09a03db VectorOps: Fix 256-bit FADDV path
Avoids falling down to the SVE-128 path.
2026-07-10 05:50:22 -04:00
LC b5660c8a92 MiscOps: Avoid stack misalignment in ProcessorID
This needs to be an add.
2026-07-10 05:41:54 -04:00
LC 54263bb5a7 ALUOps: Fix 32-bit MulH case
These need to be 64-bit multiply and ubfx. Thankfully this case wasn't
actually hit in practice.
2026-07-10 05:39:14 -04:00
LC caa030714b IR: Remove unused VUShraI
Given that this is currently unused and that we don't have the signed
equivalent implemented, we can just remove this for now.
2026-07-10 04:15:52 -04:00
LC 6488dcbb01 ALUOps: Remove unused macros
These are now unused.
2026-07-10 03:30:03 -04:00
LC 1b1e46ff6c VectorOps: Join identical branches in VFMLS/VFNMLS
Same thing, just a little less redundant.
2026-07-09 16:33:26 -04:00
Ryan Houdek 5f2455c502 Merge pull request #5665 from lioncash/blendop
[SVE256] Handle 256-bit blend operations much more efficiently
2026-07-08 16:43:20 -07:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
Simon Scherer 655102fc7d JIT/ALUOps: Fix operand overlapping bug for pdep 2026-07-08 10:09:47 +02:00
LC 5145324806 [SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC d3a85e14d9 VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek 27a5f09185 Merge pull request #5592 from ShadowCurse/fixes
JIT: Arm64: fix the loop in CacheLineClear/Clean
2026-06-23 17:15:07 -07:00
Egor Lazarchuk e3e9777ee6 JIT: Arm64: fix the loop in CacheLineClear/Clean
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
2026-06-24 00:38:48 +01:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC 1d3403fdc2 JIT: Amend op typos in implementations
Mostly benign, but ensures that they're correct in the event any of
their IR definitions change.
2026-06-20 13:16:12 -04:00
Daniel Lu 03009912ac JIT: Avoid clobbering guest rdx while raising generated faults 2026-05-22 13:56:40 -07:00
LC c98cef0da1 Merge pull request #5506 from Sonicadvance1/154
HostFeatures: Only enable `dc zva` optimization on Ampere CPUs
2026-05-19 22:29:11 -04:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00