Commit Graph
2781 Commits
Author SHA1 Message Date
Ryan Houdek 19dba4a3ec LookupCache: Add L1 entry on shrink
If we're shrinking the L1 cache then we just deleted the entry that we
just looked up. Add it back to ensure we don't get yet another lookup
for this entry.
2026-08-26 12:18:31 -07:00
Ryan Houdek 49c1cda5c0 Merge pull request #5856 from Hojun-Cho/lookupcache_l1_grow_clear
LookupCache: Clear L1 entries on dynamic cache growth
2026-08-26 12:12:44 -07:00
Hojun-Cho 7e3e4f81ac LookupCache: Clear L1 entries on dynamic cache growth
Growing widens L1PointerMask without touching the table, so an entry whose
address has the new mask bit set sits where InvalidateCache no longer looks,
and a later shrink hands it back to the JIT. Wipe [0, old) on grow, the way
the shrink wipes [new, MAX); the still-valid entries go with it.
2026-08-26 20:51:25 +09:00
Pierre-Loup A. Griffais e740af05b0 DiskCache: bucket/key bookkeeping
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.

Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.

Use printable characters in FOZ key names as intended.
2026-08-25 22:21:09 -07:00
Ryan Houdek 6646a5cc72 OpcodeDispatcher: Fixes SHLD by 16 behaviour
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!

Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Ryan Houdek 2a76d3b153 Merge pull request #5809 from simon902/fcomi_flag_zero
FCOMI fails to clear OF, SF, AF
2026-08-25 13:22:28 -07:00
Ryan Houdek 631b8c3f58 FEXCore/HostFeatures: Adds HostFeatures hashing support
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.

Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
2026-08-24 20:50:35 -07:00
Ryan Houdek b83dd97762 FEXCore/Context: Removes InitialRIP/RSP from CreateThread
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.

Just a smidge of cleanup, NFC.
2026-08-24 18:32:35 -07:00
Ryan Houdek ad3939a44f Avoid double offset relocations
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.

We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.

So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.

Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.

Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.

All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
2026-08-24 12:17:11 -07:00
LC c9b23eb0e7 Merge pull request #5845 from Plagman/plagman/thread_priority_mr
DiskCache: make Writer thread low-priority
2026-08-23 18:09:20 -04:00
Pierre-Loup A. Griffais c3fb6ccaaa DiskCache: make Writer thread low-priority 2026-08-23 14:01:32 -07:00
Pierre-Loup A. Griffais b47a36a47b DiskCache: add some SHM stats 2026-08-23 13:51:01 -07:00
LC af9b438eec Merge pull request #5841 from OFFTKP/blsmsk
unittests/ASM: Test BLSR/BLSMSK CF flag
2026-08-22 23:13:27 -04:00
Ryan Houdek db3817a260 FEXCore: Fixes ever shrinking JIT code buffer
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.

Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
2026-08-22 18:31:22 -07:00
Paris Oplopoios 635befb4c8 FEXCore: Fix CF calculation for BLSMSK and BLSR for 32-bit operands 2026-08-22 15:47:50 +03:00
Ryan Houdek dfcbdac347 JIT: Fixes some alignment logic
Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.

- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
  - Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
  - Previous query between alignments would leave us with up to 12 bytes
    unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
2026-08-21 14:55:03 -07:00
Ryan Houdek 02f8ab5f87 SharedCodeBufferManager: Leak less internal details about implementation
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.

Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.

Should be no functional change.
2026-08-21 14:06:12 -07:00
Pierre-Loup A. Griffais e3f208f61f DiskCache: offload Store to a WorkQueueThread
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.

Add some bounds checking to index population to protect against corruption.
2026-08-21 11:16:07 -07:00
Ryan Houdek c9c5a75b76 FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.

Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.

Shouldn't be any real behaviour change.
2026-08-20 16:35:28 -07:00
Pierre-Loup A. Griffais 561c32b45d Disk Cache initial implementation
Serializes code blocks to disk - only blocks coming from known regions, for now

Disabled by default, key and versioning still needs work, but works for testing
2026-08-19 18:20:49 -07:00
Ryan Houdek 9377bac5e7 Merge pull request #5825 from javelina-pkwy/fix/siglongjmp-deadlock
FEXCore: prevent deadlock when branching to MAX_UINT64
2026-08-17 17:55:39 -07:00
Justin Becker e0bbfdda84 Cleaner change 2026-08-17 12:05:24 -07:00
Ryan Houdek a3d609ebb1 Merge pull request #5820 from javelina-pkwy/lea-reg-reg
Decoder: fix illegal LEA encoding
2026-08-17 09:58:47 -07:00
Justin Becker 6f29dfcbb8 Probe before taking lock in Compile*() 2026-08-13 16:53:05 -07:00
Ryan Houdek 6734c9ed3e SharedCodeBufferManager: Allocate JIT space atomically.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.

As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
2026-08-13 14:58:01 -07:00
Jacek Caban 2158309c51 InterpreterFallbacks: Remove unused template
Fixes -Wunused-template warning.
2026-08-10 15:05:26 +02:00
LC a4e04dd7b4 Merge pull request #5808 from Sonicadvance1/197
FEXCore: Fixes SourceOutline description
2026-08-09 06:19:41 -04:00
Simon Scherer 7f0bdf8d63 OpcodeDispatcher: Zero OF, SF and AF for FCOMI and FCOMIF64 2026-08-06 16:29:29 +02:00
Ryan Houdek 92e43c25d4 Merge pull request #5807 from Sonicadvance1/196
CPUID: Adds a few new bits
2026-08-05 19:56:21 -07:00
Ryan Houdek 0122ef9e83 FEXCore: Fixes SourceOutline description 2026-08-05 19:53:33 -07:00
Ryan Houdek 9365e6240b CPUID: Adds a few new bits
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.

Closes #5805 although their (ML?) write-up was incorrect.
2026-08-05 14:20:56 -07:00
Simon Scherer f8e3571c66 JIT: Skip redundant rounding in Vector_F64ToI32 2026-08-05 13:20:28 +02:00
Simon Scherer 6ebcf65451 JIT: Fix f64->i32 precision loss in non-SVE path in Vector_F64ToI32 2026-08-05 13:17:58 +02:00
Justin Becker 34a87cc94c Decoder: fix illegal LEA encoding 2026-07-27 15:48:35 -07:00
FrontMage 151b4d4c2d Frontend: Prioritize instruction fetch faults 2026-07-25 09:01:21 +08:00
FrontMage fcf9fd77d7 FEXCore: Isolate multiblock error state per block 2026-07-23 16:52:19 +08:00
LC 10a0fe2e71 SharedCodeBufferManager: Make AllocateNew() signature consistent with declaration 2026-07-23 01:42:41 -04:00
LC c9add0d292 SharedCodeBufferManager: Hoist prctl define into util header
Same behavior, but just moves the potential define to be alongside all
of the others in the wrapper header.
2026-07-23 01:40:49 -04:00
LC c239d09ea0 SharedCodeBufferManager: Add missing header
Ensures the page size define is always visible.
2026-07-23 00:16:56 -04:00
Ryan Houdek d2c92808f5 FEXCore: Split out CodeBuffer management to its own file
NFC

- Renames CodeBufferManager to SharedCodeBufferManager to be more
  explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
  about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
  of the CPUBackend code

Makes it easier to parse ownership and lifetime semantics of these
buffers.
2026-07-20 18:09:29 -07:00
LC 2464633431 Merge pull request #5776 from Sonicadvance1/190
JIT: Remove JIT detection string
2026-07-20 21:07:23 -04:00
Ryan Houdek fe1ac1bc1d JIT: Remove JIT detection string
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
2026-07-20 17:44:15 -07:00
Ryan Houdek 9edd27b214 JIT: Rename temporary CPU buffer allocator
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.

NFC
2026-07-20 17:32:03 -07:00
Ryan Houdek eb7e02ea1d Merge pull request #5772 from lioncash/pass
PassManager: Simplify initialization interface
2026-07-19 16:28:07 -07:00
LC fa80d11960 PassManager: Remove SyscallHandler member
This isn't used anymore, so we can get rid of it to further simplify
initialization.
2026-07-21 11:28:12 -04:00
LC 2893d2b64f PassManager: Simplify pass initialization
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
2026-07-21 11:28:09 -04:00
LC 22bd10f3b1 CPUBackend: Remove unnecessary reinterpret_casts
This both take a void*, so the casting is unnecessary to begin with,
since this would occur anyway without it. We can also avoid a
duplication to reduce line noise.
2026-07-21 08:22:17 -04:00
Ryan Houdek f374b4775a Merge pull request #5769 from lioncash/bound
Core: Remove unnecessary bounds check in GenerateIR()
2026-07-18 21:35:34 -07:00
LC ff213bbc5e Core: Move vars closer to usage scope in GenerateIR()
Makes it so their purpose is more easily seen
2026-07-21 04:44:47 -04:00
LC 56a4ca6e6a Core: Remove unnecessary bounds check in GenerateIR()
We already check the bounds in the loop prior to calling at().
2026-07-21 04:40:22 -04:00