Commit Graph
3921 Commits
Author SHA1 Message Date
Ryan Houdek 66cad978c3 FEXCore: Removes syscall optimization
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.

Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.

This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.

Bumps the DiskCache version again because it causes codegen to change.
2026-08-31 19:23:18 -07:00
Ryan Houdek 8cf2bf7ad8 Merge pull request #5884 from Plagman/plagman/lookup_opt_mr
DiskCache: Lookup optimizations
2026-08-31 14:05:34 -07:00
Pierre-Loup A. Griffais c0b24b2f2e DiskCache: on Linux, block signals in Writer thread
Clean up thread flags while we're at it.
2026-08-30 22:27:05 -07:00
Pierre-Loup A. Griffais 01bc89b82e DiskCache: Lookup optimizations
- quicker misses from storing GuestHash/GuestSize in index
 - use GuestSize to pull less data from disk on hit
2026-08-30 16:50:11 -07:00
Ryan Houdek eca6569bd3 FEXCore: Expose DiskCache version to the frontend
With a comment about includes currently being incorrectly shared to the
frontend and needs to get resolved.
2026-08-30 14:26:43 -07:00
Pierre-Loup A. Griffais b3f902166b DiskCache: get rid of more extra copies/allocs on Lookup
Reorganize disk format a bit so that entrypoints and guest pages can be used
as is from the original blob allocation, with some in-place relocation.
2026-08-30 01:52:42 -07:00
Lioncache 7a8b9dd341 Config: Fix assertion being hit in config generation 2026-08-28 16:51:11 -04:00
LC 8ccc8dab49 Merge pull request #5865 from Sonicadvance1/221
FEXCore/Config:  Annotate all config options that can affect codegen
2026-08-28 17:59:11 -04:00
Pierre-Loup A. Griffais 8b9237fd52 DiskCache: Apply relocations from disk blob
Removes one allocation in Lookup path.
2026-08-27 22:45:09 -07:00
Ryan Houdek fd180a16d3 Merge pull request #5861 from Plagman/plagman/cache_map_mr
DiskCache: use memory-mapped IO for reads when possible
2026-08-27 22:18:13 -07:00
Pierre-Loup A. Griffais 2a82f58195 DiskCache: use memory-mapped IO for reads when possible
On Wine, use a unixlib call to get a quality mapping that can track a growing
file.
2026-08-27 21:15:22 -07:00
Ryan Houdek 34ffd2d6e2 FEXCore/Config: Annotate all config options that can affect codegen
Instead serializing the world, allow the config to be data driven. A
couple host feature options aren't serialized as they get explained
elsewhere.
2026-08-27 15:08:20 -07:00
LC 62c9c130c4 Merge pull request #5864 from simon902/pf2id_saturation
Fix pf2id overflow saturation
2026-08-27 09:16:04 -04:00
LC fca23d88c3 Merge pull request #5863 from simon902/vfaddp-mmx-width
Fix PFACC producing wrong results due to missing MMX-sized path in VFAddP
2026-08-27 09:15:30 -04:00
Simon Scherer b21a8dafa7 OpcodeDispatcher: Fix pf2id overflow saturation 2026-08-27 11:50:45 +02:00
Simon Scherer f0135eb332 JIT/VectorOps: Separate 64-bit VFAddP path for MMX-sized operands 2026-08-27 10:36:22 +02:00
Simon Scherer 5ab2c723ff OpcodeDispatcher: For PF2IWOp use VSQXTN instead of VUnZip to saturate values falling outside the 16-bit range 2026-08-27 09:36:17 +02:00
Pierre-Loup A. Griffais 753bce5e73 DiskCache: clean up bucket key building a bit
Less manual math like that.
2026-08-26 17:35:54 -07:00
LC baca86bc3f Merge pull request #5859 from Sonicadvance1/220
Config: Filter the `CPUFeatureRegisters` option on serialization
2026-08-26 15:33:59 -04:00
Ryan Houdek ef3a6963f3 Config: Filter the CPUFeatureRegisters option on serialization
This is unfiltered data that ends up in HostFeatures. If it was
serialized when set then it would effectively never allow the code
serialization to be used.
2026-08-26 12:22:08 -07:00
Ryan Houdek 19dba4a3ec LookupCache: Add L1 entry on shrink
If we're shrinking the L1 cache then we just deleted the entry that we
just looked up. Add it back to ensure we don't get yet another lookup
for this entry.
2026-08-26 12:18:31 -07:00
Ryan Houdek 49c1cda5c0 Merge pull request #5856 from Hojun-Cho/lookupcache_l1_grow_clear
LookupCache: Clear L1 entries on dynamic cache growth
2026-08-26 12:12:44 -07:00
Hojun-Cho 7e3e4f81ac LookupCache: Clear L1 entries on dynamic cache growth
Growing widens L1PointerMask without touching the table, so an entry whose
address has the new mask bit set sits where InvalidateCache no longer looks,
and a later shrink hands it back to the JIT. Wipe [0, old) on grow, the way
the shrink wipes [new, MAX); the still-valid entries go with it.
2026-08-26 20:51:25 +09:00
Pierre-Loup A. Griffais e740af05b0 DiskCache: bucket/key bookkeeping
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.

Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.

Use printable characters in FOZ key names as intended.
2026-08-25 22:21:09 -07:00
LC 4ed80fd071 Merge pull request #5855 from Sonicadvance1/218
HostFeatures: Fixes crash in mingw build
2026-08-25 20:37:52 -04:00
Ryan Houdek 59116a06b9 HostFeatures: Fixes crash in mingw build
Turns out packed enum classes without specifying an underlying type
causes problems. Declare its underlying type as uint32_t to match
everything else here.
2026-08-25 17:19:35 -07:00
Ryan Houdek 6646a5cc72 OpcodeDispatcher: Fixes SHLD by 16 behaviour
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!

Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Ryan Houdek 2a76d3b153 Merge pull request #5809 from simon902/fcomi_flag_zero
FCOMI fails to clear OF, SF, AF
2026-08-25 13:22:28 -07:00
Ryan Houdek 631b8c3f58 FEXCore/HostFeatures: Adds HostFeatures hashing support
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.

Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
2026-08-24 20:50:35 -07:00
Ryan Houdek b83dd97762 FEXCore/Context: Removes InitialRIP/RSP from CreateThread
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.

Just a smidge of cleanup, NFC.
2026-08-24 18:32:35 -07:00
Ryan Houdek 26f6cda589 FEXCore/Config: Support Serializing
Serializes every option, even ones that are set to default to ensure
validation that if any config option value is added or changed that they
are captured.

Skips a handful of options that are either meta options, environment
options that don't matter, or HostFeatures which is handled elsewhere.
(HostFeatures will be controlling bucketing rather than the remaining
options).

This serialization is currently 1413 bytes and generates in 17580ns on
my A1A. So it's not the fastest, definitely don't want to be generating
it constantly per process.
2026-08-24 16:10:42 -07:00
Ryan Houdek ad3939a44f Avoid double offset relocations
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.

We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.

So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.

Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.

Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.

All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
2026-08-24 12:17:11 -07:00
LC c9b23eb0e7 Merge pull request #5845 from Plagman/plagman/thread_priority_mr
DiskCache: make Writer thread low-priority
2026-08-23 18:09:20 -04:00
Pierre-Loup A. Griffais c3fb6ccaaa DiskCache: make Writer thread low-priority 2026-08-23 14:01:32 -07:00
Pierre-Loup A. Griffais b47a36a47b DiskCache: add some SHM stats 2026-08-23 13:51:01 -07:00
LC af9b438eec Merge pull request #5841 from OFFTKP/blsmsk
unittests/ASM: Test BLSR/BLSMSK CF flag
2026-08-22 23:13:27 -04:00
Ryan Houdek db3817a260 FEXCore: Fixes ever shrinking JIT code buffer
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.

Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
2026-08-22 18:31:22 -07:00
Paris Oplopoios 635befb4c8 FEXCore: Fix CF calculation for BLSMSK and BLSR for 32-bit operands 2026-08-22 15:47:50 +03:00
Ryan Houdek dfcbdac347 JIT: Fixes some alignment logic
Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.

- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
  - Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
  - Previous query between alignments would leave us with up to 12 bytes
    unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
2026-08-21 14:55:03 -07:00
Ryan Houdek 02f8ab5f87 SharedCodeBufferManager: Leak less internal details about implementation
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.

Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.

Should be no functional change.
2026-08-21 14:06:12 -07:00
Pierre-Loup A. Griffais e3f208f61f DiskCache: offload Store to a WorkQueueThread
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.

Add some bounds checking to index population to protect against corruption.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais 008d990f20 Utils: add WorkQueueThread
Straightforward queue for arbitrary work.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais efc5eaf9c7 Utils/File: add PRead(), PWrite() and Size()
Because PRead/PWrite don't have the same side effects on the file cursor
between Linux and Windows (:/), make them all-or-nothing.
2026-08-21 11:16:07 -07:00
Ryan Houdek c9c5a75b76 FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.

Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.

Shouldn't be any real behaviour change.
2026-08-20 16:35:28 -07:00
Pierre-Loup A. Griffais 561c32b45d Disk Cache initial implementation
Serializes code blocks to disk - only blocks coming from known regions, for now

Disabled by default, key and versioning still needs work, but works for testing
2026-08-19 18:20:49 -07:00
Ryan Houdek 9377bac5e7 Merge pull request #5825 from javelina-pkwy/fix/siglongjmp-deadlock
FEXCore: prevent deadlock when branching to MAX_UINT64
2026-08-17 17:55:39 -07:00
Justin Becker e0bbfdda84 Cleaner change 2026-08-17 12:05:24 -07:00
Ryan Houdek a3d609ebb1 Merge pull request #5820 from javelina-pkwy/lea-reg-reg
Decoder: fix illegal LEA encoding
2026-08-17 09:58:47 -07:00
Justin Becker 6f29dfcbb8 Probe before taking lock in Compile*() 2026-08-13 16:53:05 -07:00
Ryan Houdek 6734c9ed3e SharedCodeBufferManager: Allocate JIT space atomically.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.

As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
2026-08-13 14:58:01 -07:00