Commit Graph
2469 Commits
Author SHA1 Message Date
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
Lioncache c8f3772753 X87Tables: Handle aliases for FSTP 2026-02-26 19:08:48 -05:00
Lioncache b046942a28 X87Tables: Handle aliases for FXCH 2026-02-26 17:44:12 -05:00
Ryan Houdek 4c49036b8a Merge pull request #5332 from lioncash/fcomp
X87Tables: Add handling for FCOMP DE D0 aliases
2026-02-26 13:49:31 -08:00
Lioncache 6909d49683 X87Tables: Add handling for FCOMP DE D0 aliases
Handles the single remaining alias for FCOMP.
2026-02-26 16:14:55 -05:00
Ryan Houdek 32471953ed Merge pull request #5226 from neobrain/fix_jit_consistency
JIT: Ensure code cache consistency by zero-initializing padding bytes
2026-02-26 13:13:10 -08:00
crueter 33a01f2773 Core: Use long commit hash instead of abbreviating
Fixes #5230

Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.

Note: I have no idea how fmt will handle that format string with an
array of uint8_t.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-26 10:08:41 -05:00
Paris Oplopoios c98a30f82b Fix FALSE_OS size 2026-02-24 23:51:25 +02:00
Paris Oplopoios 427206abf7 Fix TRUE_US size 2026-02-24 23:50:40 +02:00
Paris Oplopoios 2eb4267e6e Fix TRUE_US case 2026-02-24 23:36:09 +02:00
Paris Oplopoios f547b63bca Fix NGT_US vector case 2026-02-24 22:33:46 +02:00
Paris Oplopoios fd71928e3a Fix EQ_UQ cases 2026-02-24 22:15:51 +02:00
Ryan Houdek b275068569 FEXCore: Fixes VEX float compare operations
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.

8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.

Both scalar and vector wide.
Fixes #5326
2026-02-23 15:45:44 -08:00
Ryan Houdek 9d86c3270c IR: Implement new VOrn operation 2026-02-23 14:16:45 -08:00
Tony Wasserka 89b976121f JIT: Ensure code cache consistency by zero-initializing padding bytes 2026-02-23 16:49:02 +01:00
LC f5b08bbea6 Merge pull request #5321 from Sonicadvance1/94
AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
2026-02-22 20:57:15 -05:00
LC ca2bdbca19 Merge pull request #5319 from neobrain/fix_cc_codebuffer_size
FEXCore: Default to a larger CodeBuffer size when code caching is enabled
2026-02-21 19:08:06 -05:00
LC 8dfdfdf6ef Merge pull request #5318 from neobrain/fix_cc_validation
CodeCache: Fix validation of large binaries
2026-02-21 19:07:52 -05:00
Ryan Houdek d6d4f84c3b AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
2026-02-21 15:24:13 -08:00
Ryan Houdek 12fcf96e93 IR: Adds a ContextClear operation
This will be useful for zeroing parts of the context at CLZero alignment
and sizes.
2026-02-21 15:24:13 -08:00
Tony Wasserka cec52253fb FEXCore: Default to a larger CodeBuffer size when code caching is enabled
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.

Additionally, FEXOfflineCompiler obviously should never discard compiled code.

Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
2026-02-20 19:07:48 +01:00
Tony Wasserka 31d5805681 CodeCache: Ensure CodeBuffer is large enough for validation run
Previously, CodeBuffer regrowth would discard parts of the code compiled
during the reference run.
2026-02-20 18:49:58 +01:00
Tony Wasserka fda087e24f FEXCore: Add helper function to query available CodeBuffer size 2026-02-20 18:49:07 +01:00
Tony Wasserka 1d8935d0ab CodeCache: Drop unnecessary CodeBuffer size check
The CodeBuffer is resized as required right below.
2026-02-20 18:49:07 +01:00
Ryan Houdek b49730f255 JIT: Support avoiding spilling CPU flags
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.

Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.

Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.

I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
2026-02-19 17:18:30 -08:00
Tony Wasserka 49a37c7d6f Merge pull request #5313 from neobrain/refactor_header_cleanup
Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
2026-02-17 10:31:14 +00:00
Tony Wasserka 7bdf53291b CodeCache: Avoid unnecessary copying of CodeBuffer contents 2026-02-16 11:07:32 +01:00
Tony Wasserka 44a6bfd874 Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
CodeCache.h has no dependency on SourcecodeResolver.h.

This reverts commit f2bbc0eccd.
2026-02-16 10:29:58 +01:00
Ryan Houdek 36ec090ee3 Merge pull request #5300 from pmatos/feat/MirrorsEdge
Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
2026-02-11 16:49:29 -08:00
Paulo Matos 11a5229f37 Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
This was mainly targetted to MirrorsEdge but hopefully will help other workloads
that hit these softfloat conversions in a similar way.
2026-02-11 18:39:20 +01:00
Ryan Houdek cf3c41acc7 FEXCore: Fixup ARPL implementation
Some subtle details were wrong in the implementation, and some minor
changes to be more optimal.
2026-02-10 14:29:32 -08:00
FrontMage 9bf9c00474 x86: implement ARPL (0x63) in 32-bit mode 2026-02-10 14:12:52 -08:00
FrontMage 51afcc5ca8 x87: fix dispatch table sizes after adding FCOM/FCOMP 2026-02-10 18:08:41 +08:00
FrontMage 7a0ed4c462 x87: support 0xDC D0/D8 FCOM/FCOMP st(i) 2026-02-10 18:08:41 +08:00
LC aabddede8f Merge pull request #5279 from Sonicadvance1/75
FEXCore: Disable 48-bit VA optimization
2026-02-04 23:59:40 -05:00
Ryan Houdek 217d039bb7 Merge pull request #5277 from Sonicadvance1/73
CPUID: Add option for hiding hybrid Big.Little CPUs
2026-02-03 11:46:30 -08:00
Ryan Houdek 2fc664f12a FEXCore: Early exit for invalid VEX.vvvv on 32-bit
Otherwise we hit an assert in FEXCore backend with code discovery
hitting things that look like AVX.

Fixes a crash in Uplay.

Also adds a test to just ensure that the instruction faults out and is
captured instead of crashing in FEX itself.
2026-02-02 17:44:50 -08:00
Ryan Houdek 3d989dbaff FEXCore: Disable 48-bit VA optimization
This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.

Not that big of a deal, it's only a minor optimization anyway.

Fixes #5227
2026-02-02 15:17:28 -08:00
Ryan Houdek 8bfa6b817f CPUID: Add option for hiding hybrid Big.Little CPUs
Required for Denuvo?
2026-02-02 12:59:44 -08:00
Ryan Houdek c6031a7806 Softfloat: Remove weird tail padding from X80SoftFloat
These are expected to match x87 registers in side. This was always a bit
weird.
2026-01-28 19:40:29 -08:00
Tony Wasserka ba2b0ef809 Enable automatic code formatting for X86Tables.h 2026-01-23 15:39:54 +01:00
Paulo Matos d6fd82c1ab Add Zydis x86/x86-64 disassembler support
Integrate Zydis as an optional dependency to enable x86/x86-64 guest
instruction disassembly during JIT compilation.

Build with -DENABLE_ZYDIS=TRUE.
Use FEX_X86DISASSEMBLE=1 at runtime to output guest x86 instructions
for each compiled block.
2026-01-14 15:40:28 +01:00
Ryan Houdek 6b33613bb0 CPUID: Adds AmpereOneC identifier 2026-01-12 20:04:56 -08:00
Ryan Houdek d7c2b8f513 FEXCore: Cleanup pointers structure
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.

NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
2026-01-07 12:55:34 -08:00
Ryan Houdek 817a927e31 FEXCore: Fixes circular dependency with thunk callback
Fixes crash in thunks that use callbacks, introduced in #5148.

The dispatcher would call the syscallhandler to get the VDSO thunk
callback. But due to reordering initialization, the VDSO thunk would
have not been loaded at that point. This would cause thunks that use
callbacks to crash with a nullptr exception.

Instead, defer the thunk callback pointer loading until the thread
starts executing, and load the pointer in to our thread state's pointer
struct instead.

Didn't get caught in my initial test sweep since I didn't run a Wine
game with thunks.
2026-01-07 12:03:19 -08:00
Ryan Houdek ac3cabec07 Relocations: Disable 6-byte size optimization in InsertGuestRIPMove
This wasn't handling negatives correctly which was causing xalia.exe to
assert. Disable for now rather than further changing logic, with a TODO that it
should get fixed in the future.
2026-01-07 09:19:13 -08:00
Billy Laws 1ec8c8763e OpcodeDispatcher: Explicitly calculate flags after _TelemetrySetValue
Opcode handlers are written with the assumption that LoadSource will not
touch flags and this would be an annoying assumption to change. As this
is such an edge case anyway just don't defer flags and force a load of
the saved value before _TelemetrySetValue (which are implicitly saved
before it).

Fixes the following snippet in upc.exe:
AND        word ptr [ESP + ECX*0x1 + 0x80000000],DX
BTR        CX,DX
ADC        CX,word ptr SS:[EAX + ECX*0x1 + 0x80000000]
2026-01-03 03:52:51 +00:00
Billy Laws 6b583ee697 Relocations: Switch to robin_map to improve lookup perf 2025-12-31 15:12:05 +00:00
Ryan Houdek 0b92d431f5 Frontend: Only decode REX if it is at the correct location
Otherwise it is a nop
2025-12-29 17:57:12 -08:00