Commit Graph
2485 Commits
Author SHA1 Message Date
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
Lioncache c8f3772753 X87Tables: Handle aliases for FSTP 2026-02-26 19:08:48 -05:00
Lioncache b046942a28 X87Tables: Handle aliases for FXCH 2026-02-26 17:44:12 -05:00
Ryan Houdek 4c49036b8a Merge pull request #5332 from lioncash/fcomp
X87Tables: Add handling for FCOMP DE D0 aliases
2026-02-26 13:49:31 -08:00
Lioncache 6909d49683 X87Tables: Add handling for FCOMP DE D0 aliases
Handles the single remaining alias for FCOMP.
2026-02-26 16:14:55 -05:00
Ryan Houdek 32471953ed Merge pull request #5226 from neobrain/fix_jit_consistency
JIT: Ensure code cache consistency by zero-initializing padding bytes
2026-02-26 13:13:10 -08:00
crueter 33a01f2773 Core: Use long commit hash instead of abbreviating
Fixes #5230

Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.

Note: I have no idea how fmt will handle that format string with an
array of uint8_t.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-26 10:08:41 -05:00
Paris Oplopoios c98a30f82b Fix FALSE_OS size 2026-02-24 23:51:25 +02:00
Paris Oplopoios 427206abf7 Fix TRUE_US size 2026-02-24 23:50:40 +02:00
Paris Oplopoios 2eb4267e6e Fix TRUE_US case 2026-02-24 23:36:09 +02:00
Paris Oplopoios f547b63bca Fix NGT_US vector case 2026-02-24 22:33:46 +02:00
Paris Oplopoios fd71928e3a Fix EQ_UQ cases 2026-02-24 22:15:51 +02:00
Ryan Houdek b275068569 FEXCore: Fixes VEX float compare operations
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.

8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.

Both scalar and vector wide.
Fixes #5326
2026-02-23 15:45:44 -08:00
Ryan Houdek 9d86c3270c IR: Implement new VOrn operation 2026-02-23 14:16:45 -08:00
Tony Wasserka 89b976121f JIT: Ensure code cache consistency by zero-initializing padding bytes 2026-02-23 16:49:02 +01:00
LC f5b08bbea6 Merge pull request #5321 from Sonicadvance1/94
AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
2026-02-22 20:57:15 -05:00
LC ca2bdbca19 Merge pull request #5319 from neobrain/fix_cc_codebuffer_size
FEXCore: Default to a larger CodeBuffer size when code caching is enabled
2026-02-21 19:08:06 -05:00
LC 8dfdfdf6ef Merge pull request #5318 from neobrain/fix_cc_validation
CodeCache: Fix validation of large binaries
2026-02-21 19:07:52 -05:00
Ryan Houdek d6d4f84c3b AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
2026-02-21 15:24:13 -08:00
Ryan Houdek 12fcf96e93 IR: Adds a ContextClear operation
This will be useful for zeroing parts of the context at CLZero alignment
and sizes.
2026-02-21 15:24:13 -08:00
Tony Wasserka cec52253fb FEXCore: Default to a larger CodeBuffer size when code caching is enabled
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.

Additionally, FEXOfflineCompiler obviously should never discard compiled code.

Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
2026-02-20 19:07:48 +01:00
Tony Wasserka 31d5805681 CodeCache: Ensure CodeBuffer is large enough for validation run
Previously, CodeBuffer regrowth would discard parts of the code compiled
during the reference run.
2026-02-20 18:49:58 +01:00
Tony Wasserka fda087e24f FEXCore: Add helper function to query available CodeBuffer size 2026-02-20 18:49:07 +01:00
Tony Wasserka 1d8935d0ab CodeCache: Drop unnecessary CodeBuffer size check
The CodeBuffer is resized as required right below.
2026-02-20 18:49:07 +01:00
Ryan Houdek b49730f255 JIT: Support avoiding spilling CPU flags
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.

Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.

Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.

I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
2026-02-19 17:18:30 -08:00
Tony Wasserka 49a37c7d6f Merge pull request #5313 from neobrain/refactor_header_cleanup
Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
2026-02-17 10:31:14 +00:00
Tony Wasserka 7bdf53291b CodeCache: Avoid unnecessary copying of CodeBuffer contents 2026-02-16 11:07:32 +01:00
Tony Wasserka 44a6bfd874 Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
CodeCache.h has no dependency on SourcecodeResolver.h.

This reverts commit f2bbc0eccd.
2026-02-16 10:29:58 +01:00
Ryan Houdek 36ec090ee3 Merge pull request #5300 from pmatos/feat/MirrorsEdge
Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
2026-02-11 16:49:29 -08:00
Paulo Matos 11a5229f37 Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
This was mainly targetted to MirrorsEdge but hopefully will help other workloads
that hit these softfloat conversions in a similar way.
2026-02-11 18:39:20 +01:00
Ryan Houdek cf3c41acc7 FEXCore: Fixup ARPL implementation
Some subtle details were wrong in the implementation, and some minor
changes to be more optimal.
2026-02-10 14:29:32 -08:00
FrontMage 9bf9c00474 x86: implement ARPL (0x63) in 32-bit mode 2026-02-10 14:12:52 -08:00
FrontMage 51afcc5ca8 x87: fix dispatch table sizes after adding FCOM/FCOMP 2026-02-10 18:08:41 +08:00