Commit Graph
3870 Commits
Author SHA1 Message Date
Ryan Houdek 6734c9ed3e SharedCodeBufferManager: Allocate JIT space atomically.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.

As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
2026-08-13 14:58:01 -07:00
Ryan Houdek 30a81484d7 FEXCore/unittests: Adds tests for atomic bitmap allocator 2026-08-12 14:32:31 -07:00
Ryan Houdek 00a6b046a8 FEXCore/Utils: Implements a new atomic segmented bitmap allocator
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.

In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.

There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
2026-08-12 14:32:31 -07:00
Ryan Houdek 40940ae0b3 FEXCore/unittests: Adds atomic bitset unittest
Hammers the API in a couple of ways to make sure it works.
2026-08-10 14:12:10 -07:00
Ryan Houdek 1bab28dad4 FEXCore: Adds a lock-free atomic bitset that supports contiguous range allocations
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word

These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.

Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
  those many contiguous bits within a single atomic word
  - Allocation can fail if there are cross-word contiguous bits of the
    size available
  - Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
  - Searches for contiguous bits of free space between multiple atomic
    words.
  - If found, will attempt to allocate tracking which bits were allocated
  - If allocation fails, unwind bits already acquired and continue
    scanning

Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
  - If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
  overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
  acquire semantics for scanning can technically overlook some
  allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.

Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.

The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.
2026-08-10 14:12:09 -07:00
Ryan Houdek 4838265589 FEXCore/MathUtils: Adds helper for alignment by power of 2 size
Useful for removing integer division instructions when we know the
source value is aligned to be power of two. As integer division is quite
slow, we want to use this when possible.
2026-08-10 14:08:07 -07:00
Jacek Caban 363bf85b5b LongJump: Silence -Winline-asm warnings on ARM64EC
Fixes disallowed registers warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban de16c18961 SignalScopeGuards: Mark template function helpers as inline
Fixes -Wunused-template warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban 2158309c51 InterpreterFallbacks: Remove unused template
Fixes -Wunused-template warning.
2026-08-10 15:05:26 +02:00
LC 7966bfb077 Merge pull request #5816 from Sonicadvance1/199
Misc: Adds some missing headers
2026-08-09 06:19:59 -04:00
LC a4e04dd7b4 Merge pull request #5808 from Sonicadvance1/197
FEXCore: Fixes SourceOutline description
2026-08-09 06:19:41 -04:00
Ryan Houdek b161a74365 Misc: Adds some missing headers
Newer compiler and libraries got angry that these were missing.
2026-08-07 14:52:00 -07:00
Ryan Houdek 92e43c25d4 Merge pull request #5807 from Sonicadvance1/196
CPUID: Adds a few new bits
2026-08-05 19:56:21 -07:00
Ryan Houdek 0122ef9e83 FEXCore: Fixes SourceOutline description 2026-08-05 19:53:33 -07:00
Ryan Houdek 9365e6240b CPUID: Adds a few new bits
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.

Closes #5805 although their (ML?) write-up was incorrect.
2026-08-05 14:20:56 -07:00
Simon Scherer f8e3571c66 JIT: Skip redundant rounding in Vector_F64ToI32 2026-08-05 13:20:28 +02:00
Simon Scherer 6ebcf65451 JIT: Fix f64->i32 precision loss in non-SVE path in Vector_F64ToI32 2026-08-05 13:17:58 +02:00
Martin Storsjö 08031a2767 Add missing includes
This fixes compilation with libc++ 23, which has removed a number
of unnecessary transitive includes in its headers.

Include <cstdlib> in StringConv.h for std::strtoll and std::strtoull.

Include <cstdlib> for the declarations of malloc/free/realloc/calloc
in Alloc.cpp. (Without this, the functions we define end up with
C++ name mangling.)

Include <stdarg.h> in IO.cpp for va_start/va_end.
2026-07-28 23:17:43 +03:00
FrontMage 151b4d4c2d Frontend: Prioritize instruction fetch faults 2026-07-25 09:01:21 +08:00
Ryan Houdek d028c7942b Merge pull request #5782 from lioncash/validation
IRValidation: Minor cleanups
2026-07-23 15:44:48 -07:00
LC 585286a617 IRValidation: Remove unused members from BlockInfo
HasExit is assigned to but never used, but we check this condition a
different way right after leaving the main loop anyway.
2026-07-24 16:00:20 -04:00
LC 4904fd43e9 IRValidation: Make BlockInfo private
This isn't used outside the context of the pass.
2026-07-24 16:00:20 -04:00
LC 9e3f287c1f IRValidation: Use C instead of CW
This op isn't mutated anywhere in the pass.
2026-07-24 16:00:20 -04:00
LC f5e4e26e08 IRValidation: Move var closer to usage
Same behavior, just more compact.
2026-07-24 16:00:20 -04:00
LC 6c5a39e164 IRValidation: Turn ORs with true into assignment
These are just unconditional setting to true anyway.
2026-07-24 16:00:17 -04:00
FrontMage fcf9fd77d7 FEXCore: Isolate multiblock error state per block 2026-07-23 16:52:19 +08:00
Ryan Houdek 0589d9b872 Merge pull request #5779 from lioncash/x87
x87StackOptimizationPass: Minor cleanup
2026-07-22 18:58:17 -07:00
LC 6d4c80adff x87StackOptimizationPass: Remove IR member
This is only used in the store helpers, so we can just pass it in
directly
2026-07-23 02:26:46 -04:00
LC 56dd470528 x87StackOptimizationPass: Remove unnecesary return in Run()
It's a void function, so we don't need this at the end
2026-07-23 02:21:26 -04:00
LC 6741f53d87 x87StackOptimizationPass: Mark getValidMask()/getInvalidMask() as const
These don't modify instance state.
2026-07-23 02:18:33 -04:00
LC 19550c5417 x87StackOptimizationPass: Pass by const reference in setTop()
Avoids redundant copies. Just a minor codegen saving.
2026-07-23 02:17:02 -04:00
LC 10a0fe2e71 SharedCodeBufferManager: Make AllocateNew() signature consistent with declaration 2026-07-23 01:42:41 -04:00
LC c9add0d292 SharedCodeBufferManager: Hoist prctl define into util header
Same behavior, but just moves the potential define to be alongside all
of the others in the wrapper header.
2026-07-23 01:40:49 -04:00
LC c239d09ea0 SharedCodeBufferManager: Add missing header
Ensures the page size define is always visible.
2026-07-23 00:16:56 -04:00
Ryan Houdek d2c92808f5 FEXCore: Split out CodeBuffer management to its own file
NFC

- Renames CodeBufferManager to SharedCodeBufferManager to be more
  explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
  about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
  of the CPUBackend code

Makes it easier to parse ownership and lifetime semantics of these
buffers.
2026-07-20 18:09:29 -07:00
LC 2464633431 Merge pull request #5776 from Sonicadvance1/190
JIT: Remove JIT detection string
2026-07-20 21:07:23 -04:00
Ryan Houdek fe1ac1bc1d JIT: Remove JIT detection string
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
2026-07-20 17:44:15 -07:00
Ryan Houdek 9edd27b214 JIT: Rename temporary CPU buffer allocator
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.

NFC
2026-07-20 17:32:03 -07:00
Ryan Houdek eb7e02ea1d Merge pull request #5772 from lioncash/pass
PassManager: Simplify initialization interface
2026-07-19 16:28:07 -07:00
LC 53befc68c9 PassManager: Ensure GetPass() only queries the underlying pass mappings
Previously this would create an entry in the map if it didn't exist.
2026-07-21 12:10:26 -04:00
LC 19f95d89ec PassManager: Add basic documentation 2026-07-21 12:10:26 -04:00
LC ecb9b3b7f8 PassManager: Constrain GetPass() template to Pass-derived objects
Makes the particular conversion types constrained to catch any trivial
misuses.
2026-07-21 12:10:26 -04:00
LC d619e36523 PassManager: Pass string by const reference where applicable
Gets rid of potential extraneous copies. We also add handling for cases
where two passes with the same name are unintentionally added.
Previously we'd blindly overwrite the mapping.
2026-07-21 12:09:16 -04:00
LC fa80d11960 PassManager: Remove SyscallHandler member
This isn't used anymore, so we can get rid of it to further simplify
initialization.
2026-07-21 11:28:12 -04:00
LC 2893d2b64f PassManager: Simplify pass initialization
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
2026-07-21 11:28:09 -04:00
Ryan Houdek 3bd4d244a4 Merge pull request #5771 from lioncash/fmt
Externals: Update fmt to 12.2.0
2026-07-19 11:50:58 -07:00
LC 9888de25fe Externals: Update fmt to 12.2.0
Keeps fmt up to date.
2026-07-21 09:32:23 -04:00
LC 22bd10f3b1 CPUBackend: Remove unnecessary reinterpret_casts
This both take a void*, so the casting is unnecessary to begin with,
since this would occur anyway without it. We can also avoid a
duplication to reduce line noise.
2026-07-21 08:22:17 -04:00
Ryan Houdek f374b4775a Merge pull request #5769 from lioncash/bound
Core: Remove unnecessary bounds check in GenerateIR()
2026-07-18 21:35:34 -07:00
LC ff213bbc5e Core: Move vars closer to usage scope in GenerateIR()
Makes it so their purpose is more easily seen
2026-07-21 04:44:47 -04:00