Commit Graph
14901 Commits
Author SHA1 Message Date
LC af9b438eec Merge pull request #5841 from OFFTKP/blsmsk
unittests/ASM: Test BLSR/BLSMSK CF flag
2026-08-22 23:13:27 -04:00
LC fa9bbf081b Merge pull request #5843 from Sonicadvance1/210
FEXCore: Fixes ever shrinking JIT code buffer
2026-08-22 23:05:34 -04:00
Ryan Houdek db3817a260 FEXCore: Fixes ever shrinking JIT code buffer
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.

Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
2026-08-22 18:31:22 -07:00
Paris Oplopoios 635befb4c8 FEXCore: Fix CF calculation for BLSMSK and BLSR for 32-bit operands 2026-08-22 15:47:50 +03:00
Paris Oplopoios a69daa2524 unittests/ASM: Test BLSR/BLSMSK CF flag 2026-08-22 15:34:13 +03:00
LC b563703701 Merge pull request #5839 from Sonicadvance1/208
FEX: Only fsync on assert
2026-08-21 19:56:07 -04:00
Ryan Houdek 0717f4689b FEX: Only fsync on assert
The rest of the messages should be buffered, no reason to sync those.
2026-08-21 16:26:14 -07:00
LC bf30f6af4c Merge pull request #5838 from Sonicadvance1/207
JIT: Fixes some alignment logic
2026-08-21 18:07:39 -04:00
Ryan Houdek dfcbdac347 JIT: Fixes some alignment logic
Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.

- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
  - Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
  - Previous query between alignments would leave us with up to 12 bytes
    unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
2026-08-21 14:55:03 -07:00
LC 8d12f3d5aa Merge pull request #5837 from Sonicadvance1/206
SharedCodeBufferManager: Leak less internal details about implementation
2026-08-21 17:31:30 -04:00
Ryan Houdek 02f8ab5f87 SharedCodeBufferManager: Leak less internal details about implementation
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.

Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.

Should be no functional change.
2026-08-21 14:06:12 -07:00
LC b2b6b263bb Merge pull request #5835 from Plagman/plagman/cache_thread_mr
DiskCache: offload Store to a WorkQueueThread
2026-08-21 14:36:39 -04:00
Pierre-Loup A. Griffais e3f208f61f DiskCache: offload Store to a WorkQueueThread
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.

Add some bounds checking to index population to protect against corruption.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais 008d990f20 Utils: add WorkQueueThread
Straightforward queue for arbitrary work.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais c038bb5794 Windows: add Thread implementation
So we can make a worker thread that the guest (hopefully) won't see.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais efc5eaf9c7 Utils/File: add PRead(), PWrite() and Size()
Because PRead/PWrite don't have the same side effects on the file cursor
between Linux and Windows (:/), make them all-or-nothing.
2026-08-21 11:16:07 -07:00
Ryan Houdek c1e9d19809 Merge pull request #5836 from cjacek/int-3
Windows: Handle interrupt 3 in HandleGuestException
2026-08-21 09:30:01 -07:00
Jacek Caban 85efb47a3b Windows: Handle interrupt 3 in HandleGuestException 2026-08-21 13:16:43 +02:00
LC bc69693b6c Merge pull request #5834 from Sonicadvance1/205
FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
2026-08-20 19:47:47 -04:00
Ryan Houdek c9c5a75b76 FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.

Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.

Shouldn't be any real behaviour change.
2026-08-20 16:35:28 -07:00
Ryan Houdek f50279a2e7 Merge pull request #5832 from Plagman/plagman/cache_mr
Disk Cache initial implementation
2026-08-20 12:06:59 -07:00
Pierre-Loup A. Griffais 561c32b45d Disk Cache initial implementation
Serializes code blocks to disk - only blocks coming from known regions, for now

Disabled by default, key and versioning still needs work, but works for testing
2026-08-19 18:20:49 -07:00
Ryan Houdek f42dc71972 Merge pull request #5831 from cjacek/int-assert
Windows: Handle 0x2c interrupt in HandleGuestException
2026-08-18 16:39:39 -07:00
Jacek Caban ea67665bbf Windows: Handle 0x2c interrupt in HandleGuestException 2026-08-19 00:41:07 +02:00
LC c3b4d4b7bb Merge pull request #5829 from Sonicadvance1/204
unittests: Disable siglongjmp_branch_invalid on 32-bit
2026-08-17 23:21:41 -04:00
Ryan Houdek 492dac719f unittests: Disable siglongjmp_branch_invalid on 32-bit
This test is trying to execute "invalid" code from the last two pages of
the address space to the first two pages of the address space. But
failed to noticed that the last two pages of the 32-bit x86 address
space are actually valid, usually containing VDSO things.
2026-08-17 18:22:03 -07:00
Ryan Houdek 9377bac5e7 Merge pull request #5825 from javelina-pkwy/fix/siglongjmp-deadlock
FEXCore: prevent deadlock when branching to MAX_UINT64
2026-08-17 17:55:39 -07:00
Justin Becker e0bbfdda84 Cleaner change 2026-08-17 12:05:24 -07:00
Ryan Houdek 9618b5adef Merge pull request #5828 from Sonicadvance1/203
Steam/CompatTool: Become more picky about configs
2026-08-17 10:22:19 -07:00
Ryan Houdek a3d609ebb1 Merge pull request #5820 from javelina-pkwy/lea-reg-reg
Decoder: fix illegal LEA encoding
2026-08-17 09:58:47 -07:00
Ryan Houdek 2b0d94536f Steam/CompatTool: Become more picky about configs
If `STEAM_COMPAT_FEX_CONFIG` is missing options, then instead of having
an opinion about what those options should be, just leave them unset.
This allows FEX's regular default option handling to kick in for missing
configuration options.

Where previously if an option was missing from the config, it would
default to boolean false, which may or may not be the default depending
on option.
2026-08-17 09:34:25 -07:00
LC 73ab3bc56d Merge pull request #5826 from cjacek/crt-printf
Windows/CRT: Add printf and puts stubs
2026-08-16 10:07:21 -04:00
Jacek Caban 412c49a8ce Windows/CRT: Add printf and puts stubs
Fixes PE builds with VIXL disassembler enabled.
2026-08-16 15:14:36 +02:00
Justin Becker 6f29dfcbb8 Probe before taking lock in Compile*() 2026-08-13 16:53:05 -07:00
LC f3ab82a73f Merge pull request #5823 from Sonicadvance1/202
SharedCodeBufferManager: Allocate JIT space atomically.
2026-08-13 18:15:59 -04:00
Ryan Houdek 6734c9ed3e SharedCodeBufferManager: Allocate JIT space atomically.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.

As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
2026-08-13 14:58:01 -07:00
Ryan Houdek 71afe47675 Merge pull request #5821 from Sonicadvance1/200
Thunks: Adds some new PV paths
2026-08-12 22:47:56 -07:00
LC 7075377a63 Merge pull request #5822 from Sonicadvance1/201
clang-format: Slight whitespace difference
2026-08-12 20:13:16 -04:00
Ryan Houdek 7c1036df09 clang-format: Slight whitespace difference 2026-08-12 16:55:38 -07:00
Ryan Houdek a312347589 Thunks: Adds some new PV paths
Slight PV behaviour changes meant we missed this.
2026-08-12 16:54:01 -07:00
LC f386c62dba Merge pull request #5815 from Sonicadvance1/198
FEXCore/Utils: Implements a new atomic segmented bitmap allocator
2026-08-12 17:58:19 -04:00
Ryan Houdek 30a81484d7 FEXCore/unittests: Adds tests for atomic bitmap allocator 2026-08-12 14:32:31 -07:00
Ryan Houdek 00a6b046a8 FEXCore/Utils: Implements a new atomic segmented bitmap allocator
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.

In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.

There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
2026-08-12 14:32:31 -07:00
Justin Becker c156498c5c Add 16 bit and 32 bit variants 2026-08-11 16:19:49 -07:00
LC adea3e410f Merge pull request #5793 from Sonicadvance1/193
FEXCore: Adds a lock-free atomic bitset that supports contiguous range allocations
2026-08-10 17:24:37 -04:00
Ryan Houdek 40940ae0b3 FEXCore/unittests: Adds atomic bitset unittest
Hammers the API in a couple of ways to make sure it works.
2026-08-10 14:12:10 -07:00
Ryan Houdek 1bab28dad4 FEXCore: Adds a lock-free atomic bitset that supports contiguous range allocations
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word

These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.

Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
  those many contiguous bits within a single atomic word
  - Allocation can fail if there are cross-word contiguous bits of the
    size available
  - Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
  - Searches for contiguous bits of free space between multiple atomic
    words.
  - If found, will attempt to allocate tracking which bits were allocated
  - If allocation fails, unwind bits already acquired and continue
    scanning

Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
  - If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
  overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
  acquire semantics for scanning can technically overlook some
  allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.

Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.

The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.
2026-08-10 14:12:09 -07:00
Ryan Houdek 4838265589 FEXCore/MathUtils: Adds helper for alignment by power of 2 size
Useful for removing integer division instructions when we know the
source value is aligned to be power of two. As integer division is quite
slow, we want to use this when possible.
2026-08-10 14:08:07 -07:00
Ryan Houdek f6d20a1a88 Merge pull request #5819 from cjacek/clang-warnings
Fix warnings in llvm-mingw builds
2026-08-10 10:37:40 -07:00
Jacek Caban 11be444d45 toolchain_mingw: Don't use -static-libgcc -static-libstdc++
Those are not supported by Clang and cause -Wunused-command-line-argument warnings.
They are also redundant when -static is used.
2026-08-10 15:05:26 +02:00