It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.
Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.
Shouldn't be any real behaviour change.
This test is trying to execute "invalid" code from the last two pages of
the address space to the first two pages of the address space. But
failed to noticed that the last two pages of the 32-bit x86 address
space are actually valid, usually containing VDSO things.
If `STEAM_COMPAT_FEX_CONFIG` is missing options, then instead of having
an opinion about what those options should be, just leave them unset.
This allows FEX's regular default option handling to kick in for missing
configuration options.
Where previously if an option was missing from the config, it would
default to boolean false, which may or may not be the default depending
on option.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.
As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.
In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.
There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word
These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.
Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
those many contiguous bits within a single atomic word
- Allocation can fail if there are cross-word contiguous bits of the
size available
- Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
- Searches for contiguous bits of free space between multiple atomic
words.
- If found, will attempt to allocate tracking which bits were allocated
- If allocation fails, unwind bits already acquired and continue
scanning
Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
- If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
acquire semantics for scanning can technically overlook some
allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.
Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.
The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.
Useful for removing integer division instructions when we know the
source value is aligned to be power of two. As integer division is quite
slow, we want to use this when possible.
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.
Closes#5805 although their (ML?) write-up was incorrect.
While it would be better to fix the error in the source, it shouldn't be
a case of blocking release. Print the line that was failed to parse and
then continue onwards.
In particular hit by `ERROR:root:Failure to parse Thread shared code buffer management`
NFC
- Renames CodeBufferManager to SharedCodeBufferManager to be more
explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
of the CPUBackend code
Makes it easier to parse ownership and lifetime semantics of these
buffers.
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.
NFC
This used to be used for the intrusively allocated `LiveVMARegion` but
that is all handled internally to the object now, making this
unnecessary. It was always receiving zero and doing nothing so just
remove it.