The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.
Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.
This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.
Bumps the DiskCache version again because it causes codegen to change.
This is unfiltered data that ends up in HostFeatures. If it was
serialized when set then it would effectively never allow the code
serialization to be used.
If we're shrinking the L1 cache then we just deleted the entry that we
just looked up. Add it back to ensure we don't get yet another lookup
for this entry.
Growing widens L1PointerMask without touching the table, so an entry whose
address has the new mask bit set sits where InvalidateCache no longer looks,
and a later shrink hands it back to the JIT. Wipe [0, old) on grow, the way
the shrink wipes [new, MAX); the still-valid entries go with it.
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.
Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.
Use printable characters in FOZ key names as intended.
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!
Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.
Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.
Just a smidge of cleanup, NFC.
Serializes every option, even ones that are set to default to ensure
validation that if any config option value is added or changed that they
are captured.
Skips a handful of options that are either meta options, environment
options that don't matter, or HostFeatures which is handled elsewhere.
(HostFeatures will be controlling bucketing rather than the remaining
options).
This serialization is currently 1413 bytes and generates in 17580ns on
my A1A. So it's not the fastest, definitely don't want to be generating
it constantly per process.
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.
We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.
So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.
Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.
Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.
All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.
Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.
- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
- Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
- Previous query between alignments would leave us with up to 12 bytes
unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.
Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.
Should be no functional change.
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.
Add some bounds checking to index population to protect against corruption.
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.
Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.
Shouldn't be any real behaviour change.
Serializes code blocks to disk - only blocks coming from known regions, for now
Disabled by default, key and versioning still needs work, but works for testing
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.
As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.
In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.
There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word
These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.
Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
those many contiguous bits within a single atomic word
- Allocation can fail if there are cross-word contiguous bits of the
size available
- Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
- Searches for contiguous bits of free space between multiple atomic
words.
- If found, will attempt to allocate tracking which bits were allocated
- If allocation fails, unwind bits already acquired and continue
scanning
Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
- If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
acquire semantics for scanning can technically overlook some
allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.
Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.
The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.