We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.
Just a smidge of cleanup, NFC.
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.
We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.
So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.
Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.
Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.
All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.
Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.
Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.
Should be no functional change.
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.
Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.
Shouldn't be any real behaviour change.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.
As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.
It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
Fixes#5230
Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.
Note: I have no idea how fmt will handle that format string with an
array of uint8_t.
Signed-off-by: crueter <crueter@eden-emu.dev>
This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.
Not that big of a deal, it's only a minor optimization anyway.
Fixes#5227
Instead of just a trivial pad being on or off, support a tri-state
on/off/auto where on will always pad, off will never pad, and auto will
pad only if code caching is enabled.
Further augment this by allowing a byte-width to be passed in, which can
be used with pointers to force a 48-bit VA width to only ever pad to
three instructions, reducing the common worst-case situation from 4
instructions to 3. This works because we're not going to expose a VA
width larger than 47-bit to the guest.
Fixes the handful of use-cases that explicitly chose their NOP padding,
and a bug in Arm64Relocations.cpp where it was incorrectly asking to not
receive padding even though it requires it.
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.