The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.
Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.
Should be no functional change.
Serializes code blocks to disk - only blocks coming from known regions, for now
Disabled by default, key and versioning still needs work, but works for testing
NFC
- Renames CodeBufferManager to SharedCodeBufferManager to be more
explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
of the CPUBackend code
Makes it easier to parse ownership and lifetime semantics of these
buffers.
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
the modifying thread. Old versions of the CodeBuffer persist as read-only
data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
decrease the reference count and eventually trigger deallocation of the
old version
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.
In addition, there a couple of clang-tidy fixes which should be NFC.
When #2722 implemented this initially and #4271 switched over to signed
int16_t there was assumptions made that int16_t was a reasonable
trade-off in encoding size versus needing to deal with 8-bit values
being too small in some cases.
In the common case we are almost always encoding 8-bit values because
instructions are typically linear (and less than 15-bytes in size), but
16-bit was chosen because optimizing JIT and multiple instructions that
don't cause exceptions can add up to larger than 8-bit.
Instead of hardcoding 16-bit values, implement a variable length integer
class where ~96.8% of values are 8-bit encoded, and the remaining 3.19% are encoded using 16-bit.
Due to some constraints that #4271 put in place, we can basically
guarantee currently that branch targets are within 16-bit. The VL class
does support 32-bit and 64-bit as well so if we change behaviour then
nothing needs to change.
Some stats when running Sonic Mania with multiblock enabled.
Encoded integers: 3,504,907
Encoded 8-bit: 3,393,095 (96.8%)
Encoded 16-bit: 111,812 (3.19%)
Encoded 32/64-bit: 0
Encoded Size: 3,615,181 bytes (3.44MiB)
Fixed encoded size: 7,007,604 bytes (6.68MiB)
Definitely worth using and saves the headache of large RIP/PC offsets
causing problems.
With multiblock enabled, host code generated from guest code with a
lower address may be placed after host code generated from guest code
with a higher address in a multiblock. As each guest RIP reconstruction
entry is always relative to the one before it the offset needs to be
signed to allow this.
guest instruction
Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
IRListView is now purely a view type. Instead, ownership is managed on-demand
by a separate interface (IRStorageBase). Materialization of IRListViews to
owning types is moved to this interface as well.
This also avoids unneeded copies of the data.
This is no longer necessary to be part of the public API. Moves the
header internally.
Needed to pass through `IsAddressInCodeBuffer` from CPUBackend through
the Context object, but otherwise no functional change.