We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.
If we packed this would get to a further 62 bytes but probably not worth
it.
Unity games crash with TSO disabled due to the SPSC GfxDevice
ThreadedStreamBuffer read/write pointer updates missing acq/rel
semantics. Rather than attempting to use heuristics to match cases like
this, which could be overzealous and hit more accesses than necessary, just
target the problem directly as this is consist across 32/64 bit Unity versions
for at least the past 10 years. Gate this behind the existing Unity mono
hacks to avoid false-positives in non-Unity games.
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
A wild use case of union over a variant because we don't want to
increase the encoding size from 128-bit to 256-bit (because of padding).
The type of operation is encoded with the table operation type, so a
variant is unnecessary and we get to keep the 128-bit encoding.
This allows us to have "recursive" x86 table descriptions. But in
reality this is going to only be one layer deep. As this will allow the
Frontend decoder to select instruction encodings based on arch bitness
once the tables are generated correctly.
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.
Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
Now executable page tracking is implemented, this limit is technically
unnecessary and effectively never hit in normal code. Keep a reasonable
limit however to avoid accidentally inlining tail calls and exploring
dead branches in obfuscated code.
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.
Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.
This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.
Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
These three options are mutually exclusive with each other and could
potentially result in invalid encodings of the table on accident.
Change over to a 2-bit bitfield to encode if the operand that consumes
the VEX option is none, destination, 1st src, or 2nd src.
This ensures the table can't ever be incorrectly encoded.
With the prior approach, backwards jumps into existing blocks would
explore the overlapping part rather than splitting the block, generating
needless code and wasting time decoding. Similarly, the current block
wouldn't be split when it is extended to overlap with a pending jump target.
Solve this by tracking blocks in a sorted vector and splitting existing blocks
on jumps when appropriate, in order to avoid any possibility of overlapping
blocks, which would break the lookup, misaligned and zero instruction blocks
are disallowed.