With the change from #4538 I had accidentally broken x87 reduced
precision.
This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.
Fixes Steam when x87 reduced precision is enabled.
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.
Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed
As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
- 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
- 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
- 7.3% -> 5.2% code buffer space used for RIP reconstruction
Fixes#4535
A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher
This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.
This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
TMP4 was used before we passed in a tmp register. Now use that temp
register.
Also return the amount of stack used on the push function. This will be
used in a bit.
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.
In addition, there a couple of clang-tidy fixes which should be NFC.
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.
This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
Ensures that blocks always start with the same state independently of predecessors
which allows independent compilation of blocks.
Starting in the X87 state is better than starting in MMX state because
MMX state is more work to initialize.
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.
Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.
So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.
This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.
This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.
This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
It is common practice for games to use CPUID as an instruction barrier
for various reasons. Ensure that we respect this by adding support for
an instruction barrier.
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.
Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
In the Push/Pop CalleeSavedRegisters these vectors were getting created
on the heap, allocating memory and then just iterating them.
Just use a std::array which makes it stop allocating memory and saves
the number of instructions.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
These three options are mutually exclusive with each other and could
potentially result in invalid encodings of the table on accident.
Change over to a 2-bit bitfield to encode if the operand that consumes
the VEX option is none, destination, 1st src, or 2nd src.
This ensures the table can't ever be incorrectly encoded.