I don't think these are worth it, and also currently they don't trigger ever.
n=100:
Difference at 95.0% confidence
-0.00245961 +/- 0.00139573
-0.524468% +/- 0.297615%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is expensive and only needed for spilling, so only do it for spilling. This
complicates the RA a bit but speeds us up on average since most blocks
don't spill. Total results of this change (including the prep commits that
slowed things down temporarily):
Difference at 95.0% confidence
-1.71952% +/- 0.455996%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
if we spill for SRA, we don't need/want to execute this code path. this will be
load bearing by the end of this series.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.
Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.
In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.
Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.
Difference at 95.0% confidence
-0.00138911 +/- 0.00104724
-0.418608% +/- 0.315587%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.
Difference at 95.0% confidence
-0.00474273 +/- 0.00119189
-1.40908% +/- 0.354114%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.
This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!
But now as we are no longer utilizing it, it is time to remove it.