In b8dd5d95b ("OpcodeDispatcher: optimize X87FTWTag"), we optimized
X87FTWTag using an efficient Morton interleave operation. Here, we do the
inverse, optimizing SetX87FTW using an efficient Morton deinterleave
operation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.
This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!
But now as we are no longer utilizing it, it is time to remove it.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
When the shift size is exactly 16bytes, then it turns in to a move.
If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.
Fixes#3791.
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes#4319.
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.