Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.
This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.
Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
These get stored based on the current rotation of TOP. So we need to be
a bit careful with how we do this storing. A smidge of overhead, but
nothing unexpected.
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.
If we packed this would get to a further 62 bytes but probably not worth
it.
x87 80-bit loads, both BCD and regular tword were loading 128-bits of
data when reduced precision was enabled. This was an oversight from the
previous fix a while ago.
Adds a specific reduced precision test for this, and updates the current
test to ensure stores are still tested as well.
This is now cached internally in the X87 pass, and all operations
outside of that that use it are rare so can afford loading/storing
directly from context after flushing x87 regs.
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.
Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
Similar to the previous far jmp, if the CS changes operating mode then
things will still explode with other FEX asserts. But this gets another
change out of my stashes.
These instructions go hand-in-hand obviously so they get implemented as
a pair.
now obsolete!
Results for the whole series are excellent:
Difference at 95.0% confidence
-3.97603% +/- 0.254656%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In b8dd5d95b ("OpcodeDispatcher: optimize X87FTWTag"), we optimized
X87FTWTag using an efficient Morton interleave operation. Here, we do the
inverse, optimizing SetX87FTW using an efficient Morton deinterleave
operation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.
This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!
But now as we are no longer utilizing it, it is time to remove it.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.