Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.
Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.
8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.
Both scalar and vector wide.
Fixes#5326
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.
The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
Enables memcpy optimization of 80bit floats on reduced precision.
Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.
If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.
Correct the implementation to use the correct register size for FMA.
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.
Before:
text data bss dec hex filename
4160559 1471360 4336824 9968743 981c67 Bin/FEX
After:
text data bss dec hex filename
4159927 1471360 4336824 9968111 9819ef Bin/FEX
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.
Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.
This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.
Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
These get stored based on the current rotation of TOP. So we need to be
a bit careful with how we do this storing. A smidge of overhead, but
nothing unexpected.
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.
If we packed this would get to a further 62 bytes but probably not worth
it.
x87 80-bit loads, both BCD and regular tword were loading 128-bits of
data when reduced precision was enabled. This was an oversight from the
previous fix a while ago.
Adds a specific reduced precision test for this, and updates the current
test to ensure stores are still tested as well.
This is now cached internally in the X87 pass, and all operations
outside of that that use it are rare so can afford loading/storing
directly from context after flushing x87 regs.
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.
Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands