Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.
The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
Enables memcpy optimization of 80bit floats on reduced precision.
Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.
If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.
Correct the implementation to use the correct register size for FMA.
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.
Before:
text data bss dec hex filename
4160559 1471360 4336824 9968743 981c67 Bin/FEX
After:
text data bss dec hex filename
4159927 1471360 4336824 9968111 9819ef Bin/FEX
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.
Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.
This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.
Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
These get stored based on the current rotation of TOP. So we need to be
a bit careful with how we do this storing. A smidge of overhead, but
nothing unexpected.
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.
If we packed this would get to a further 62 bytes but probably not worth
it.
x87 80-bit loads, both BCD and regular tword were loading 128-bits of
data when reduced precision was enabled. This was an oversight from the
previous fix a while ago.
Adds a specific reduced precision test for this, and updates the current
test to ensure stores are still tested as well.
This is now cached internally in the X87 pass, and all operations
outside of that that use it are rare so can afford loading/storing
directly from context after flushing x87 regs.
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.
Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
Similar to the previous far jmp, if the CS changes operating mode then
things will still explode with other FEX asserts. But this gets another
change out of my stashes.
These instructions go hand-in-hand obviously so they get implemented as
a pair.