vblendvps, vblendvpd, vpblendvb all broke because of failing to zext the
128-bit register correctly.
vinsertps broke because it was accidentally using the wrong
implementation.
Extends each of their unittests to handle these cases.
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.
Also removed the IR node F80XTRACTStack which is not needed anymore.
These had completely forgotten to load segment prefixes on the element
loadstore operations. This had caused a crash on the Linux version of
"Halls of Tormet" (steamid 2218750).
This game has TLS vector data which it was loading with a movhps and hit
this path.
Only doing the single table for review purposes. Once reviewed I will
hammer out the remaining tables.
Similar to #3320, most of the OpcodeDispatcher tables can be consteval
and made to be a compile time constant. This just requires shuffling the
code slightly. The idea is to get almost all of the table setup out of
the `InstallOpcodeHandlers` function and instead only install the
handlers that change based on 32-bit or 64-bit, just like the x86 tables
we also did.
This is one of the most expensive headers in FEX, averaging 1.2 seconds
of compile time per include.
We include this header 31 times around the project. Remove
the five instances where it is unused to help this issue. Next step
would be to make the header lighter if possible.
When the shift amount is >= 16-bytes then we need to zero the register.
We had a bug where we were assigning `Result.High` to itself, which
effectively made the top 128-bits of the ymm register not modify itself.
Adds a unit test to ensure that doesn't happen again.
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.
PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".
The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.
Fixes Aperture Desk Job in FEX.