Commit Graph
849 Commits
Author SHA1 Message Date
Ryan Houdek 223e0f4e53 Merge pull request #5558 from lioncash/blend
[SVE256] Handle SSE insertions for blends
2026-06-16 22:34:32 -07:00
LC 0258fcb116 [SVE256] Handle SSE insertions for blends
See #3799
2026-06-17 01:05:46 -04:00
LC 3272aa3f08 [SVE256] Handle SSE insertions for ROUNDPD/ROUNDPS
See #3799
2026-06-17 00:42:20 -04:00
LC 32b11603d8 [SVE256] Handle SSE insertions for PMOVSX/PMOVZX ops 2026-06-17 00:04:03 -04:00
LC 538fd2672d [SVE256] Handle SSE insertions for PSADBW
See #3799
2026-06-17 00:04:03 -04:00
LC 3e37724e3e [SVE256] Handle SSE insertions for PHADDSW
See #379
2026-06-17 00:04:03 -04:00
LC 1b249ba76b [SVE256] Handle SSE insertion for PHSUBD/PHSUBW/PHSUBSW 2026-06-17 00:04:03 -04:00
LC df1295fbd0 [SVE256] Handle SSE insertions for HSUBPD/HSUBPS
See #3799
2026-06-17 00:04:03 -04:00
LC dcd71fe126 [SVE256] Handle SSE insertions for PMULHW/PMULHRSW
See #3799
2026-06-17 00:04:00 -04:00
LC 610ee5db76 [SVE256] Handle SSE insertions for PMADDUBSW
See #3799
2026-06-16 23:02:25 -04:00
Ryan Houdek 9d5494d9f0 Merge pull request #5555 from lioncash/alu
[SVE256] Vector: Handle SSE insertion properly for various ALU operations
2026-06-16 19:56:38 -07:00
Ryan Houdek dd44bc8d00 Merge pull request #5554 from lioncash/vmov
OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
2026-06-16 19:47:01 -07:00
Ryan Houdek 110c7cb62b Merge pull request #5553 from lioncash/bind
OpcodeDispatcher: Make use of Bind consistently
2026-06-16 19:44:44 -07:00
LC 1d8b6df630 [SVE256] Vector: Handle SSE insertion properly for various ALU operations
See #3799 for the bulk of the issue explanation. Ensures that emulated
SSE operation on aarch64 don't end up obliterating the upper 128-bit
lane when SVE-256 is present (Adv. SIMD operations zero-extend)

Knocks out quite a few SSE instructions right off the jump.
2026-06-16 22:26:28 -04:00
LC 195058752e OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
Will make removing the TODO in LoadSource regarding partial loads a
little easier.
2026-06-16 17:47:01 -04:00
LC c62805e86d OpcodeDispatcher: Make use of Bind consistently
We had a few places that were using Bind, and a few other places
that were using specializations as a means to composing the instruction
tables. Instead, we can just use Bind consistently, which lets us tidy
up a bunch of the implementations (and gets rid of some unnecessary
codegen).
2026-06-16 15:07:59 -04:00
LC d15b175c33 OpcodeDispatcher: Move a few stray literal accesses to Literal()
Same core behavior, but ensures that the immediates are valid literals
when assertions are enabled.
2026-06-16 11:50:48 -04:00
Ryan Houdek fef5a98602 Fix build failure. 2026-05-22 15:33:41 -07:00
Daniel Lu 5cce65cdfa OpcodeDispatcher: Decode INVD and WBINVD through privileged op handling 2026-05-22 15:25:22 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Paris Oplopoios c98a30f82b Fix FALSE_OS size 2026-02-24 23:51:25 +02:00
Paris Oplopoios 427206abf7 Fix TRUE_US size 2026-02-24 23:50:40 +02:00
Paris Oplopoios 2eb4267e6e Fix TRUE_US case 2026-02-24 23:36:09 +02:00
Paris Oplopoios f547b63bca Fix NGT_US vector case 2026-02-24 22:33:46 +02:00
Paris Oplopoios fd71928e3a Fix EQ_UQ cases 2026-02-24 22:15:51 +02:00
Ryan Houdek b275068569 FEXCore: Fixes VEX float compare operations
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.

8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.

Both scalar and vector wide.
Fixes #5326
2026-02-23 15:45:44 -08:00
Ryan Houdek d6d4f84c3b AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
2026-02-21 15:24:13 -08:00
Ryan Houdek 217bbf423b FEXCore: Switch constant emission to default to NoPad
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.

The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
2025-12-29 11:45:51 -08:00
Ryan Houdek b794b9ed2c Core/Vector: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 3fd86a953b Core/AVX_128: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 5e782cc1c2 Core/Flags: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 0ff3fb7f47 Core/X87: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek aa631c5585 Core/X87F64: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek f63ba7e3be OpcodeDispatcher/Vector: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 54dca47e09 OpcodeDispatcher/X87: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 3025a10808 FEXCore: Revert literal optimization from #4884
This no longer does anything due to #5123
2025-12-22 10:00:55 -08:00
Billy Laws 2878583627 OpcodeDispatcher: Support relocated operand type variants
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
2025-12-22 16:19:02 +00:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
Ryan Houdek 854a741ea4 OpcodeDispatcher: Fix #4982
Forgot to move the OpcodeDispatcher
2025-10-17 14:01:47 -07:00
Ryan Houdek 5f390c16be OpcodeDispatcher: Fixes Scalar FMA size calculation
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.

Correct the implementation to use the correct register size for FMA.
2025-10-09 15:18:26 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00