Justin Becker
5ec147109e
Implement peephole optimization for vpcmpeqd + vgatherdps
...
This peephole removes unnecessary masking operations when the mask is
provably all 1s. This improves performance in GB7 'Asset Compression' by
about 2% with SVE and about 4% without SVE.
2026-10-05 15:17:58 -07:00
Justin Becker
079a69dde2
Style nits
2026-09-30 15:22:22 -07:00
Justin Becker
6788bcf264
x87: complete FXAM implementation
2026-09-30 15:07:32 -07:00
Justin Becker
af8770005b
Clang format
2026-09-22 12:53:05 -07:00
Justin Becker
e2f0cf41d8
More cleanup
2026-09-22 12:53:05 -07:00
Justin Becker
638329278c
Inlined PCMPXSTRX
...
Improves performance for SSE4.2 String operations by ~10x by emitting inline ASM instead of jumping to a C++ helper.
Performance numbers from a simple microbenchmark vs the existing C++ implementation:
| | Equal Any | Ranges | Equal Each | Equal Ordered |
|---------------|-----------|--------|------------|---------------|
| **pcmpestri** | 13.40× | 9.71× | 19.38× | 14.07× |
| **pcmpestrm** | 13.12× | 9.54× | 18.73× | 13.80× |
| **pcmpistri** | 11.75× | 9.19× | 17.63× | 12.80× |
| **pcmpistrm** | 11.39× | 8.83× | 16.31× | 12.26× |
2026-09-22 12:50:02 -07:00
Simon Scherer
93def42061
OpcodeDispatcher: Fix RestoreX87State to rotate MM regs by Top of Stack
2026-09-15 16:38:36 -07:00
Ryan Houdek
1523142540
OpcodeDispatcher: Removes DST_RAX from CMP/TEST
2026-09-15 15:00:03 -07:00
Ryan Houdek
5743282944
OpcodeDispatcher: Removes DST_RAX from SBB
2026-09-15 15:00:02 -07:00
Ryan Houdek
9d936a8989
OpcodeDispatcher: Removes DST_RAX from ADC
2026-09-15 15:00:02 -07:00
Ryan Houdek
1e95181b34
OpcodeDispatcher: Removes DST_RAX from ALUOp operations.
2026-09-15 15:00:01 -07:00
Ryan Houdek
e8084e3f11
OpcodeDispatcher: Remove DST_RAX from FNSTSW
2026-09-15 15:00:00 -07:00
Ryan Houdek
ea263166c5
OpcodeDispatcher: Removes SRC_RCX from shift/rotate ops
...
These are all similar to each other with similar code. Hit them at the
same time.
2026-09-15 11:56:09 -07:00
Ryan Houdek
58e0e1a03a
OpcodeDispatcher: Removes SRC_RAX from XCHG
2026-09-14 16:03:31 -07:00
Justin Becker
89a13cd5fd
AVX-VNNI
2026-09-10 15:56:03 -07:00
Ryan Houdek
38c9f14dbd
OpcodeDispatcher: Explicitly handle no-nop vector moves
...
Just be very explicit about this rather than questioning why MMX moves
can't be a nop in the regular vector move code.
Doesn't change codegen so binarycacheversion doesn't need to change.
2026-09-04 12:19:21 -07:00
Simon Scherer
8ed4c52d58
OpcodeDispatcher: Don't skip the MMX state transition for same-register writes in MOVVectorUnalignedOp
2026-09-04 17:00:45 +02:00
Simon Scherer
7236c2bf0a
Fix x87 tag word not updating on FST ST(0) self-store
2026-09-04 11:51:12 +02:00
Simon Scherer
d1b2333195
OpcodeDispatcher: Clear C1 to 0 for X87ModifySTP (fdecstp and fincstp)
2026-09-04 07:22:05 +02:00
jubecker
61d94ac117
Potential improvement for *PMULHRSW
2026-09-02 14:31:18 -07:00
Simon Scherer
4a7fda4516
OpcodeDispatcher: Preserve IE exception flag for FCOMI and FTST
2026-09-02 07:42:04 +02:00
LC
62c9c130c4
Merge pull request #5864 from simon902/pf2id_saturation
...
Fix pf2id overflow saturation
2026-08-27 09:16:04 -04:00
Simon Scherer
b21a8dafa7
OpcodeDispatcher: Fix pf2id overflow saturation
2026-08-27 11:50:45 +02:00
Simon Scherer
5ab2c723ff
OpcodeDispatcher: For PF2IWOp use VSQXTN instead of VUnZip to saturate values falling outside the 16-bit range
2026-08-27 09:36:17 +02:00
Ryan Houdek
6646a5cc72
OpcodeDispatcher: Fixes SHLD by 16 behaviour
...
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!
Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Simon Scherer
7f0bdf8d63
OpcodeDispatcher: Zero OF, SF and AF for FCOMI and FCOMIF64
2026-08-06 16:29:29 +02:00
Ryan Houdek
c4a5ac892f
AVX128: Optimize 256-bit vmovmaskpd as well
...
Similar to #5757 , but once the elements have been zipped together, we
can treat it identically to the 128-bit 32-bit element path.
Closes #3782
2026-07-17 13:15:47 -07:00
moonfloww
3680282b30
new vmovmsk from 11 to 7 ins.
2026-07-16 13:57:12 +02:00
LC
2934b01d58
Vector: Indicate 128-bit zero vector in DefaultX87State()
...
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
2026-07-13 14:07:12 -04:00
LC
fb2cdc8541
Crypto: Clarify zero vector size in SHA1RNDS4Op()
...
This ends up zeroing out the whole 128-bit vector.
2026-07-12 18:49:27 -04:00
LC
9d3c388664
Crypto: Make use of XAR in SHA1NEXTE when available
...
Lets us shave an instruction off on hardware that supports XAR.
Closes #5730
2026-07-12 16:02:12 -04:00
LC
c7f52bcee4
Vector: Centralize masking in InsertScalarFCMPOp
...
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.
We can move the masking in the AVX variant too, just to be consistent.
2026-07-11 16:07:19 -04:00
Ryan Houdek
37814111de
Merge pull request #5692 from lioncash/unary
...
Vector: Use unary handler for scalar unary insertions
2026-07-10 10:26:44 -07:00
LC
401e542fcd
Vector: Use unary handler for scalar unary insertions
...
Same behavior, but just uses a more proper handler.
2026-07-10 07:27:54 -04:00
LC
210514f74b
Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp
2026-07-10 07:24:39 -04:00
Ryan Houdek
3370d9af15
Merge pull request #5670 from simon902/MOVDoverride
...
Fix movd when prefixed with 0x66
2026-07-09 13:02:08 -07:00
Ryan Houdek
9306de79ad
Merge pull request #5667 from simon902/CVTTSS2SIOverride
...
Fix cvttss2si when prefixed with 0x66
2026-07-09 12:52:00 -07:00
LC
9b7c9f0fb6
Vector: Trim one instruction off insertq
...
We can fold a bitwise not and and pair into a bic
2026-07-09 14:58:04 -04:00
LC
710b85be70
IR: Remove need to specify element size for vector bitwise ops
...
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.
Makes using these ops a little bit less noisy.
2026-07-09 13:26:51 -04:00
Simon Scherer
fe08b96844
OpcodeDispatcher: Fix cvttss2si when prefixed with 0x66
2026-07-09 11:58:18 +02:00
Simon Scherer
4d78901420
OpcodeDispatcher: Fix movd when prefixed with 0x66
2026-07-09 11:46:48 +02:00
LC
6bc67609a3
[SVE256] Handle 256-bit blend operations much more efficiently
...
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC
7fd9b897c2
[SVE256] Handle 256-bit AES operations
...
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.
Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
LC
d11b19fd2b
[SVE256] Ensure insertion behavior for PCLMUL SSE operations
...
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC
684c568033
[SVE256] Ensure insertion behavior for SHA SSE operations
...
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC
ab4fb7b3ad
[SVE256] Ensure insertion behavior for AES operations on SSE
...
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
LC
c6b0f360fe
AVX: Make use of table swapping constant to trim down relevant ops
...
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
Paulo Matos
37b010795e
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 15:00:16 -07:00
LC
2cb4f8b6f5
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
...
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC
600e2ddecf
Vector: Only signify 128-bit vector loads in UCOMISxOp
...
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00