Commit Graph
966 Commits
Author SHA1 Message Date
Justin Becker e2f0cf41d8 More cleanup 2026-09-22 12:53:05 -07:00
Justin Becker 638329278c Inlined PCMPXSTRX
Improves performance for SSE4.2 String operations by ~10x by emitting inline ASM instead of jumping to a C++ helper.

Performance numbers from a simple microbenchmark vs the existing C++ implementation:
|               | Equal Any | Ranges | Equal Each | Equal Ordered |
|---------------|-----------|--------|------------|---------------|
| **pcmpestri** | 13.40×    | 9.71×  | 19.38×     | 14.07×        |
| **pcmpestrm** | 13.12×    | 9.54×  | 18.73×     | 13.80×        |
| **pcmpistri** | 11.75×    | 9.19×  | 17.63×     | 12.80×        |
| **pcmpistrm** | 11.39×    | 8.83×  | 16.31×     | 12.26×        |
2026-09-22 12:50:02 -07:00
Simon Scherer 93def42061 OpcodeDispatcher: Fix RestoreX87State to rotate MM regs by Top of Stack 2026-09-15 16:38:36 -07:00
Ryan Houdek 1523142540 OpcodeDispatcher: Removes DST_RAX from CMP/TEST 2026-09-15 15:00:03 -07:00
Ryan Houdek 5743282944 OpcodeDispatcher: Removes DST_RAX from SBB 2026-09-15 15:00:02 -07:00
Ryan Houdek 9d936a8989 OpcodeDispatcher: Removes DST_RAX from ADC 2026-09-15 15:00:02 -07:00
Ryan Houdek 1e95181b34 OpcodeDispatcher: Removes DST_RAX from ALUOp operations. 2026-09-15 15:00:01 -07:00
Ryan Houdek e8084e3f11 OpcodeDispatcher: Remove DST_RAX from FNSTSW 2026-09-15 15:00:00 -07:00
Ryan Houdek ea263166c5 OpcodeDispatcher: Removes SRC_RCX from shift/rotate ops
These are all similar to each other with similar code. Hit them at the
same time.
2026-09-15 11:56:09 -07:00
Ryan Houdek 58e0e1a03a OpcodeDispatcher: Removes SRC_RAX from XCHG 2026-09-14 16:03:31 -07:00
Justin Becker 89a13cd5fd AVX-VNNI 2026-09-10 15:56:03 -07:00
Ryan Houdek 38c9f14dbd OpcodeDispatcher: Explicitly handle no-nop vector moves
Just be very explicit about this rather than questioning why MMX moves
can't be a nop in the regular vector move code.
Doesn't change codegen so binarycacheversion doesn't need to change.
2026-09-04 12:19:21 -07:00
Simon Scherer 8ed4c52d58 OpcodeDispatcher: Don't skip the MMX state transition for same-register writes in MOVVectorUnalignedOp 2026-09-04 17:00:45 +02:00
Simon Scherer 7236c2bf0a Fix x87 tag word not updating on FST ST(0) self-store 2026-09-04 11:51:12 +02:00
Simon Scherer d1b2333195 OpcodeDispatcher: Clear C1 to 0 for X87ModifySTP (fdecstp and fincstp) 2026-09-04 07:22:05 +02:00
jubecker 61d94ac117 Potential improvement for *PMULHRSW 2026-09-02 14:31:18 -07:00
Simon Scherer 4a7fda4516 OpcodeDispatcher: Preserve IE exception flag for FCOMI and FTST 2026-09-02 07:42:04 +02:00
LC 62c9c130c4 Merge pull request #5864 from simon902/pf2id_saturation
Fix pf2id overflow saturation
2026-08-27 09:16:04 -04:00
Simon Scherer b21a8dafa7 OpcodeDispatcher: Fix pf2id overflow saturation 2026-08-27 11:50:45 +02:00
Simon Scherer 5ab2c723ff OpcodeDispatcher: For PF2IWOp use VSQXTN instead of VUnZip to saturate values falling outside the 16-bit range 2026-08-27 09:36:17 +02:00
Ryan Houdek 6646a5cc72 OpcodeDispatcher: Fixes SHLD by 16 behaviour
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!

Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Simon Scherer 7f0bdf8d63 OpcodeDispatcher: Zero OF, SF and AF for FCOMI and FCOMIF64 2026-08-06 16:29:29 +02:00
Ryan Houdek c4a5ac892f AVX128: Optimize 256-bit vmovmaskpd as well
Similar to #5757, but once the elements have been zipped together, we
can treat it identically to the 128-bit 32-bit element path.

Closes #3782
2026-07-17 13:15:47 -07:00
moonfloww 3680282b30 new vmovmsk from 11 to 7 ins. 2026-07-16 13:57:12 +02:00
LC 2934b01d58 Vector: Indicate 128-bit zero vector in DefaultX87State()
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
2026-07-13 14:07:12 -04:00
LC fb2cdc8541 Crypto: Clarify zero vector size in SHA1RNDS4Op()
This ends up zeroing out the whole 128-bit vector.
2026-07-12 18:49:27 -04:00
LC 9d3c388664 Crypto: Make use of XAR in SHA1NEXTE when available
Lets us shave an instruction off on hardware that supports XAR.

Closes #5730
2026-07-12 16:02:12 -04:00
LC c7f52bcee4 Vector: Centralize masking in InsertScalarFCMPOp
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.

We can move the masking in the AVX variant too, just to be consistent.
2026-07-11 16:07:19 -04:00
Ryan Houdek 37814111de Merge pull request #5692 from lioncash/unary
Vector: Use unary handler for scalar unary insertions
2026-07-10 10:26:44 -07:00
LC 401e542fcd Vector: Use unary handler for scalar unary insertions
Same behavior, but just uses a more proper handler.
2026-07-10 07:27:54 -04:00
LC 210514f74b Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp 2026-07-10 07:24:39 -04:00
Ryan Houdek 3370d9af15 Merge pull request #5670 from simon902/MOVDoverride
Fix movd when prefixed with 0x66
2026-07-09 13:02:08 -07:00
Ryan Houdek 9306de79ad Merge pull request #5667 from simon902/CVTTSS2SIOverride
Fix cvttss2si when prefixed with 0x66
2026-07-09 12:52:00 -07:00
LC 9b7c9f0fb6 Vector: Trim one instruction off insertq
We can fold a bitwise not and and pair into a bic
2026-07-09 14:58:04 -04:00
LC 710b85be70 IR: Remove need to specify element size for vector bitwise ops
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.

Makes using these ops a little bit less noisy.
2026-07-09 13:26:51 -04:00
Simon Scherer fe08b96844 OpcodeDispatcher: Fix cvttss2si when prefixed with 0x66 2026-07-09 11:58:18 +02:00
Simon Scherer 4d78901420 OpcodeDispatcher: Fix movd when prefixed with 0x66 2026-07-09 11:46:48 +02:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00