Commit Graph
952 Commits
Author SHA1 Message Date
Simon Scherer d1b2333195 OpcodeDispatcher: Clear C1 to 0 for X87ModifySTP (fdecstp and fincstp) 2026-09-04 07:22:05 +02:00
jubecker 61d94ac117 Potential improvement for *PMULHRSW 2026-09-02 14:31:18 -07:00
Simon Scherer 4a7fda4516 OpcodeDispatcher: Preserve IE exception flag for FCOMI and FTST 2026-09-02 07:42:04 +02:00
LC 62c9c130c4 Merge pull request #5864 from simon902/pf2id_saturation
Fix pf2id overflow saturation
2026-08-27 09:16:04 -04:00
Simon Scherer b21a8dafa7 OpcodeDispatcher: Fix pf2id overflow saturation 2026-08-27 11:50:45 +02:00
Simon Scherer 5ab2c723ff OpcodeDispatcher: For PF2IWOp use VSQXTN instead of VUnZip to saturate values falling outside the 16-bit range 2026-08-27 09:36:17 +02:00
Ryan Houdek 6646a5cc72 OpcodeDispatcher: Fixes SHLD by 16 behaviour
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!

Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Simon Scherer 7f0bdf8d63 OpcodeDispatcher: Zero OF, SF and AF for FCOMI and FCOMIF64 2026-08-06 16:29:29 +02:00
Ryan Houdek c4a5ac892f AVX128: Optimize 256-bit vmovmaskpd as well
Similar to #5757, but once the elements have been zipped together, we
can treat it identically to the 128-bit 32-bit element path.

Closes #3782
2026-07-17 13:15:47 -07:00
moonfloww 3680282b30 new vmovmsk from 11 to 7 ins. 2026-07-16 13:57:12 +02:00
LC 2934b01d58 Vector: Indicate 128-bit zero vector in DefaultX87State()
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
2026-07-13 14:07:12 -04:00
LC fb2cdc8541 Crypto: Clarify zero vector size in SHA1RNDS4Op()
This ends up zeroing out the whole 128-bit vector.
2026-07-12 18:49:27 -04:00
LC 9d3c388664 Crypto: Make use of XAR in SHA1NEXTE when available
Lets us shave an instruction off on hardware that supports XAR.

Closes #5730
2026-07-12 16:02:12 -04:00
LC c7f52bcee4 Vector: Centralize masking in InsertScalarFCMPOp
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.

We can move the masking in the AVX variant too, just to be consistent.
2026-07-11 16:07:19 -04:00
Ryan Houdek 37814111de Merge pull request #5692 from lioncash/unary
Vector: Use unary handler for scalar unary insertions
2026-07-10 10:26:44 -07:00
LC 401e542fcd Vector: Use unary handler for scalar unary insertions
Same behavior, but just uses a more proper handler.
2026-07-10 07:27:54 -04:00
LC 210514f74b Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp 2026-07-10 07:24:39 -04:00
Ryan Houdek 3370d9af15 Merge pull request #5670 from simon902/MOVDoverride
Fix movd when prefixed with 0x66
2026-07-09 13:02:08 -07:00
Ryan Houdek 9306de79ad Merge pull request #5667 from simon902/CVTTSS2SIOverride
Fix cvttss2si when prefixed with 0x66
2026-07-09 12:52:00 -07:00
LC 9b7c9f0fb6 Vector: Trim one instruction off insertq
We can fold a bitwise not and and pair into a bic
2026-07-09 14:58:04 -04:00
LC 710b85be70 IR: Remove need to specify element size for vector bitwise ops
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.

Makes using these ops a little bit less noisy.
2026-07-09 13:26:51 -04:00
Simon Scherer fe08b96844 OpcodeDispatcher: Fix cvttss2si when prefixed with 0x66 2026-07-09 11:58:18 +02:00
Simon Scherer 4d78901420 OpcodeDispatcher: Fix movd when prefixed with 0x66 2026-07-09 11:46:48 +02:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
LC 20647f2287 AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
Ryan Houdek ab4e0f653a Merge pull request #5600 from lioncash/palign
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek fda023e7dc Merge pull request #5601 from lioncash/pshufb
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00