LC
a6e7fba433
AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
...
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.
In some cases, this can be quite drastic, like with:
vpblendw ymm0, ymm1, ymm2, 0b00000001
being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC
8905e39439
OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
...
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC
df41b85827
OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
...
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC
9fa3b9345e
AVX: Remove unnecessary moves from PINSRX ops
...
These are old paths still around from when StoreResult used to
automatically perform truncating moves.
These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
Ryan Houdek
1619374252
Merge pull request #5587 from lioncash/pd
...
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
LC
42af6c8508
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
...
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC
b7df1bc259
[SVE256] Remove unnecessary move in VCVTPS2PD
...
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC
b40f9db735
[SVE256] Remove heavy handed moves from scalar compares
...
(See #3799 )
I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)
This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC
8989ce1766
[SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
...
See #3799
2026-06-19 18:53:50 -04:00
Ryan Houdek
9d0c05d9cc
Merge pull request #5575 from lioncash/movq2dq
...
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek
f6a68cb7fd
Merge pull request #5574 from lioncash/sse4a
...
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
LC
faa121e9ef
Vector: Move MOVQ2DQ over to Bind
...
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC
f48759e83a
[SVE256] Handle SSE insertions for MOVQ2DQ
...
See #3799
2026-06-19 17:42:21 -04:00
LC
2eca733603
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
...
See #3799
2026-06-19 17:05:59 -04:00
LC
14b65cec43
[SVE256] Handle SSE insertions for CVTPI2PD
...
See #3799
CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
LC
edd044752d
[SVE256] Handle SSE insertions for PMADDWD
...
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek
3a23bb4b73
Merge pull request #5569 from lioncash/cmp
...
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00
LC
99baa4f3d9
[SVE256] Handle SSE insertions for CMPSD/CMPSS
...
See #3799
2026-06-17 21:05:31 -04:00
LC
f64d4c571b
[SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
...
See #3799
2026-06-17 20:38:58 -04:00
LC
9372fa169a
[SVE256] Handle SSE insertions for MOVSD/MOVSS
...
See #3799
2026-06-17 20:09:17 -04:00
LC
daaa6ec129
[SVE256] Handle SSE insertions for aligned and unaligned moves
...
See #3799
2026-06-17 19:39:06 -04:00
LC
88afc22d5b
[SVE256] Handle SSE insertions for MOVNTDQA
...
See #3799
2026-06-17 19:38:57 -04:00
LC
b7ea9e30df
[SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
...
See #3799
Gets a few of the moves out of the way.
2026-06-17 17:39:13 -04:00
LC
844c3bb197
[SVE256] Handle SSE insertions for XOR special case
...
See #3799
Ensures that our special case maintains insertion behavior
2026-06-17 16:33:05 -04:00
LC
208c6d3eac
[SVE256] Handle SSE insertions for vector unary ops
...
See #3799
2026-06-17 14:43:56 -04:00
LC
886a2e74ac
[SVE256] Handle SSE insertions for pack ops
2026-06-17 13:36:18 -04:00
LC
db1d90ec9d
[SVE256] Handle SSE insertions for shuffles
...
See #3799
2026-06-17 08:44:45 -04:00
LC
124ce8420a
[SVE256] Handle SSE insertions for PINSR(B,D,Q,W)
...
See #3799
2026-06-17 08:22:26 -04:00
LC
4433eaf242
[SVE256] Handle SSE insertions for INSERTPS
...
See #3799
2026-06-17 08:11:49 -04:00
LC
5ba070f600
[SVE256] Handle SSE insertions for PSIGN(B,D,W)
...
See #3799
2026-06-17 07:59:33 -04:00
LC
2d1a42aa00
[SVE256] Handle SSE insertions for shifts
...
See #3799
2026-06-17 06:06:08 -04:00
LC
74d9f5a3e2
[SVE256] Handle SSE insertions for MOVDDUP
...
See #3799
2026-06-17 05:13:08 -04:00
LC
634fbb5a73
[SVE256] Handle SSE insertions for Float->Int/Int->Float conversions
...
See #3799
2026-06-17 04:44:39 -04:00
LC
e16948bf80
[SVE256] Handle SSE insertions for CVTPD2PS/CVTPS2PD
...
See #3799
2026-06-17 03:51:36 -04:00
LC
76c9833ee5
[SVE256] Handle SSE insertions for CMPPD/CMPPS
...
See #3799
2026-06-17 03:33:26 -04:00
Ryan Houdek
99662b70ff
Merge pull request #5560 from lioncash/psad
...
[SVE256] Handle SSE insertions for more misc ops
2026-06-17 00:14:30 -07:00
LC
fcde9eabbf
[SVE256] Handle SSE insertions for PALIGNR
...
See #3799
2026-06-17 02:13:37 -04:00
LC
4ef15951a8
[SVE256] Handle SSE insertions for PACKSS/PACKUS ops
...
See #3799
2026-06-17 02:08:24 -04:00
LC
1bc51c2290
[SVE256] Handle SSE insertions for PMULUDQ
...
See #3799
2026-06-17 02:00:11 -04:00
LC
be1025901a
[SVE256] Handle SSE insertions for ADDSUBPD/ADDSUBPS
...
See #3799
2026-06-17 01:54:55 -04:00
Ryan Houdek
1a606de29f
Merge pull request #5559 from lioncash/phmin
...
[SVE256] Handle SSE insertions for PHMINPOSUW, DPPD, and DPPS
2026-06-16 22:50:53 -07:00
LC
454c0b31cb
[SVE256] Handle SSE insertions for MPSADBW
...
See #3799
2026-06-17 01:47:21 -04:00
Ryan Houdek
223e0f4e53
Merge pull request #5558 from lioncash/blend
...
[SVE256] Handle SSE insertions for blends
2026-06-16 22:34:32 -07:00
LC
3a84091945
[SVE256] Handle SSE insertions for DPPD/DPPS
...
See #3799
2026-06-17 01:32:36 -04:00
LC
a0e8f1097f
[SVE256] Handle SSE insertions for PHMINPOSUW
...
See #3799
2026-06-17 01:21:19 -04:00
LC
0258fcb116
[SVE256] Handle SSE insertions for blends
...
See #3799
2026-06-17 01:05:46 -04:00
LC
3272aa3f08
[SVE256] Handle SSE insertions for ROUNDPD/ROUNDPS
...
See #3799
2026-06-17 00:42:20 -04:00
LC
32b11603d8
[SVE256] Handle SSE insertions for PMOVSX/PMOVZX ops
2026-06-17 00:04:03 -04:00
LC
538fd2672d
[SVE256] Handle SSE insertions for PSADBW
...
See #3799
2026-06-17 00:04:03 -04:00
LC
3e37724e3e
[SVE256] Handle SSE insertions for PHADDSW
...
See #379
2026-06-17 00:04:03 -04:00