LC
ee2fb57f4e
AVX: Handle full broadcast in VDPPS
...
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC
11fe95d8ea
AVX: Simplify trivial case of VDPPS
...
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
LC
d555ee8bcc
Vector: Fix typo in VPERMQOp
...
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
LC
8102a0974a
AVX: Handle easily broadcastable permutations in VPERMQ
...
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.
e.g.
0b00'00'00'01
0b00'01'01'01
0b11'11'00'11
are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek
d5be15c90e
Merge pull request #5605 from lioncash/permq
...
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC
e91efc6694
AVX: Skip identity insertions in VPERMQ
...
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
LC
a2889a09e3
Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
...
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC
cf098a0de6
AVX: Handle transpose cases in VPERMQ
...
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
LC
20647f2287
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
...
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
Ryan Houdek
ab4e0f653a
Merge pull request #5600 from lioncash/palign
...
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek
fda023e7dc
Merge pull request #5601 from lioncash/pshufb
...
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00
LC
a2e4209f4f
AVX: Reduce moves in 256-bit VPSHUFB
...
Just a minor reduction by avoiding insertion overhead.
2026-06-22 10:47:49 -04:00
LC
73d6716828
AVX: Shave some moves off 256-bit VPALIGNR
...
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
2026-06-22 10:09:13 -04:00
LC
1fb2419be2
Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
...
No behavior change, just a reduction in noise.
2026-06-22 09:47:46 -04:00
LC
3394808c06
AVX: Wire up helper to VPERMILPD
...
We can just leverage the shuffle handler for this, since VPSHUFD
essentially functions like VPERMILPD
2026-06-22 09:03:22 -04:00
LC
c252b58a15
AVX: Wire up lane helper for VPERMILPS imm variant
...
Makes for some more trivial savings. Will need handling for VPERMILPD
added separately, since selector behavior is different.
2026-06-22 05:21:10 -04:00
LC
39ae8c3ea0
AVX: Wire up lane helper for VSHUFPD/VSHUFPS
...
Also allows collapsing quite a bit of emitted code, like with
the shuffles in #5594
2026-06-22 04:46:22 -04:00
LC
eb8c2d964c
Vector: Factor out 128-bit path in SHUFOpImpl
...
We can leverage this for the 256-bit path
2026-06-22 03:47:20 -04:00
LC
bd9cf9ca11
AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
...
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.
Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.
For example:
vpshufd ymm0, ymm1, 0b00000011
drops from 50 instructions to 9
2026-06-22 00:16:39 -04:00
LC
9ab0920f01
OpcodeDispatcher: Fix typo in comment
...
It's the bits in general, not just the even ones (whoops).
2026-06-21 19:16:38 -04:00
LC
a6e7fba433
AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
...
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.
In some cases, this can be quite drastic, like with:
vpblendw ymm0, ymm1, ymm2, 0b00000001
being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC
8905e39439
OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
...
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC
df41b85827
OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
...
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC
9fa3b9345e
AVX: Remove unnecessary moves from PINSRX ops
...
These are old paths still around from when StoreResult used to
automatically perform truncating moves.
These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
Ryan Houdek
1619374252
Merge pull request #5587 from lioncash/pd
...
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
LC
42af6c8508
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
...
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC
b7df1bc259
[SVE256] Remove unnecessary move in VCVTPS2PD
...
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC
b40f9db735
[SVE256] Remove heavy handed moves from scalar compares
...
(See #3799 )
I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)
This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC
8989ce1766
[SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
...
See #3799
2026-06-19 18:53:50 -04:00
Ryan Houdek
9d0c05d9cc
Merge pull request #5575 from lioncash/movq2dq
...
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek
f6a68cb7fd
Merge pull request #5574 from lioncash/sse4a
...
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
LC
faa121e9ef
Vector: Move MOVQ2DQ over to Bind
...
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC
f48759e83a
[SVE256] Handle SSE insertions for MOVQ2DQ
...
See #3799
2026-06-19 17:42:21 -04:00
LC
2eca733603
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
...
See #3799
2026-06-19 17:05:59 -04:00
LC
14b65cec43
[SVE256] Handle SSE insertions for CVTPI2PD
...
See #3799
CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
LC
edd044752d
[SVE256] Handle SSE insertions for PMADDWD
...
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek
3a23bb4b73
Merge pull request #5569 from lioncash/cmp
...
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00
LC
99baa4f3d9
[SVE256] Handle SSE insertions for CMPSD/CMPSS
...
See #3799
2026-06-17 21:05:31 -04:00
LC
f64d4c571b
[SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
...
See #3799
2026-06-17 20:38:58 -04:00
LC
9372fa169a
[SVE256] Handle SSE insertions for MOVSD/MOVSS
...
See #3799
2026-06-17 20:09:17 -04:00
LC
daaa6ec129
[SVE256] Handle SSE insertions for aligned and unaligned moves
...
See #3799
2026-06-17 19:39:06 -04:00
LC
88afc22d5b
[SVE256] Handle SSE insertions for MOVNTDQA
...
See #3799
2026-06-17 19:38:57 -04:00
LC
b7ea9e30df
[SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
...
See #3799
Gets a few of the moves out of the way.
2026-06-17 17:39:13 -04:00
LC
844c3bb197
[SVE256] Handle SSE insertions for XOR special case
...
See #3799
Ensures that our special case maintains insertion behavior
2026-06-17 16:33:05 -04:00
LC
208c6d3eac
[SVE256] Handle SSE insertions for vector unary ops
...
See #3799
2026-06-17 14:43:56 -04:00
LC
886a2e74ac
[SVE256] Handle SSE insertions for pack ops
2026-06-17 13:36:18 -04:00
LC
db1d90ec9d
[SVE256] Handle SSE insertions for shuffles
...
See #3799
2026-06-17 08:44:45 -04:00
LC
124ce8420a
[SVE256] Handle SSE insertions for PINSR(B,D,Q,W)
...
See #3799
2026-06-17 08:22:26 -04:00
LC
4433eaf242
[SVE256] Handle SSE insertions for INSERTPS
...
See #3799
2026-06-17 08:11:49 -04:00
LC
5ba070f600
[SVE256] Handle SSE insertions for PSIGN(B,D,W)
...
See #3799
2026-06-17 07:59:33 -04:00