Commit Graph
2634 Commits
Author SHA1 Message Date
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek d93997c1cb Merge pull request #5602 from lioncash/same
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek e19aa975c8 Merge pull request #5597 from simon902/BTOpTypo
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00
Simon Scherer 3f02dd0a36 OpcodeDispatcher: Update BTOp comment to clarify AMD vs Intel flag behavior 2026-06-25 07:36:32 +02:00
Ryan Houdek ab4e0f653a Merge pull request #5600 from lioncash/palign
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek fda023e7dc Merge pull request #5601 from lioncash/pshufb
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00
Simon Scherer e03187852b OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved. 2026-06-24 16:11:13 +02:00
LC 20647f2287 AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
LC a2e4209f4f AVX: Reduce moves in 256-bit VPSHUFB
Just a minor reduction by avoiding insertion overhead.
2026-06-22 10:47:49 -04:00
LC 73d6716828 AVX: Shave some moves off 256-bit VPALIGNR
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
2026-06-22 10:09:13 -04:00
LC 1fb2419be2 Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
No behavior change, just a reduction in noise.
2026-06-22 09:47:46 -04:00
LC 3394808c06 AVX: Wire up helper to VPERMILPD
We can just leverage the shuffle handler for this, since VPSHUFD
essentially functions like VPERMILPD
2026-06-22 09:03:22 -04:00
LC c252b58a15 AVX: Wire up lane helper for VPERMILPS imm variant
Makes for some more trivial savings. Will need handling for VPERMILPD
added separately, since selector behavior is different.
2026-06-22 05:21:10 -04:00
LC 39ae8c3ea0 AVX: Wire up lane helper for VSHUFPD/VSHUFPS
Also allows collapsing quite a bit of emitted code, like with
the shuffles in #5594
2026-06-22 04:46:22 -04:00
LC eb8c2d964c Vector: Factor out 128-bit path in SHUFOpImpl
We can leverage this for the 256-bit path
2026-06-22 03:47:20 -04:00
LC bd9cf9ca11 AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.

Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.

For example:

vpshufd ymm0, ymm1, 0b00000011

drops from 50 instructions to 9
2026-06-22 00:16:39 -04:00
Ryan Houdek 27a5f09185 Merge pull request #5592 from ShadowCurse/fixes
JIT: Arm64: fix the loop in CacheLineClear/Clean
2026-06-23 17:15:07 -07:00
Ryan Houdek 01b0b4e653 Merge pull request #5593 from lioncash/dup
VectorOps: Avoid dup if able in VInsElement 128-bit element path
2026-06-23 17:10:35 -07:00
Egor Lazarchuk e3e9777ee6 JIT: Arm64: fix the loop in CacheLineClear/Clean
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
2026-06-24 00:38:48 +01:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC 9ab0920f01 OpcodeDispatcher: Fix typo in comment
It's the bits in general, not just the even ones (whoops).
2026-06-21 19:16:38 -04:00
LC a6e7fba433 AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.

In some cases, this can be quite drastic, like with:

vpblendw ymm0, ymm1, ymm2, 0b00000001

being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC 8905e39439 OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC df41b85827 OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC 9fa3b9345e AVX: Remove unnecessary moves from PINSRX ops
These are old paths still around from when StoreResult used to
automatically perform truncating moves.

These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
Ryan Houdek 1619374252 Merge pull request #5587 from lioncash/pd
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
LC 42af6c8508 [SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC b7df1bc259 [SVE256] Remove unnecessary move in VCVTPS2PD
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC b40f9db735 [SVE256] Remove heavy handed moves from scalar compares
(See #3799)

I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)

This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
Ryan Houdek 55c90cfc38 Merge pull request #5579 from lioncash/typo
JIT: Amend op typos in implementations
2026-06-21 20:04:19 -07:00
LC 1d3403fdc2 JIT: Amend op typos in implementations
Mostly benign, but ensures that they're correct in the event any of
their IR definitions change.
2026-06-20 13:16:12 -04:00
LC 53301b0f56 Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
Same thing, just a little less verbose.
2026-06-20 12:32:23 -04:00
LC 8989ce1766 [SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
See #3799
2026-06-19 18:53:50 -04:00
Ryan Houdek 9d0c05d9cc Merge pull request #5575 from lioncash/movq2dq
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek f6a68cb7fd Merge pull request #5574 from lioncash/sse4a
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
LC faa121e9ef Vector: Move MOVQ2DQ over to Bind
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC f48759e83a [SVE256] Handle SSE insertions for MOVQ2DQ
See #3799
2026-06-19 17:42:21 -04:00
LC 2eca733603 [SVE256] Handle SSE insertions for EXTRQ/INSERTQ
See #3799
2026-06-19 17:05:59 -04:00
LC 14b65cec43 [SVE256] Handle SSE insertions for CVTPI2PD
See #3799

CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
Simon Scherer 9fa3221687 FEXCore: Fix incorrect RSP update for 16bit leave 2026-06-19 14:23:16 +02:00
LC edd044752d [SVE256] Handle SSE insertions for PMADDWD
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek 3a23bb4b73 Merge pull request #5569 from lioncash/cmp
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00