Commit Graph
2673 Commits
Author SHA1 Message Date
Ryan Houdek 9d18ecc5cb FEXCore: Fixes a crash with multiblock if ProcessorID IR op is encountered
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.

Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
2026-07-08 17:41:03 -07:00
Ryan Houdek 5f2455c502 Merge pull request #5665 from lioncash/blendop
[SVE256] Handle 256-bit blend operations much more efficiently
2026-07-08 16:43:20 -07:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 8a8827c980 Merge pull request #5664 from simon902/CMPXCHGZeroing
OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as first operand
2026-07-08 15:14:12 -04:00
Simon Scherer 3d65c030a8 OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as destination operand and remove incorrect comment. 2026-07-08 14:58:32 +02:00
Simon Scherer 655102fc7d JIT/ALUOps: Fix operand overlapping bug for pdep 2026-07-08 10:09:47 +02:00
Ryan Houdek b90c9836cb Merge pull request #5662 from lioncash/alias
OpcodeDispatcher: Remove asterisk from BMI source args
2026-07-07 11:22:59 -07:00
LC 718f2e01f7 OpcodeDispatcher: Remove asterisk from BMI source args
Keeps it consistent with the rest of the code and prevents breakages
whenever the Ref alias gets turned into its own value type.
2026-07-07 14:07:38 -04:00
Tony Wasserka 0ac6b3e8f3 CodeCache: Mark executable memory as EC code on ARM64EC
See bd5b817c3a.
2026-07-07 15:55:27 +02:00
Tony Wasserka 852e93aa74 Core: Fix dereference of invalid iterator 2026-07-06 15:07:37 +02:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
LC 5145324806 [SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
LC d3a85e14d9 VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek d93997c1cb Merge pull request #5602 from lioncash/same
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek e19aa975c8 Merge pull request #5597 from simon902/BTOpTypo
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00