Ryan Houdek
9d18ecc5cb
FEXCore: Fixes a crash with multiblock if ProcessorID IR op is encountered
...
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.
Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
2026-07-08 17:41:03 -07:00
Ryan Houdek
5f2455c502
Merge pull request #5665 from lioncash/blendop
...
[SVE256] Handle 256-bit blend operations much more efficiently
2026-07-08 16:43:20 -07:00
LC
6bc67609a3
[SVE256] Handle 256-bit blend operations much more efficiently
...
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC
8a8827c980
Merge pull request #5664 from simon902/CMPXCHGZeroing
...
OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as first operand
2026-07-08 15:14:12 -04:00
Simon Scherer
3d65c030a8
OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as destination operand and remove incorrect comment.
2026-07-08 14:58:32 +02:00
Simon Scherer
655102fc7d
JIT/ALUOps: Fix operand overlapping bug for pdep
2026-07-08 10:09:47 +02:00
Ryan Houdek
b90c9836cb
Merge pull request #5662 from lioncash/alias
...
OpcodeDispatcher: Remove asterisk from BMI source args
2026-07-07 11:22:59 -07:00
LC
718f2e01f7
OpcodeDispatcher: Remove asterisk from BMI source args
...
Keeps it consistent with the rest of the code and prevents breakages
whenever the Ref alias gets turned into its own value type.
2026-07-07 14:07:38 -04:00
Tony Wasserka
0ac6b3e8f3
CodeCache: Mark executable memory as EC code on ARM64EC
...
See bd5b817c3a .
2026-07-07 15:55:27 +02:00
Tony Wasserka
852e93aa74
Core: Fix dereference of invalid iterator
2026-07-06 15:07:37 +02:00
LC
7fd9b897c2
[SVE256] Handle 256-bit AES operations
...
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.
Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
LC
5145324806
[SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
...
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
LC
d11b19fd2b
[SVE256] Ensure insertion behavior for PCLMUL SSE operations
...
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC
684c568033
[SVE256] Ensure insertion behavior for SHA SSE operations
...
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC
ab4fb7b3ad
[SVE256] Ensure insertion behavior for AES operations on SSE
...
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
Ryan Houdek
6bcadde658
Merge pull request #5644 from lioncash/vmov
...
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
LC
c1e29f9013
VectorOps: Eliminate unnecessary moves in VMov if applicable
...
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC
4c27dfd5eb
VectorOps: Avoid temporary if able in 256-bit VFRecp
...
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC
201216ba54
HostFeatures: Put SVE support querying into single function
...
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
Ryan Houdek
16f90b33f3
Merge pull request #5641 from lioncash/minmax
...
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
LC
d3a85e14d9
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
...
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
LC
3d289f4489
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
...
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.
Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
LC
c6b0f360fe
AVX: Make use of table swapping constant to trim down relevant ops
...
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC
ef19242be3
IR: Add constant for swapping midsections of 256-bit vectors around
...
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
Paulo Matos
37b010795e
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 15:00:16 -07:00
LC
2cb4f8b6f5
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
...
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC
600e2ddecf
Vector: Only signify 128-bit vector loads in UCOMISxOp
...
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
LC
2ec2c39cf1
AVX: Lessen codegen for VMOVMSKPD
...
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
LC
41fc57f46c
AVX: Lessen codegen for 256-bit VMOVMSKPS
...
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
Ryan Houdek
9f2e982944
Merge pull request #5617 from simon902/vcvtps2ph
...
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC
3e60aa5738
Merge pull request #5610 from Sonicadvance1/171
...
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer
4564325bc7
OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph
2026-06-28 13:17:06 +02:00
LC
a5ecb71993
AVX: Handle two field insertions in VPERMQ
...
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC
6d86dca20b
AVX: Slightly trim codegen for VPSADW 256-bit case
...
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
Ryan Houdek
c85e426cc5
Merge pull request #5622 from lioncash/perm
...
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
Ryan Houdek
b23fa30099
Merge pull request #5618 from wsxarcher/fixsmcfullvector
...
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
wsxarcher
b21c49352e
Core: Start a new block for the next op in full SMC check
2026-06-28 18:07:28 +02:00
LC
e24f84504f
AVX: Handle trivial UZP/ZIP operations in VPERMQ
...
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC
ee2fb57f4e
AVX: Handle full broadcast in VDPPS
...
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC
11fe95d8ea
AVX: Simplify trivial case of VDPPS
...
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek
c0251dc8be
FEXCore: Pass host type that changes codegen to FEXCore
...
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
LC
d555ee8bcc
Vector: Fix typo in VPERMQOp
...
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Ryan Houdek
64392b2d45
OpcodeDispatcher: Fixes CRC32 with high 8-bit register
...
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC
8102a0974a
AVX: Handle easily broadcastable permutations in VPERMQ
...
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.
e.g.
0b00'00'00'01
0b00'01'01'01
0b11'11'00'11
are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek
d5be15c90e
Merge pull request #5605 from lioncash/permq
...
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC
e91efc6694
AVX: Skip identity insertions in VPERMQ
...
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
LC
a2889a09e3
Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
...
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC
cf098a0de6
AVX: Handle transpose cases in VPERMQ
...
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek
d93997c1cb
Merge pull request #5602 from lioncash/same
...
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek
e19aa975c8
Merge pull request #5597 from simon902/BTOpTypo
...
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00