Commit Graph
486 Commits
Author SHA1 Message Date
Ryan Houdek 548fd9daf8 OpcodeDispatcher: Implement support for SSE4.1 NT load 2024-07-10 23:07:37 -07:00
Ryan Houdek f831f5a0e1 AVX128: Implement support for NT Load 2024-07-10 23:07:14 -07:00
Ryan Houdek 72d6c8ebd6 Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00
Mai af6a0be832 Merge pull request #3842 from Sonicadvance1/fix_f64_to_i32
VCVT{T,}PD2DQ fixes and optimization
2024-07-09 03:49:31 -04:00
Ryan Houdek b9c214e6e8 OpcodeDispatcher: Use new IR op for vcvt{t,}pd2dq
Also fixes a bug where it was failing to zero the upper bits of the
destination register in the AVX128 implementation. Which the updated
unit tests now check against.

Fixes a minor precision issue that was reported in #2995. We still don't
return correct values for overflow. x86 always returns maximum negative
int32_t on overflow, ARM will return maximum negative or positive
depending on sign of the double.
2024-07-09 00:38:47 -07:00
Ryan Houdek b3a7a973a1 AVX128: Extends 32-bit indexes path for 128-bit operations
The codepath from #3826 was only targeting 256-bit sized operations.
This missed the vpgatherdq/vgatherdpd 128-bit operations. By extending
the codepath to understand 128-bit operations, we now hit these
instruction variants.

With this PR, we now have SVE128 codepaths that handle ALL variants of
x86 gather instructions! There are zero ASIMD fallbacks used in this
case!

Of course depending on the instruction, the performance still leaves a
lot to be desired, and there is no way to emulate x86 TSO behaviour
without an ASIMD fallback, which we will likely need to add as a
fallback at some point.

Based on #3836 until that is merged.
2024-07-08 18:44:07 -07:00
Ryan Houdek 4afbfcae17 AVX128: Optimize the vpgatherdd/vgatherdps cases that would fall back to ASIMD
With the introduction of the wide gathers in #3828 this has opened new
avenues for optimizing these cases that would typically fall back to
ASIMD. In the cases that 32-bit SVE scaling doesn't fit, we can instead
sign extend the elements in to double-width address registers.

This then feeds naturally in to the SVE path even though we end up
needing to allocate 512-bits worth of address registers. This ends up
being significantly better than the ASIMD path still.

Relies on #3828 to be merged first
Fixes #3829
2024-07-08 18:12:28 -07:00
Ryan Houdek ec7c8fd922 AVX128: Optimize QPS/QD variant of gather loads!
SVE has a special version of their gather instruction that gets similar
behaviour to x86's VGATHERQPS/VPGATHERQD instructions.

The quirk of these instructions that the previous SVE implementation
didn't handle and required ASIMD fallback, was that most gather
instructions require the data element size and address element size to
match. This x86 instruction uses a 64-bit address size while loading 32-bit
elements. This matches this specific variant of the SVE instruction, but
the data is zero-extended once loaded, requiring us to shuffle the data
after it is loaded.

This isn't the worst but the implementation is different enough that
stuffing it in to the other gather load will cause headaches.

Basically gets 32 instruction variants to use the SVE version!

Fixes #3827
2024-07-08 17:19:18 -07:00
Ryan Houdek df40515087 AVX128: Extend 32-bit address indices when possible
When loading 256-bits of data with only 128-bits of address indices, we
can sign extend the source indices to be 64-bit. Thus falling down the
ideal path for SVE where each 128-bit lane is loading the data to
addresses in a 1:1 element ratio.

This means we use the SVE path more often because of this.

Based on top of #3825 because the prescaling behaviour was introduced
there. This implements its own prescaling when the sign extension occurs
because ARM's SSHLL{,2} instruction gives us that for free.

This additionally fixes a bug where we were accidentally loading the top
128-bit half of the addresses for gathers when it was unnecessary, and
on the AVX256 side it was duplicating and doing some additional work
when it shouldn't have.

It'll be good to walk the commits when looking at this one, as there are
a couple of incremental changes that are easier to follow that way.

Fixes #3806
2024-07-06 18:32:35 -07:00
Ryan Houdek 0f9abe68b9 AVX128: Fixes accidentally loading high addr register when unnnecessary
Was missing a clamp on the high half when encounting a 128-bit gather
instruction. Was causing us to unconditionally load the top half when it
was unncessary.
2024-07-06 18:32:35 -07:00
Ryan Houdek 0d4414fdd0 AVX128: Removes templated AddrElementSize and add as argument
NFC
2024-07-06 18:32:35 -07:00
Ryan Houdek 9bad09c45f Merge pull request #3823 from alyssarosenzweig/bug/shl-var-small
Fix CF with small shifts
2024-07-06 01:33:57 -07:00
Ryan Houdek 11a494d7b3 AVX128: Prescale addresses in gathers if possible
If the host supports SVE128, if the address element size and data size is 64-bit, and the scale is not one of the two that is supported by SVE; Then prescale the addresses.
64-bit address overflow masks the top bits so is well defined that we
can scale the vector elements and still execute the SVE code path in
that case. Removing the ASIMD code paths from a lot of gathers.

Fixes #3805
2024-07-05 16:47:11 -07:00
Alyssa Rosenzweig 0f0e402db4 OpcodeDispatcher: fix CF with 8/16-bit immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:24:34 -04:00
Alyssa Rosenzweig adc709db2f OpcodeDispatcher: drop remnants of deferred flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 17:22:41 -04:00
Alyssa Rosenzweig 395573720d OpcodeDispatcher: drop pointless flag defers for shifts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 0e62759d24 OpcodeDispatcher: stop deferring logical
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 926b6c3117 OpcodeDispatcher: don't defer mul flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig c9f9304ba5 OpcodeDispatcher: stop deferring obscure bitwise
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig fabd6be5af OpcodeDispatcher: drop SUB defer
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig a38205069b OpcodeDispatcher: fix SBB carry flag
do it the naive way, just applying the x86 definitions of SBB.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Ryan Houdek 0d06e3e47d Revert "OpcodeDispatcher: add cache"
This reverts commit 46676ca376.
2024-07-02 20:24:57 -07:00
Ryan Houdek 472a373861 Merge pull request #3786 from Sonicadvance1/non_temporal_stores
OpcodeDispatcher: Implement support for non-temporal vector stores
2024-07-01 18:57:38 -07:00
Ryan Houdek a451420911 Merge pull request #3783 from Sonicadvance1/optimize_vector_zeroregister
OpcodeDispatcher: Optimize x86 canonical vector zero register
2024-07-01 18:57:31 -07:00
Ryan Houdek babde31bf0 AVX128: Fixes vmovlhps
We didn't have a unit test for this and we weren't implementing it at
all.
We treated it as vmovhps/vmovhpd accidentally. Once again caught by the
libaom Intrinsics unit tests.
2024-07-01 13:54:11 -07:00
Ryan Houdek 5821054d91 Merge pull request #3789 from Sonicadvance1/avx128_minor_pshufb_opt
AVX128: Minor optimization to 256-bit vpshufb
2024-06-30 15:45:11 -07:00
Ryan Houdek 4626145374 Merge pull request #3792 from Sonicadvance1/avx128_fix_scalar_fma
AVX128: Fixes scalar FMA accidentally using vector wide
2024-06-30 15:36:09 -07:00
Ryan Houdek 1393dc2a5b AVX128: Fixes scalar FMA accidentally using vector wide 2024-06-30 14:36:33 -07:00
Ryan Houdek cffae9cb0f AVX128: Minor optimization to 256-bit vpshufb 2024-06-30 13:41:03 -07:00
Ryan Houdek 7d05610da7 OpcodeDispatcher: Optimize x86 canonical vector zero register
The canonical way to generate a zero register vector in x86 is to xor
itself. Capture this can convert it to canonical zero register instead.

Can get zero-cycle renamed on latest CPUs.
2024-06-29 22:21:53 -07:00
Ryan Houdek f4ff1b0688 OpcodeDispatcher: Implement support for non-temporal vector stores
x86 doesn't have a lot of non-temporal vector stores but we do have a
few of them.

- MMX: MOVNTQ
- SSE2: MOVNTDQ, MOVNTPS, MOVNTPD
- AVX: VMOVNTDQ (128-bit & 256-bit), VMOVNTPD

Additionally SSE4a adds 32-bit and 64-bit scalar vector non-temporal
stores, which we keep as regular stores. Since ARM doesn't have matching
semantics for those.

Additionally SSE4.1 adds non-temporal vector LOADS which this doesn't
touch.
- SSE4.1: MOVNTDQA
- AVX: VMOVNTDQA (128-bit)
- AVX2: VMOVNTDQA (256-bit)

Fixes #3364
2024-06-29 22:05:56 -07:00
Ryan Houdek ebfa65fedc AVX128: Minor optimization to vmov{l,h}{ps,pd} 2024-06-29 19:27:16 -07:00
Ryan Houdek aba7a3a830 AVX128: Fixes vblendps lower and upper selector 2024-06-27 17:20:39 -07:00
Ryan Houdek 9027d1eee7 AVX128: Fixes bug in vector immediate shift 2024-06-27 16:22:14 -07:00
Alyssa Rosenzweig f9b53c6b51 AVX_128: save a move in vzeroall
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-27 10:30:25 -04:00
Ryan Houdek 4d56fec5f1 AVX128: Work around glibc fault testing 2024-06-26 16:49:00 -07:00
Ryan Houdek 975069825e AVX128: Fix a real bug with VCVTPS2PH 2024-06-26 16:49:00 -07:00
Mai a031a49546 Merge pull request #3767 from Sonicadvance1/avx128_fix_wide_shift
AVX128: Fixes wide shifts
2024-06-26 17:29:09 -04:00
Ryan Houdek f277025c9a AVX128: Fixes wide shifts
During refactoring this was missed and rerunning unittests locally
caught it. 256-bit operations get their shift only from the lower half
of the vector register.
2024-06-26 14:16:39 -07:00
Ryan Houdek 3a89df9bed AVX128: Implement support for F16C 2024-06-26 14:05:12 -07:00
Ryan Houdek f6a0866fbb IR: Split Vector_FToF2 in to VFCVTL2 and VCVTFN2
I forgot in the narrowing case we need to be careful about insert. No IR
op used Vector_FToF2 with narrowing.
2024-06-26 14:03:41 -07:00
Ryan Houdek 756fa2ecc5 Merge pull request #3766 from alyssarosenzweig/opt/f16c-round
Optimize vcvtps2ph
2024-06-26 14:03:24 -07:00
Alyssa Rosenzweig d2324f4a93 OpcodeDispatcher: optimize vcvtps2ph
We can avoid a LOT of pointless work with some dedicated IR ops for specifically
overriding the round mode.

Small behaviour change here: we no longer reset FTZ. I think this is a bug fix?
But if it's not it's not hard to fix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 16:46:21 -04:00
Ryan Houdek a4fa3a460e OpcodeDispatcher: Implement AVX gathers with SVE256
Just to ensure we still have feature parity.
2024-06-26 16:00:53 -04:00
Ryan Houdek 77ba708933 AVX128: Implement support for gather load instructions
This is the last family of instructions that we needed to implement for
AVX2 to be properly advertised!
2024-06-26 16:00:53 -04:00
Alyssa Rosenzweig d1d41f5645 Merge pull request #3763 from alyssarosenzweig/rclse/less-aggressive
Remove RCLSE
2024-06-26 15:14:14 -04:00
Ryan Houdek 94fd100fc7 Merge pull request #3719 from lioncash/f16c
OpcodeDispatcher: Handle F16C operations
2024-06-26 12:12:13 -07:00
Lioncache cd5a809ec9 OpcodeDispatcher: Handle VCVTPS2PH 2024-06-26 15:05:03 -04:00
Lioncache 045a8efbeb OpcodeDispatcher: Handle VCVTPH2PS
Fairly straightforward, since we already have handling for half-float conversions.
2024-06-26 15:05:00 -04:00
Alyssa Rosenzweig 46676ca376 OpcodeDispatcher: add cache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 14:49:05 -04:00