Commit Graph
54 Commits
Author SHA1 Message Date
Lioncache 2f5fae7677 OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
Zero-extension will occur automatically when necessary upon storing.
2023-08-22 23:06:53 -04:00
Lioncache 8f8062eb4e OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
These zero-extensions will occur automatically when applicable.
2023-08-22 22:50:10 -04:00
Lioncache e5f5629ffc OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Ryan Houdek ead141fd90 Merge pull request #2962 from lioncash/variable
OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
2023-08-22 18:52:47 -07:00
Lioncache 2b071e282e OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
We already zero-extend on a store if necessary.
2023-08-22 21:20:34 -04:00
Lioncache fa17d9fae9 OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
We already do a zero-extend if necessary in StoreResult.

This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Ryan Houdek c795d42d21 OpcodeDispatcher: Optimize phminposuw
I would now consider the XMM version of this to be optimal.

Thanks to @rygorous for giving the idea for how to optimize this!
2023-08-22 16:29:06 -07:00
Mai 6c7933e7b1 Merge pull request #2957 from Sonicadvance1/optimize_pfnacc
OpcodeDispatcher: Optimize PFNACC
2023-08-22 10:10:11 -04:00
Ryan Houdek 364f084604 Merge pull request #2956 from Sonicadvance1/optimize_hsubp
OpcodeDispatcher: Optimize hsubp
2023-08-21 20:47:50 -07:00
Ryan Houdek ffa8f1e3dc Merge pull request #2955 from lioncash/sign
OpcodeDispatcher: Remove redundant move from VPSIGN
2023-08-21 20:47:41 -07:00
Ryan Houdek ad6738939b OpcodeDispatcher: Optimize PFNACC
Turns out this can be even more optimal.
2023-08-21 20:38:18 -07:00
Ryan Houdek dcb3e4ee86 OpcodeDispatcher: Optimize hsubp
This makes the SSE version optimal.
This dramatically improves the AVX version as well.
2023-08-21 20:22:14 -07:00
Lioncache dbbe6288de OpcodeDispatcher: Remove redundant move from VPSIGN
StoreResult will already zero-extend if the vector is 128-bit.
2023-08-21 23:11:07 -04:00
Ryan Houdek 1563398d2c OpcodeDispatcher: Optimize pmuludq
MMX version was already optimal, SSE version is now also.
AVX version is significantly improved.
2023-08-21 20:07:35 -07:00
Lioncache 920a0fb132 OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
Our store will already zero-extend if the vector is 128-bit.
2023-08-21 22:39:07 -04:00
Lioncache ce8169794f OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
We already zero-extend on a store if we have 256-bit vectors and the stored
vector is 128-bit.
2023-08-21 22:13:29 -04:00
Mai 185e3bfcb6 Merge pull request #2950 from Sonicadvance1/optimize_pmaddwd
OpcodeDispatcher: Optimize pmaddwd
2023-08-21 21:39:30 -04:00
Mai 3c49b3238a Merge pull request #2949 from Sonicadvance1/optimize_phsub
OpcodeDispatcher: Optimize phsub
2023-08-21 21:39:02 -04:00
Ryan Houdek 2d7a3a578e Merge pull request #2931 from Sonicadvance1/optimize_psign
Optimize PSIGN and VBSL
2023-08-21 18:30:55 -07:00
Ryan Houdek 3a2a576c35 Merge pull request #2951 from lioncash/shift
OpcodeDispatcher: Handle zero immediate shifts better
2023-08-21 18:20:18 -07:00
Ryan Houdek 869136b907 OpcodeDispatcher: Optimize pmaddwd
This is actually fairly trivial looking at it.
2023-08-21 17:53:32 -07:00
Lioncache af8b6766d8 OpcodeDispatcher: Handle zero immediate shifts better
In the SSE and lower cases, we don't need to do anything,
since the value is already in the destination.
2023-08-21 20:46:41 -04:00
Ryan Houdek a29076244f OpcodeDispatcher: Optimizes mpsadbw
Two optimizations here:
1) The final VInsElement was generating three instructions
   - This itself could have been change to vzip, which would have
     removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
   instruction off the calculation.
   - addp odd elements and even elemnts first
   - Then transpose those elements
   - Then use one final addp to generate the result in the correct
     order.

The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
2023-08-21 16:56:41 -07:00
Ryan Houdek 445c43792b OpcodeDispatcher: Optimize phsub
This was...surprisingly bad. I blame myself entirely.
2023-08-21 16:54:35 -07:00
Ryan Houdek b35f4df798 OpcodeDispatcher: Optimize SSE pmaddubsw
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
2023-08-21 16:21:01 -07:00
Ryan Houdek a3bf952f2b IR: Implements support for wide scalar shifts
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.

With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).

This is a significant improvement even on platforms that only support
128-bit SVE.
2023-08-20 19:16:40 -07:00
Ryan Houdek c5f5a03c68 OpcodeDispatcher: Optimize PSIGN
This dramatically improves the performance of the PSIGN instructions.
2023-08-20 13:38:23 -07:00
Lioncache c3778a9729 OpcodeDispatcher: Improve SHA1MSG1 output
We can simplify these inserts down to a single EXT
2023-08-20 12:50:07 -04:00
Ryan Houdek 92c3014aaa OpcodeDispatcher: Minor optimization around clearing flags
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.

Back to back bfi is actually less optimal due to dependency tracking.

With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
2023-08-19 20:14:37 -07:00
Ryan Houdek b973c193be Merge pull request #2939 from lioncash/round
OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
2023-08-19 19:29:01 -07:00
Ryan Houdek 4c409ea47d Merge pull request #2938 from lioncash/fcmp
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
2023-08-19 19:27:24 -07:00
Ryan Houdek ac53913c37 Merge pull request #2937 from lioncash/scalarunary
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
2023-08-19 19:22:10 -07:00
Ryan Houdek 2224c23c79 Merge pull request #2934 from lioncash/scalarfp
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
2023-08-19 19:20:20 -07:00
Ryan Houdek 6ce380d7a9 Merge pull request #2936 from lioncash/scalaralu
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorScalarALUOp
2023-08-19 19:09:06 -07:00
Lioncache 5152854b98 OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
2023-08-19 22:07:25 -04:00
Ryan Houdek 8dade7eea1 Merge pull request #2935 from lioncash/scalarfp2
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
2023-08-19 19:05:52 -07:00
Lioncache 6ba42e5cf1 OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
When dealing with scalar source registers, we can opt to not zero-extend
the vector and just perform the scalar operation and then insert the result.
2023-08-19 21:27:59 -04:00
Lioncache 343b00818d OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
We dealing with scalar vector sources, we don't need to zero-extend
the vector, and we can just use it as is.
2023-08-19 21:20:03 -04:00
Lioncache 09addb217a OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
2023-08-19 20:09:19 -04:00
Lioncache 4d1f002dea OpcodeDispatcher: Remove unnecessary moves in AVXVectorScalarALUOp
Same thing as the SSE variant, but for AVX.
2023-08-19 19:22:42 -04:00
Lioncache 1158ad7b2a OpcodeDispatcher: Remove unnecessary moves in VectorScalarALUOp
We can explicitly specify the vector width when working with a
vector source, so that we don't do any unnecessary zero-extending
on the element.
2023-08-19 19:15:13 -04:00
Lioncache 6907fdca6b OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
2023-08-19 18:46:08 -04:00
Lioncache c31329609f OpcodeDispatcher: Unify handling code for MOVSD and MOVSS
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
2023-08-19 17:44:45 -04:00
Lioncache 1fe8470933 OpcodeDispatcher: Remove extraneous moves from VMOVSS/VMOVSD xmm to mem case
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
2023-08-19 17:35:52 -04:00
Lioncache 99b5aaa426 OpcodeDispatcher: Remove extraneous moves in VMOVSS/VMOVSD register case
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
2023-08-19 17:35:12 -04:00
Lioncache bbed4d73ed OpcodeDispatcher: Improve VPERMQ/VPERMPD broadcast cases
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
2023-08-19 01:08:30 -04:00
Ryan Houdek f09d9af3db Merge pull request #2922 from lioncash/psrld
OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
2023-08-17 17:05:35 -07:00
Lioncache 9e54ec2724 OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
2023-08-17 19:33:09 -04:00
Ryan Houdek 461ca6fe7c Merge pull request #2921 from lioncash/shift
OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
2023-08-17 16:03:16 -07:00
Lioncache 5a1f32c339 OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
PSLLIImpl already checks for and handles a shift value of zero.
2023-08-17 18:47:50 -04:00