Lioncache
d99bcbf01b
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
...
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Lioncache
3858e4124b
OpcodeDispatcher: Remove redundant moves in AVX blend special cases
...
Zero-extension will happen if necessary upon storing.
2023-08-23 00:08:33 -04:00
Mai
819fe110da
Merge pull request #2967 from Sonicadvance1/optimize_storeelement
...
OpcodeDispatcher: Optimize MOVHP{S,D}
2023-08-23 00:01:25 -04:00
Ryan Houdek
adfd6787c0
Merge pull request #2969 from lioncash/insert
...
OpcodeDispatcher: Remove unnecessary moves from AVX inserts
2023-08-22 20:51:55 -07:00
Ryan Houdek
0ee2579a5e
OpcodeDispatcher: Optimize MOVHP{S,D}
...
Loads can turn in to element Loads.
Stores can turn in to element stores.
These four instruction variants are now optimal.
2023-08-22 20:42:24 -07:00
Mai
bb2f7107cd
Merge pull request #2963 from Sonicadvance1/optimize_loadelement
...
OpcodeDispatcher: Optimize MOVLP{S,D} loads
2023-08-22 23:41:59 -04:00
Lioncache
c33f3ff8df
OpcodeDispatcher: Remove unnecessary moves from AVX inserts
...
We already zero-extend on stores when necessary.
2023-08-22 23:29:40 -04:00
Ryan Houdek
de239cde67
OpcodeDispatcher: Optimize MOVLP{S,D} loads
...
This now uses the new load element IR operation and makes these
instructions optimal.
LRPCPC3 will introduce instructions in the future for TSO emulation to
help these operations, but that doesn't exist today.
2023-08-22 20:15:16 -07:00
Lioncache
2f5fae7677
OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
...
Zero-extension will occur automatically when necessary upon storing.
2023-08-22 23:06:53 -04:00
Lioncache
8f8062eb4e
OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
...
These zero-extensions will occur automatically when applicable.
2023-08-22 22:50:10 -04:00
Lioncache
e5f5629ffc
OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
...
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Ryan Houdek
ead141fd90
Merge pull request #2962 from lioncash/variable
...
OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
2023-08-22 18:52:47 -07:00
Lioncache
2b071e282e
OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
...
We already zero-extend on a store if necessary.
2023-08-22 21:20:34 -04:00
Lioncache
fa17d9fae9
OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
...
We already do a zero-extend if necessary in StoreResult.
This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Ryan Houdek
c795d42d21
OpcodeDispatcher: Optimize phminposuw
...
I would now consider the XMM version of this to be optimal.
Thanks to @rygorous for giving the idea for how to optimize this!
2023-08-22 16:29:06 -07:00
Mai
6c7933e7b1
Merge pull request #2957 from Sonicadvance1/optimize_pfnacc
...
OpcodeDispatcher: Optimize PFNACC
2023-08-22 10:10:11 -04:00
Ryan Houdek
364f084604
Merge pull request #2956 from Sonicadvance1/optimize_hsubp
...
OpcodeDispatcher: Optimize hsubp
2023-08-21 20:47:50 -07:00
Ryan Houdek
ffa8f1e3dc
Merge pull request #2955 from lioncash/sign
...
OpcodeDispatcher: Remove redundant move from VPSIGN
2023-08-21 20:47:41 -07:00
Ryan Houdek
ad6738939b
OpcodeDispatcher: Optimize PFNACC
...
Turns out this can be even more optimal.
2023-08-21 20:38:18 -07:00
Ryan Houdek
dcb3e4ee86
OpcodeDispatcher: Optimize hsubp
...
This makes the SSE version optimal.
This dramatically improves the AVX version as well.
2023-08-21 20:22:14 -07:00
Lioncache
dbbe6288de
OpcodeDispatcher: Remove redundant move from VPSIGN
...
StoreResult will already zero-extend if the vector is 128-bit.
2023-08-21 23:11:07 -04:00
Ryan Houdek
1563398d2c
OpcodeDispatcher: Optimize pmuludq
...
MMX version was already optimal, SSE version is now also.
AVX version is significantly improved.
2023-08-21 20:07:35 -07:00
Lioncache
920a0fb132
OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
...
Our store will already zero-extend if the vector is 128-bit.
2023-08-21 22:39:07 -04:00
Lioncache
ce8169794f
OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
...
We already zero-extend on a store if we have 256-bit vectors and the stored
vector is 128-bit.
2023-08-21 22:13:29 -04:00
Mai
185e3bfcb6
Merge pull request #2950 from Sonicadvance1/optimize_pmaddwd
...
OpcodeDispatcher: Optimize pmaddwd
2023-08-21 21:39:30 -04:00
Mai
3c49b3238a
Merge pull request #2949 from Sonicadvance1/optimize_phsub
...
OpcodeDispatcher: Optimize phsub
2023-08-21 21:39:02 -04:00
Ryan Houdek
2d7a3a578e
Merge pull request #2931 from Sonicadvance1/optimize_psign
...
Optimize PSIGN and VBSL
2023-08-21 18:30:55 -07:00
Ryan Houdek
3a2a576c35
Merge pull request #2951 from lioncash/shift
...
OpcodeDispatcher: Handle zero immediate shifts better
2023-08-21 18:20:18 -07:00
Ryan Houdek
869136b907
OpcodeDispatcher: Optimize pmaddwd
...
This is actually fairly trivial looking at it.
2023-08-21 17:53:32 -07:00
Lioncache
af8b6766d8
OpcodeDispatcher: Handle zero immediate shifts better
...
In the SSE and lower cases, we don't need to do anything,
since the value is already in the destination.
2023-08-21 20:46:41 -04:00
Ryan Houdek
a29076244f
OpcodeDispatcher: Optimizes mpsadbw
...
Two optimizations here:
1) The final VInsElement was generating three instructions
- This itself could have been change to vzip, which would have
removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
instruction off the calculation.
- addp odd elements and even elemnts first
- Then transpose those elements
- Then use one final addp to generate the result in the correct
order.
The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
2023-08-21 16:56:41 -07:00
Ryan Houdek
445c43792b
OpcodeDispatcher: Optimize phsub
...
This was...surprisingly bad. I blame myself entirely.
2023-08-21 16:54:35 -07:00
Ryan Houdek
b35f4df798
OpcodeDispatcher: Optimize SSE pmaddubsw
...
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
2023-08-21 16:21:01 -07:00
Ryan Houdek
a3bf952f2b
IR: Implements support for wide scalar shifts
...
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.
With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).
This is a significant improvement even on platforms that only support
128-bit SVE.
2023-08-20 19:16:40 -07:00
Ryan Houdek
c5f5a03c68
OpcodeDispatcher: Optimize PSIGN
...
This dramatically improves the performance of the PSIGN instructions.
2023-08-20 13:38:23 -07:00
Lioncache
c3778a9729
OpcodeDispatcher: Improve SHA1MSG1 output
...
We can simplify these inserts down to a single EXT
2023-08-20 12:50:07 -04:00
Ryan Houdek
92c3014aaa
OpcodeDispatcher: Minor optimization around clearing flags
...
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.
Back to back bfi is actually less optimal due to dependency tracking.
With #2911 , this is a total win since this hits an edge case with
constant loading that #2911 fixes.
2023-08-19 20:14:37 -07:00
Ryan Houdek
b973c193be
Merge pull request #2939 from lioncash/round
...
OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
2023-08-19 19:29:01 -07:00
Ryan Houdek
4c409ea47d
Merge pull request #2938 from lioncash/fcmp
...
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
2023-08-19 19:27:24 -07:00
Ryan Houdek
ac53913c37
Merge pull request #2937 from lioncash/scalarunary
...
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
2023-08-19 19:22:10 -07:00
Ryan Houdek
2224c23c79
Merge pull request #2934 from lioncash/scalarfp
...
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
2023-08-19 19:20:20 -07:00
Ryan Houdek
6ce380d7a9
Merge pull request #2936 from lioncash/scalaralu
...
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorScalarALUOp
2023-08-19 19:09:06 -07:00
Lioncache
5152854b98
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
...
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
2023-08-19 22:07:25 -04:00
Ryan Houdek
8dade7eea1
Merge pull request #2935 from lioncash/scalarfp2
...
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
2023-08-19 19:05:52 -07:00
Lioncache
6ba42e5cf1
OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
...
When dealing with scalar source registers, we can opt to not zero-extend
the vector and just perform the scalar operation and then insert the result.
2023-08-19 21:27:59 -04:00
Lioncache
343b00818d
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
...
We dealing with scalar vector sources, we don't need to zero-extend
the vector, and we can just use it as is.
2023-08-19 21:20:03 -04:00
Lioncache
09addb217a
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
...
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
2023-08-19 20:09:19 -04:00
Lioncache
4d1f002dea
OpcodeDispatcher: Remove unnecessary moves in AVXVectorScalarALUOp
...
Same thing as the SSE variant, but for AVX.
2023-08-19 19:22:42 -04:00
Lioncache
1158ad7b2a
OpcodeDispatcher: Remove unnecessary moves in VectorScalarALUOp
...
We can explicitly specify the vector width when working with a
vector source, so that we don't do any unnecessary zero-extending
on the element.
2023-08-19 19:15:13 -04:00
Lioncache
6907fdca6b
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
...
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
2023-08-19 18:46:08 -04:00