Ryan Houdek
1cb2b084b3
OpcodeDispatcher: Generate more optimal code for scalar GPR converts
...
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
FPR and converted it in-place.
These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek
189b0da68f
JIT/Int: Add support for scalar conversion as well
2023-08-25 03:19:11 -07:00
Ryan Houdek
62156f2152
ARM64JIT: Adds support for scalar cvt
2023-08-25 02:34:30 -07:00
Ryan Houdek
80d871fb18
Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
...
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek
3731e6d88b
OpcodeDispatcher: Optimize cvtps2pd
...
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek
c441b238c7
OpcodeDispatcher: Optimize MMX conversion operation
...
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek
72ce7ddf2d
Arm64: Optimize CVT operations for 64-bit variants
...
Using 128-bit converts for 64-bit versions cuts their throughput in half
on Cortex. Ensure we use the 64-bit version when possible.
2023-08-24 15:45:07 -07:00
Ryan Houdek
a1210f892a
OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
...
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.
Needs #2994 merged first.
Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek
ba01eac467
IR: Adds support for ARM's FCMA FCADD instruction
2023-08-24 15:00:41 -07:00
Ryan Houdek
c5d147322f
HostFeatures: Adds support for FCMA
2023-08-24 15:00:41 -07:00
Ryan Houdek
565b30e15e
OpcodeDispatcher: Cache named vector constants in the block
...
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Ryan Houdek
f300196d90
OpcodeDispatcher: Optimize AddSubP{S,D}
...
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.
Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Lioncache
42ccc18606
x86_64/MemoryOps: Fix mislabeled IR op messages
2023-08-23 22:54:36 -04:00
Mai
66c6f96120
Merge pull request #2990 from Sonicadvance1/optimize_pmulh
...
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek
77b6d854b9
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
...
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Ryan Houdek
05b9651279
IR: Implements new vector multiply returning high bits
...
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Lioncache
26c81224ac
OpcodeDispatcher: Remove redundant moves from AESIMC
...
Zero-extension will occur automatically upon storing if necessary.
We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache
8a622a3c1a
OpcodeDispatcher: Remove redundant moves from VAESEnc
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache
f4848fd1a7
OpcodeDispatcher: Remove redundant moves from VAESEncLast
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache
d37ce08ae9
OpcodeDispatcher: Remove redundant move from VAESDec
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache
a6f1a9f8e8
OpcodeDispatcher: Remove redundant moves from VAESDecLast
...
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache
52ab3f6a1e
OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
...
Zero-extension will occur upon storing if necessary.
We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache
410e99ba09
OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
...
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek
6e4765d48b
Merge pull request #2989 from lioncash/ins
...
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
2023-08-23 18:09:15 -07:00
Ryan Houdek
172c8f3ba6
Merge pull request #2988 from lioncash/half
...
Arm64/ConversionOps: Add missing half-precision conversions to scalar functions
2023-08-23 17:56:23 -07:00
Lioncache
203a2b1105
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
...
If Dst and DestVector alias one another, then we don't need to
move the vector unnecessarily.
2023-08-23 20:50:21 -04:00
Ryan Houdek
4297e13fcf
Merge pull request #2986 from lioncash/ext
...
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
2023-08-23 17:40:50 -07:00
Lioncache
5ad56ad52e
Arm64/ConversionOps: Add missing half-precision operations to Float_FromGPR_S
...
Provides parity with vector operations.
2023-08-23 20:36:34 -04:00
Lioncache
24e7baf28f
Arm64/ConversionOps: Add missing half-precision conversions to Float_FToF
...
Provides parity with the vector conversion operations.
2023-08-23 20:31:49 -04:00
Lioncache
b248ae4c04
Arm64/VectorOps: Remove redundant moves from SVE SQSHL
...
We don't need to emit a move if the destination and source alias.
2023-08-23 20:12:09 -04:00
Lioncache
95bea864cf
Arm64/VectorOps: Remove redundant moves from SVE SRSHR
...
We don't need to perform a move is the destination aliases
the source vector to be shifted.
2023-08-23 20:10:51 -04:00
Lioncache
47c4507bb6
Arm64/VectorOps: Remove redundant moves from SVE VSQXTUN2
...
We don't need to perform a move if the destination aliases the lower vector.
2023-08-23 20:10:29 -04:00
Lioncache
5ea0b6db28
Arm64/VectorOps: Remove redundant moves from SVE VSQXTN2
...
We don't need to perform a move if the destination aliases the
lower vector.
2023-08-23 20:02:28 -04:00
Mai
ee10153d14
Merge pull request #2984 from Sonicadvance1/optimize_pack
...
OpcodeDispatcher: Use new IR ops for pack instructions
2023-08-23 20:02:16 -04:00
Lioncache
d0d94adabe
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
...
We don't need to do any moves here is the destination aliases the
lower bits.
2023-08-23 19:56:10 -04:00
Ryan Houdek
926b8c2c97
Merge pull request #2985 from lioncash/shift
...
Arm64/VectorOps: Remove redundant moves from SVE variable/immediate/vector shifts when possible
2023-08-23 16:41:39 -07:00
Lioncache
18ebcdc9de
Arm64/VectorOps: Remove redundant moves in VUshrNI2
...
If the destination and VectorLower alias, then we don't need
to emit a movprfx.
2023-08-23 18:52:32 -04:00
Lioncache
f31a9a52e6
Arm64/VectorOps: Remove redundant moves from SVE immediate vector shifts when possible
...
If the destination and source vector alias one another, then the
operation can largely be done in place.
2023-08-23 18:36:21 -04:00
Lioncache
03504a5f8c
Arm64/VectorOps: Remove redundant moves from SVE vector shifts when possible
...
If the destination and the vector to be shifted alias, then we can
avoid needing to move some data around.
2023-08-23 18:24:58 -04:00
Lioncache
d29b4de1ee
Arm64/VectorOps: Remove redundant moves from SVE variable vector register shifts when possible
...
In the event that the destination and the vector to be shifted
alias one another, then we can skip the movprfx, since it's not
necessary.
2023-08-23 18:24:53 -04:00
Ryan Houdek
fc4559d3c4
OpcodeDispatcher: Use new IR ops for pack instructions
...
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Ryan Houdek
c508570da0
IR: Implements VSQXT{U,}NPair operations
...
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Lioncache
d5e145c4b0
Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
...
If the destination and true vector alias one another, then we can
perform the operation in place instead of moving data around.
2023-08-23 17:54:10 -04:00
Ryan Houdek
350bca97c6
Merge pull request #2982 from lioncash/imin
...
Arm64/VectorOps: Remove redundant moves from SVE V{S,U}Min/V{S,U}Max when possible
2023-08-23 14:53:15 -07:00
Ryan Houdek
226405880f
Merge pull request #2981 from lioncash/fmin
...
Arm64/VectorOps: Remove redundant moves from SVE VFMin/VFMax when possible
2023-08-23 14:46:03 -07:00
Lioncache
37a8cb6821
Arm64/VectorOps: Remove redundant moves from SVE VSMax when possible
...
When the destination and first source alias one another, then we
can perform the operation in place instead of moving data around.
2023-08-23 17:34:52 -04:00
Lioncache
fe2c7dbf97
Arm64/VectorOps: Remove redundant moves from SVE VUMax when possible
...
When the destination and source alias one another, then we
can perform the operation in place without needing to move
data around.
2023-08-23 17:32:17 -04:00
Lioncache
787b4f37fb
Arm64/VectorOps: Remove redundant moves from SVE VSMin when possible
...
When the destination and first source alias one another, then we can
perform the operation in place without moving any data.
2023-08-23 17:30:08 -04:00
Lioncache
c3faa019f5
Arm64/VectorOps: Remove redundant moves from SVE VUMin when possible
...
If the destination and first source alias, then we can perfom the operation
in place.
2023-08-23 17:27:52 -04:00
Ryan Houdek
da098d8204
Merge pull request #2979 from lioncash/div
...
Arm64/VectorOps: Remove moves from SVE VFDiv if possible
2023-08-23 14:22:46 -07:00