Commit Graph
166 Commits
Author SHA1 Message Date
Ryan Houdek 1cb2b084b3 OpcodeDispatcher: Generate more optimal code for scalar GPR converts
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
   FPR and converted it in-place.

These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek 189b0da68f JIT/Int: Add support for scalar conversion as well 2023-08-25 03:19:11 -07:00
Ryan Houdek 62156f2152 ARM64JIT: Adds support for scalar cvt 2023-08-25 02:34:30 -07:00
Ryan Houdek 80d871fb18 Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek 3731e6d88b OpcodeDispatcher: Optimize cvtps2pd
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek 72ce7ddf2d Arm64: Optimize CVT operations for 64-bit variants
Using 128-bit converts for 64-bit versions cuts their throughput in half
on Cortex. Ensure we use the 64-bit version when possible.
2023-08-24 15:45:07 -07:00
Ryan Houdek a1210f892a OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.

Needs #2994 merged first.

Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek ba01eac467 IR: Adds support for ARM's FCMA FCADD instruction 2023-08-24 15:00:41 -07:00
Ryan Houdek c5d147322f HostFeatures: Adds support for FCMA 2023-08-24 15:00:41 -07:00
Ryan Houdek 565b30e15e OpcodeDispatcher: Cache named vector constants in the block
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.

Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.

Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Ryan Houdek f300196d90 OpcodeDispatcher: Optimize AddSubP{S,D}
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.

Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Lioncache 42ccc18606 x86_64/MemoryOps: Fix mislabeled IR op messages 2023-08-23 22:54:36 -04:00
Mai 66c6f96120 Merge pull request #2990 from Sonicadvance1/optimize_pmulh
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek 77b6d854b9 OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Ryan Houdek 05b9651279 IR: Implements new vector multiply returning high bits
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 8a622a3c1a OpcodeDispatcher: Remove redundant moves from VAESEnc
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache f4848fd1a7 OpcodeDispatcher: Remove redundant moves from VAESEncLast
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache d37ce08ae9 OpcodeDispatcher: Remove redundant move from VAESDec
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache a6f1a9f8e8 OpcodeDispatcher: Remove redundant moves from VAESDecLast
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache 410e99ba09 OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek 6e4765d48b Merge pull request #2989 from lioncash/ins
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
2023-08-23 18:09:15 -07:00
Ryan Houdek 172c8f3ba6 Merge pull request #2988 from lioncash/half
Arm64/ConversionOps: Add missing half-precision conversions to scalar functions
2023-08-23 17:56:23 -07:00
Lioncache 203a2b1105 Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
If Dst and DestVector alias one another, then we don't need to
move the vector unnecessarily.
2023-08-23 20:50:21 -04:00
Ryan Houdek 4297e13fcf Merge pull request #2986 from lioncash/ext
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
2023-08-23 17:40:50 -07:00
Lioncache 5ad56ad52e Arm64/ConversionOps: Add missing half-precision operations to Float_FromGPR_S
Provides parity with vector operations.
2023-08-23 20:36:34 -04:00
Lioncache 24e7baf28f Arm64/ConversionOps: Add missing half-precision conversions to Float_FToF
Provides parity with the vector conversion operations.
2023-08-23 20:31:49 -04:00
Lioncache b248ae4c04 Arm64/VectorOps: Remove redundant moves from SVE SQSHL
We don't need to emit a move if the destination and source alias.
2023-08-23 20:12:09 -04:00
Lioncache 95bea864cf Arm64/VectorOps: Remove redundant moves from SVE SRSHR
We don't need to perform a move is the destination aliases
the source vector to be shifted.
2023-08-23 20:10:51 -04:00
Lioncache 47c4507bb6 Arm64/VectorOps: Remove redundant moves from SVE VSQXTUN2
We don't need to perform a move if the destination aliases the lower vector.
2023-08-23 20:10:29 -04:00
Lioncache 5ea0b6db28 Arm64/VectorOps: Remove redundant moves from SVE VSQXTN2
We don't need to perform a move if the destination aliases the
lower vector.
2023-08-23 20:02:28 -04:00
Mai ee10153d14 Merge pull request #2984 from Sonicadvance1/optimize_pack
OpcodeDispatcher: Use new IR ops for pack instructions
2023-08-23 20:02:16 -04:00
Lioncache d0d94adabe Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
We don't need to do any moves here is the destination aliases the
lower bits.
2023-08-23 19:56:10 -04:00
Ryan Houdek 926b8c2c97 Merge pull request #2985 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from SVE variable/immediate/vector shifts when possible
2023-08-23 16:41:39 -07:00
Lioncache 18ebcdc9de Arm64/VectorOps: Remove redundant moves in VUshrNI2
If the destination and VectorLower alias, then we don't need
to emit a movprfx.
2023-08-23 18:52:32 -04:00
Lioncache f31a9a52e6 Arm64/VectorOps: Remove redundant moves from SVE immediate vector shifts when possible
If the destination and source vector alias one another, then the
operation can largely be done in place.
2023-08-23 18:36:21 -04:00
Lioncache 03504a5f8c Arm64/VectorOps: Remove redundant moves from SVE vector shifts when possible
If the destination and the vector to be shifted alias, then we can
avoid needing to move some data around.
2023-08-23 18:24:58 -04:00
Lioncache d29b4de1ee Arm64/VectorOps: Remove redundant moves from SVE variable vector register shifts when possible
In the event that the destination and the vector to be shifted
alias one another, then we can skip the movprfx, since it's not
necessary.
2023-08-23 18:24:53 -04:00
Ryan Houdek fc4559d3c4 OpcodeDispatcher: Use new IR ops for pack instructions
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Ryan Houdek c508570da0 IR: Implements VSQXT{U,}NPair operations
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Lioncache d5e145c4b0 Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
If the destination and true vector alias one another, then we can
perform the operation in place instead of moving data around.
2023-08-23 17:54:10 -04:00
Ryan Houdek 350bca97c6 Merge pull request #2982 from lioncash/imin
Arm64/VectorOps: Remove redundant moves from SVE V{S,U}Min/V{S,U}Max when possible
2023-08-23 14:53:15 -07:00
Ryan Houdek 226405880f Merge pull request #2981 from lioncash/fmin
Arm64/VectorOps: Remove redundant moves from SVE VFMin/VFMax when possible
2023-08-23 14:46:03 -07:00
Lioncache 37a8cb6821 Arm64/VectorOps: Remove redundant moves from SVE VSMax when possible
When the destination and first source alias one another, then we
can perform the operation in place instead of moving data around.
2023-08-23 17:34:52 -04:00
Lioncache fe2c7dbf97 Arm64/VectorOps: Remove redundant moves from SVE VUMax when possible
When the destination and source alias one another, then we
can perform the operation in place without needing to move
data around.
2023-08-23 17:32:17 -04:00
Lioncache 787b4f37fb Arm64/VectorOps: Remove redundant moves from SVE VSMin when possible
When the destination and first source alias one another, then we can
perform the operation in place without moving any data.
2023-08-23 17:30:08 -04:00
Lioncache c3faa019f5 Arm64/VectorOps: Remove redundant moves from SVE VUMin when possible
If the destination and first source alias, then we can perfom the operation
in place.
2023-08-23 17:27:52 -04:00
Ryan Houdek da098d8204 Merge pull request #2979 from lioncash/div
Arm64/VectorOps: Remove moves from SVE VFDiv if possible
2023-08-23 14:22:46 -07:00