Ryan Houdek
514a8223d9
OpcodeDispatcher: Optimizes SSE movmaskps
...
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.
I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek
8d110738ac
IR: Add option to disable vector shift range clamping
...
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.
Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.
When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Ryan Houdek
9ba46f429e
X8764: Ensure frndint uses host rounding mode
...
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.
This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.
In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.
[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00
Ryan Houdek
a76c2c57b0
OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
...
This is way more optimal!
2023-08-25 12:59:40 -07:00
Mai
bf12f08218
Merge pull request #3002 from Sonicadvance1/optimize_movmaskpd
...
OpcodeDispatcher: Optimize 128-bit movmaskpd
2023-08-25 08:50:28 -04:00
Mai
1f7d138d2a
Merge pull request #3008 from Sonicadvance1/optimize_movddup
...
OpcodeDispatcher: Optimize movddup from register
2023-08-25 08:48:40 -04:00
Mai
f36f07055a
Merge pull request #3007 from Sonicadvance1/optimize_cvtdq2pd
...
OpcodeDispatcher: Optimize cvtdq2pd from register source
2023-08-25 08:47:47 -04:00
Mai
30a1a382c4
Merge pull request #3006 from Sonicadvance1/optimize_movq
...
OpcodeDispatcher: Optimizes movq
2023-08-25 08:46:02 -04:00
Ryan Houdek
2fbcf2e4a9
OpcodeDispatcher: Optimize movddup from register
...
This is now optimal
2023-08-25 03:44:10 -07:00
Ryan Houdek
00124205e5
OpcodeDispatcher: Optimize cvtdq2pd from register source
...
This is now optimal
2023-08-25 03:39:24 -07:00
Ryan Houdek
81281e2115
OpcodeDispatcher: Optimizes movq
...
Removes a redundant move between registers and makes it optimal.
Also removes a couple redundant moves on the avx version.
2023-08-25 03:29:58 -07:00
Ryan Houdek
1cb2b084b3
OpcodeDispatcher: Generate more optimal code for scalar GPR converts
...
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
FPR and converted it in-place.
These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek
9a54898429
OpcodeDispatcher: Optimize 128-bit movmaskpd
...
I'd consider this optimal now.
Thanks to @dougallj for the optimization idea again!
2023-08-24 17:27:54 -07:00
Ryan Houdek
80d871fb18
Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
...
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek
3731e6d88b
OpcodeDispatcher: Optimize cvtps2pd
...
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek
c441b238c7
OpcodeDispatcher: Optimize MMX conversion operation
...
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek
a1210f892a
OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
...
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.
Needs #2994 merged first.
Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek
565b30e15e
OpcodeDispatcher: Cache named vector constants in the block
...
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Ryan Houdek
f300196d90
OpcodeDispatcher: Optimize AddSubP{S,D}
...
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.
Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Mai
66c6f96120
Merge pull request #2990 from Sonicadvance1/optimize_pmulh
...
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek
77b6d854b9
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
...
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Lioncache
26c81224ac
OpcodeDispatcher: Remove redundant moves from AESIMC
...
Zero-extension will occur automatically upon storing if necessary.
We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache
8a622a3c1a
OpcodeDispatcher: Remove redundant moves from VAESEnc
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache
f4848fd1a7
OpcodeDispatcher: Remove redundant moves from VAESEncLast
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache
d37ce08ae9
OpcodeDispatcher: Remove redundant move from VAESDec
...
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache
a6f1a9f8e8
OpcodeDispatcher: Remove redundant moves from VAESDecLast
...
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache
52ab3f6a1e
OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
...
Zero-extension will occur upon storing if necessary.
We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache
410e99ba09
OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
...
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek
fc4559d3c4
OpcodeDispatcher: Use new IR ops for pack instructions
...
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Mai
4b06069c0d
Merge pull request #2972 from Sonicadvance1/optimize_scalar_mov
...
OpcodeDispatcher: Optimizes scalar movd/movq
2023-08-23 16:30:45 -04:00
Ryan Houdek
b646f4b781
Merge pull request #2978 from lioncash/misc
...
OpcodeDispatcher: Remove redundant moves from remaining AVX ops
2023-08-23 13:18:25 -07:00
Ryan Houdek
8836ab8988
OpcodeDispatcher: Optimizes scalar movd/movq
...
MMX and SSE versions are now optimal.
2023-08-23 12:56:11 -07:00
Lioncache
ea9747289a
OpcodeDispatcher: Remove redundant moves from remaining AVX ops
...
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 15:31:59 -04:00
Lioncache
735e2060a3
OpcodeDispatcher: Remove redundant moves from VPACKUSOP/VPACKSSOp
...
Zero-extension will occur automatically if necessary.
2023-08-23 15:09:57 -04:00
Ryan Houdek
86ef6fe48d
Merge pull request #2976 from lioncash/mov
...
OpcodeDispatcher: Remove unnecessary moves from AVX move ops where applicable
2023-08-23 12:02:05 -07:00
Lioncache
8e7e91d61f
OpcodeDispatcher: Remove redundant moves in VMOVLPOp
...
Zero-extension will automatically occur upon storing if necessary.
2023-08-23 14:23:57 -04:00
Lioncache
bcba3700c8
OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
...
Zero-extension will automatically occur if necessary upon storing.
We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache
7d05797e82
OpcodeDispatcher: Remove redundant moves from VMOVHPOp
...
Zero-extension will automatically occur if necessary upon storing.
2023-08-23 14:14:35 -04:00
Lioncache
c409ea78bc
OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
...
Zero-extension will occur automatically upon storing if necessary.
We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Ryan Houdek
a62ba75ede
Merge pull request #2975 from lioncash/scalar
...
Arm64/ConversionOps: Add scalar support to Vector_FToI
2023-08-23 10:57:38 -07:00
Lioncache
4a7ef3da13
OpcodeDispatcher: Remove unnecessary moves in AVXVectorRound
...
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 13:41:56 -04:00
Lioncache
990b70dcd6
OpcodeDispatcher: Use scalar rounding for scalar round instructions
2023-08-23 13:34:01 -04:00
Ryan Houdek
6624f50abf
Merge pull request #2974 from lioncash/extend
...
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
2023-08-23 10:25:56 -07:00
Lioncache
d99bcbf01b
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
...
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Lioncache
f516aed4b7
OpcodeDispatcher: Remove unnecessary moves in AVXVFCMPOp
...
Zero-extension will already occur if necessary upon storing.
2023-08-23 12:37:45 -04:00
Lioncache
3858e4124b
OpcodeDispatcher: Remove redundant moves in AVX blend special cases
...
Zero-extension will happen if necessary upon storing.
2023-08-23 00:08:33 -04:00
Mai
819fe110da
Merge pull request #2967 from Sonicadvance1/optimize_storeelement
...
OpcodeDispatcher: Optimize MOVHP{S,D}
2023-08-23 00:01:25 -04:00
Ryan Houdek
adfd6787c0
Merge pull request #2969 from lioncash/insert
...
OpcodeDispatcher: Remove unnecessary moves from AVX inserts
2023-08-22 20:51:55 -07:00
Ryan Houdek
0ee2579a5e
OpcodeDispatcher: Optimize MOVHP{S,D}
...
Loads can turn in to element Loads.
Stores can turn in to element stores.
These four instruction variants are now optimal.
2023-08-22 20:42:24 -07:00
Mai
bb2f7107cd
Merge pull request #2963 from Sonicadvance1/optimize_loadelement
...
OpcodeDispatcher: Optimize MOVLP{S,D} loads
2023-08-22 23:41:59 -04:00