Commit Graph
124 Commits
Author SHA1 Message Date
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00
Ryan Houdek 8bc84c202c IR: Removes implicit sized xor 2023-08-28 19:51:14 -07:00
Ryan Houdek e9a3848602 Merge pull request #3027 from Sonicadvance1/remove_implicit_andn
IR: Removes implicit sized andn
2023-08-28 19:39:49 -07:00
Ryan Houdek 1699ec9a76 IR: Removes implicit sized andn 2023-08-28 19:16:16 -07:00
Ryan Houdek db6c8852fc IR: Removes implicit sized or 2023-08-28 19:06:05 -07:00
Ryan Houdek 65dc6f3e90 IR: Removes implicit sized lshr 2023-08-28 18:16:56 -07:00
Ryan Houdek 60c4438780 IR: Removes implicit sized lshl 2023-08-28 17:50:41 -07:00
Ryan Houdek b9e4a1423f IR: Removes sext IR helper
You hold no power here IR operation.
2023-08-28 17:03:38 -07:00
Ryan Houdek 48669b7006 IR: Removes implicit sized orlshl/orlshr 2023-08-28 05:04:32 -07:00
Ryan Houdek 5768444ce9 IR: Removes implicit sized abs 2023-08-28 05:04:32 -07:00
Ryan Houdek ce8392d5ae IR: Removes implicit sized ror 2023-08-28 05:02:01 -07:00
Ryan Houdek 386cf36cfd IR: Removes implicit sized sbfe
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.

We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek b95648a4ab IR: Removes implicit sized FindMSB 2023-08-28 05:02:01 -07:00
Ryan Houdek bf18672999 IR: Removes implicit sized FindLSB 2023-08-28 05:02:01 -07:00
Ryan Houdek 514a8223d9 OpcodeDispatcher: Optimizes SSE movmaskps
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.

I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek 8d110738ac IR: Add option to disable vector shift range clamping
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.

Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.

When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Ryan Houdek 9ba46f429e X8764: Ensure frndint uses host rounding mode
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.

This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.

In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.

[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00
Ryan Houdek a76c2c57b0 OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
This is way more optimal!
2023-08-25 12:59:40 -07:00
Mai bf12f08218 Merge pull request #3002 from Sonicadvance1/optimize_movmaskpd
OpcodeDispatcher: Optimize 128-bit movmaskpd
2023-08-25 08:50:28 -04:00
Mai 1f7d138d2a Merge pull request #3008 from Sonicadvance1/optimize_movddup
OpcodeDispatcher: Optimize movddup from register
2023-08-25 08:48:40 -04:00
Mai f36f07055a Merge pull request #3007 from Sonicadvance1/optimize_cvtdq2pd
OpcodeDispatcher: Optimize cvtdq2pd from register source
2023-08-25 08:47:47 -04:00
Mai 30a1a382c4 Merge pull request #3006 from Sonicadvance1/optimize_movq
OpcodeDispatcher: Optimizes movq
2023-08-25 08:46:02 -04:00
Ryan Houdek 2fbcf2e4a9 OpcodeDispatcher: Optimize movddup from register
This is now optimal
2023-08-25 03:44:10 -07:00
Ryan Houdek 00124205e5 OpcodeDispatcher: Optimize cvtdq2pd from register source
This is now optimal
2023-08-25 03:39:24 -07:00
Ryan Houdek 81281e2115 OpcodeDispatcher: Optimizes movq
Removes a redundant move between registers and makes it optimal.
Also removes a couple redundant moves on the avx version.
2023-08-25 03:29:58 -07:00
Ryan Houdek 1cb2b084b3 OpcodeDispatcher: Generate more optimal code for scalar GPR converts
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
   FPR and converted it in-place.

These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek 9a54898429 OpcodeDispatcher: Optimize 128-bit movmaskpd
I'd consider this optimal now.

Thanks to @dougallj for the optimization idea again!
2023-08-24 17:27:54 -07:00
Ryan Houdek 80d871fb18 Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek 3731e6d88b OpcodeDispatcher: Optimize cvtps2pd
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek a1210f892a OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.

Needs #2994 merged first.

Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek 565b30e15e OpcodeDispatcher: Cache named vector constants in the block
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.

Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.

Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Ryan Houdek f300196d90 OpcodeDispatcher: Optimize AddSubP{S,D}
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.

Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Mai 66c6f96120 Merge pull request #2990 from Sonicadvance1/optimize_pmulh
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek 77b6d854b9 OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 8a622a3c1a OpcodeDispatcher: Remove redundant moves from VAESEnc
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache f4848fd1a7 OpcodeDispatcher: Remove redundant moves from VAESEncLast
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache d37ce08ae9 OpcodeDispatcher: Remove redundant move from VAESDec
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache a6f1a9f8e8 OpcodeDispatcher: Remove redundant moves from VAESDecLast
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache 410e99ba09 OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek fc4559d3c4 OpcodeDispatcher: Use new IR ops for pack instructions
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Mai 4b06069c0d Merge pull request #2972 from Sonicadvance1/optimize_scalar_mov
OpcodeDispatcher: Optimizes scalar movd/movq
2023-08-23 16:30:45 -04:00
Ryan Houdek b646f4b781 Merge pull request #2978 from lioncash/misc
OpcodeDispatcher: Remove redundant moves from remaining AVX ops
2023-08-23 13:18:25 -07:00
Ryan Houdek 8836ab8988 OpcodeDispatcher: Optimizes scalar movd/movq
MMX and SSE versions are now optimal.
2023-08-23 12:56:11 -07:00