Ryan Houdek
594baff705
IR: Removes implicit sized CASPair
2023-08-28 07:12:31 -07:00
Ryan Houdek
6dcfd6eb73
IR: Removes non-opsize AtomicXor
2023-08-28 07:12:31 -07:00
Ryan Houdek
4f8a63459c
IR: Removes non-opsize AtomicSwap
2023-08-28 07:12:31 -07:00
Ryan Houdek
ab230bf527
IR: Removes non-opsize AtomicFetchAdd
2023-08-28 07:12:31 -07:00
Ryan Houdek
8ba9613972
IR: Removes non-opsize AtomicFetchSub
2023-08-28 07:12:31 -07:00
Ryan Houdek
9370e30af4
IR: Removes non-opsize AtomicFetchAnd
2023-08-28 07:12:31 -07:00
Ryan Houdek
25ce57ef92
IR: Removes non-opsize AtomicFetchOr
2023-08-28 07:12:31 -07:00
Ryan Houdek
436e0f6f86
IR: Removes non-opsize AtomicFetchXor
2023-08-28 07:12:31 -07:00
Ryan Houdek
f44c6f394f
IR: Removes non-opsize AtomicFetchNeg
2023-08-28 07:12:31 -07:00
Ryan Houdek
51baf594c9
OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
...
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek
48669b7006
IR: Removes implicit sized orlshl/orlshr
2023-08-28 05:04:32 -07:00
Ryan Houdek
16481c0e55
IR: Removes implicit sized rev
2023-08-28 05:04:32 -07:00
Ryan Houdek
b00310674a
IR: Removes implicit sized not
2023-08-28 05:04:32 -07:00
Ryan Houdek
5768444ce9
IR: Removes implicit sized abs
2023-08-28 05:04:32 -07:00
Ryan Houdek
1f1473eb74
IR: Removes implicit sized neg
2023-08-28 05:04:32 -07:00
Ryan Houdek
cccd7001cb
IR: Removes implicit sized ashr
2023-08-28 05:04:32 -07:00
Ryan Houdek
ce8392d5ae
IR: Removes implicit sized ror
2023-08-28 05:02:01 -07:00
Ryan Houdek
386cf36cfd
IR: Removes implicit sized sbfe
...
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.
We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek
0ddc23a5c9
IR: Removes implicit sized Popcount
2023-08-28 05:02:01 -07:00
Ryan Houdek
b95648a4ab
IR: Removes implicit sized FindMSB
2023-08-28 05:02:01 -07:00
Ryan Houdek
bf18672999
IR: Removes implicit sized FindLSB
2023-08-28 05:02:01 -07:00
Ryan Houdek
f405c1be69
IR: Removes inverted EntrypointOffset/InlineEntrypointOffset
2023-08-28 05:02:01 -07:00
Ryan Houdek
f3cd115fa7
IR: Removes implicit sized CountLeadingZeroes
2023-08-28 05:02:00 -07:00
Ryan Houdek
a14720e130
IR: Removes implicit sized FindTrailingZeroes
2023-08-28 05:02:00 -07:00
Ryan Houdek
86ed909de8
IR: Removes implicit sized DIV/REM
2023-08-28 05:02:00 -07:00
Ryan Houdek
ac7e75c06b
IR: Removes implicit sized UDIV/UREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
ec3e7ceeb5
IR: Removes implicit sized EXTR
2023-08-28 05:02:00 -07:00
Ryan Houdek
c8c8ddbd4f
IR: Removes implicit sized PDEP/PEXT
2023-08-28 05:02:00 -07:00
Ryan Houdek
62a9a075b7
Arm64: Leave a comment that 32-bit division shouldn't leave garbage in upper 64-bits
2023-08-28 05:02:00 -07:00
Ryan Houdek
ea6d068cc5
IR: Removes implicit sized LDIV/LREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
5013473ec0
IR: Removes implicit sized LUDIV/LUREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
6f2b3e76ac
Merge pull request #3013 from Sonicadvance1/32bit_sra
...
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek
1d7c280367
Merge pull request #3012 from Sonicadvance1/optimize_movmskps
...
OpcodeDispatcher: Optimizes SSE movmaskps
2023-08-27 21:29:04 -07:00
Ryan Houdek
514a8223d9
OpcodeDispatcher: Optimizes SSE movmaskps
...
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.
I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek
8d110738ac
IR: Add option to disable vector shift range clamping
...
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.
Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.
When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Ryan Houdek
e4bb0df486
IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
...
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.
This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.
Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek
572d6cd3e6
OpcodeDispatcher: Fixes ADC and SBB
2023-08-26 18:22:50 -07:00
Ryan Houdek
fcc37bf6a8
OpcodeDispatcher: Fixes RCR and ADOX 32-bit
...
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek
8f7925d06f
Arm64: Simple typo fix
2023-08-26 18:22:50 -07:00
Ryan Houdek
eace648fa9
OpcodeDispatcher: Fixes bug in UMUL
...
This was trying to operating on a 32-bit value but BFE the upper
32-bits.
Actually fixes this so it is operating on the 64-bit multiply result.
2023-08-26 18:22:50 -07:00
Ryan Houdek
a4ac21a4e4
OpcodeDispatcher: Fixes bug in GetRFLAG with CachedNZCV
...
This was only operating at byte size but it was attempting to get bit
offsets at greater than operating size.
Change operating size over to 64-bit.
2023-08-26 18:22:50 -07:00
Ryan Houdek
a01e69092d
Arm64: Ensure Bfe and Sbfe operate at 32-bit or 64-bit op size
...
For Sbfe at least it ensures the upper bits don't get filled with
garbage.
Bfe it doesn't change behaviour but best to be correct.
2023-08-26 18:22:50 -07:00
Ryan Houdek
2fde2140ef
Arm64: Ensure assert is testing correct array
2023-08-26 18:22:50 -07:00
Ryan Houdek
9ba46f429e
X8764: Ensure frndint uses host rounding mode
...
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.
This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.
In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.
[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00
Ryan Houdek
a76c2c57b0
OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
...
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek
7f63d87295
IR: Adds support for new LoadNamedVectorIndexedConstant IR
2023-08-25 12:59:40 -07:00
Mai
bf12f08218
Merge pull request #3002 from Sonicadvance1/optimize_movmaskpd
...
OpcodeDispatcher: Optimize 128-bit movmaskpd
2023-08-25 08:50:28 -04:00
Mai
1f7d138d2a
Merge pull request #3008 from Sonicadvance1/optimize_movddup
...
OpcodeDispatcher: Optimize movddup from register
2023-08-25 08:48:40 -04:00
Mai
f36f07055a
Merge pull request #3007 from Sonicadvance1/optimize_cvtdq2pd
...
OpcodeDispatcher: Optimize cvtdq2pd from register source
2023-08-25 08:47:47 -04:00
Mai
30a1a382c4
Merge pull request #3006 from Sonicadvance1/optimize_movq
...
OpcodeDispatcher: Optimizes movq
2023-08-25 08:46:02 -04:00