Ryan Houdek
c0bb6a053f
IR: Removes implicit sized {Create,Extract}ElementPair
2023-08-28 16:50:00 -07:00
Ryan Houdek
bc1e89d91d
Merge pull request #3020 from Sonicadvance1/remove_implicit_ops_pt_atomic
...
Remove implicit sized IR ops part atomic
2023-08-28 07:27:47 -07:00
Ryan Houdek
c1b4c11e54
Merge pull request #3018 from Sonicadvance1/opcodedispatcher_sizeopt
...
OpcodeDispatcher: Optimize Get{Src,Dst}Size
2023-08-28 07:13:41 -07:00
Ryan Houdek
a36427d01e
IR: Removes implicit sized CAS
2023-08-28 07:12:31 -07:00
Ryan Houdek
594baff705
IR: Removes implicit sized CASPair
2023-08-28 07:12:31 -07:00
Ryan Houdek
6dcfd6eb73
IR: Removes non-opsize AtomicXor
2023-08-28 07:12:31 -07:00
Ryan Houdek
4f8a63459c
IR: Removes non-opsize AtomicSwap
2023-08-28 07:12:31 -07:00
Ryan Houdek
ab230bf527
IR: Removes non-opsize AtomicFetchAdd
2023-08-28 07:12:31 -07:00
Ryan Houdek
8ba9613972
IR: Removes non-opsize AtomicFetchSub
2023-08-28 07:12:31 -07:00
Ryan Houdek
9370e30af4
IR: Removes non-opsize AtomicFetchAnd
2023-08-28 07:12:31 -07:00
Ryan Houdek
25ce57ef92
IR: Removes non-opsize AtomicFetchOr
2023-08-28 07:12:31 -07:00
Ryan Houdek
436e0f6f86
IR: Removes non-opsize AtomicFetchXor
2023-08-28 07:12:31 -07:00
Ryan Houdek
f44c6f394f
IR: Removes non-opsize AtomicFetchNeg
2023-08-28 07:12:31 -07:00
Ryan Houdek
bb1362b2bf
OpcodeDispatcher: Optimize Get{Src,Dst}Size
...
These functions are called a lot....a lot a lot.
Optimize these in to a couple of ALU operations instead of a whole
table lookup. Confirming with output assembly that this becomes more
optimal.
2023-08-28 05:29:56 -07:00
Ryan Houdek
51baf594c9
OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
...
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek
16481c0e55
IR: Removes implicit sized rev
2023-08-28 05:04:32 -07:00
Ryan Houdek
b00310674a
IR: Removes implicit sized not
2023-08-28 05:04:32 -07:00
Ryan Houdek
1f1473eb74
IR: Removes implicit sized neg
2023-08-28 05:04:32 -07:00
Ryan Houdek
cccd7001cb
IR: Removes implicit sized ashr
2023-08-28 05:04:32 -07:00
Ryan Houdek
ce8392d5ae
IR: Removes implicit sized ror
2023-08-28 05:02:01 -07:00
Ryan Houdek
386cf36cfd
IR: Removes implicit sized sbfe
...
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.
We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek
0ddc23a5c9
IR: Removes implicit sized Popcount
2023-08-28 05:02:01 -07:00
Ryan Houdek
b95648a4ab
IR: Removes implicit sized FindMSB
2023-08-28 05:02:01 -07:00
Ryan Houdek
bf18672999
IR: Removes implicit sized FindLSB
2023-08-28 05:02:01 -07:00
Ryan Houdek
f405c1be69
IR: Removes inverted EntrypointOffset/InlineEntrypointOffset
2023-08-28 05:02:01 -07:00
Ryan Houdek
f3cd115fa7
IR: Removes implicit sized CountLeadingZeroes
2023-08-28 05:02:00 -07:00
Ryan Houdek
a14720e130
IR: Removes implicit sized FindTrailingZeroes
2023-08-28 05:02:00 -07:00
Ryan Houdek
86ed909de8
IR: Removes implicit sized DIV/REM
2023-08-28 05:02:00 -07:00
Ryan Houdek
ac7e75c06b
IR: Removes implicit sized UDIV/UREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
ec3e7ceeb5
IR: Removes implicit sized EXTR
2023-08-28 05:02:00 -07:00
Ryan Houdek
c8c8ddbd4f
IR: Removes implicit sized PDEP/PEXT
2023-08-28 05:02:00 -07:00
Ryan Houdek
ea6d068cc5
IR: Removes implicit sized LDIV/LREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
5013473ec0
IR: Removes implicit sized LUDIV/LUREM
2023-08-28 05:02:00 -07:00
Ryan Houdek
6f2b3e76ac
Merge pull request #3013 from Sonicadvance1/32bit_sra
...
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek
e4bb0df486
IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
...
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.
This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.
Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek
572d6cd3e6
OpcodeDispatcher: Fixes ADC and SBB
2023-08-26 18:22:50 -07:00
Ryan Houdek
fcc37bf6a8
OpcodeDispatcher: Fixes RCR and ADOX 32-bit
...
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek
eace648fa9
OpcodeDispatcher: Fixes bug in UMUL
...
This was trying to operating on a 32-bit value but BFE the upper
32-bits.
Actually fixes this so it is operating on the 64-bit multiply result.
2023-08-26 18:22:50 -07:00
Ryan Houdek
a76c2c57b0
OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
...
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek
f3679a99ec
OpcodeDispatcher: Optimize nontemporal moves
...
These are now optimal.
2023-08-25 03:13:10 -07:00
Ryan Houdek
c441b238c7
OpcodeDispatcher: Optimize MMX conversion operation
...
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Lioncache
26c81224ac
OpcodeDispatcher: Remove redundant moves from AESIMC
...
Zero-extension will occur automatically upon storing if necessary.
We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache
52ab3f6a1e
OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
...
Zero-extension will occur upon storing if necessary.
We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache
bcba3700c8
OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
...
Zero-extension will automatically occur if necessary upon storing.
We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache
c409ea78bc
OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
...
Zero-extension will occur automatically upon storing if necessary.
We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Lioncache
d99bcbf01b
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
...
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Lioncache
e5f5629ffc
OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
...
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Lioncache
fa17d9fae9
OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
...
We already do a zero-extend if necessary in StoreResult.
This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Ryan Houdek
a523858f66
Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
...
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek
8b051b5e63
OpcodeDispatcher: Implement support for push IR operation
...
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.
While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
2023-08-18 14:19:16 -07:00