Ryan Houdek
cec1814a09
Merge pull request #3384 from pmatos/CDQOp-Opt
...
Optimize CDQOp
2024-01-31 17:51:23 -08:00
Mai
4d49ac7c3d
Merge pull request #3387 from alyssarosenzweig/opt/rotates
...
Optimize rotates
2024-01-31 18:20:40 -05:00
Alyssa Rosenzweig
6d13d9fb56
Merge pull request #3395 from pmatos/StaticAnalysis
...
Code cleanup - mainly dead store removal; NFC
2024-01-31 17:24:48 -04:00
Mai
ae7dc250db
Merge pull request #3386 from alyssarosenzweig/opt/shift
...
Optimize shifts a bit
2024-01-31 14:11:58 -05:00
Paulo Matos
e4560ed0c8
Code cleanup - mainly dead store removal; NFC
...
scan-build found a few dead stores that can be easily cleaned-up
2024-01-31 08:35:55 +00:00
Alyssa Rosenzweig
f3eee8f305
OpcodeDispatcher: optimize bextr's length sanitize
...
reordering the operations saves an immediate move.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
f66085f4a7
OpcodeDispatcher: optimize bextr's (1 << x) - 1
...
little algebraic trick I cribbed from llvm
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
f3175848b1
OpcodeDispatcher: use lzcnt flag gen for tzcnt
...
as far as flags go, they're identical: set ZF for zero output, set CF for output
= DestSize, undef the rest. merge the impls, so we get the optimized lzcnt impl
for tzcnt.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
3dd597a591
OpcodeDispatcher: optimize lzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
fe88b904c9
OpcodeDispatcher: fix missing SF set with blsmsk
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
338f12845d
OpcodeDispatcher: save a constant in shld
...
one weird trick
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig
b3ae81f75f
OpcodeDispatcher: allow garbage on shld shift
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig
c1a1c37980
OpcodeDispatcher: mark ideas to improve SHLD
...
a bit tricky right now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig
fb6f850bb4
OpcodeDispatcher: remove rcl sub
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
b6d8749525
OpcodeDispatcher: remove select from rcl
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
d3f1397325
OpcodeDispatcher: eliminate constants in RCR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
0a164428fa
OpcodeDispatcher: eliminate select in RCR
...
the nzcv clobber I actually came ofr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
7496175100
OpcodeDispatcher: optimize 32-bit rcl/rcr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
0616a9cef1
OpcodeDispatcher: eliminate move in rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
97f8775354
OpcodeDispatcher: optimize <32-bit rcr op1
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
c92099aa98
OpcodeDispatcher: fuse orlshl in rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
7c288b09f1
OpcodeDispatcher: rmif mask rcl smaller OF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
680af7b1b0
OpcodeDispatcher: rcr op 8x1 cleanup
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
349bc9efab
OpcodeDispatcher: unify rcr op 1bit codepaths
...
get additional opt for <32-bit
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
ad5c3cb268
OpcodeDispatcher: rmif mask for OF in rcr smaller
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
be8d37ef3d
OpcodeDispatcher: optimize 32-bit rol/ror imm
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
6ad2514bfe
OpcodeDispatcher: rmif mask rcl smaller cf
...
better on flagm. extra moves on non-flagm but, meh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
3fa6129a14
OpcodeDispatcher: rmif mask rcr smaller cf
...
and do some constant folding to do so more.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
a57cebaf58
OpcodeDispatcher: skip OF calc for constant rotate >= 2
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
34fdb14da1
OpcodeDispatcher: add and use AndConst
...
this skips the constant folding, which saves the branching in the rotate
immediate implementations.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
974baca09c
OpcodeDispatcher: allow upper garbage with rcl/rcr smaller
...
we're masking immediately to something smaller
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
f22094a493
OpcodeDispatcher: use a branch for 8/16-bit rotate flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
d979b3a1da
OpcodeDispatcher: note idea to further optimize rcl
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
6d82c957fa
OpcodeDispatcher: fuse orlshl in rcl
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Mai
fa3352004e
Merge pull request #3381 from alyssarosenzweig/opt/masking
...
Allow upper garbage on a bunch of instructions
2024-01-30 10:07:53 -05:00
Mai
58f3d3caf5
Merge pull request #3380 from alyssarosenzweig/opt/pdep
...
Optimize PDEP
2024-01-29 13:27:15 -05:00
Paulo Matos
027fbbf051
Optimize CDQOp
2024-01-29 17:18:02 +00:00
Alyssa Rosenzweig
16a54742e6
OpcodeDispatcher: optimize 32-bit tzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
ad9aa0bc87
OpcodeDispatcher: optimize 32-bit lzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
bd2b3f35a3
OpcodeDispatcher: optimize 32-bit popcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
50169ce640
OpcodeDispatcher: optimize 32-bit pext
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
baae2d68f9
OpcodeDispatcher: optimize 32-bit bextr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
ad8d038b8a
OpcodeDispatcher: optimize 32-bit blsi
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
820932e3c7
OpcodeDispatcher: optimize 32-bit blsmsk
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
6f11f2e6f4
OpcodeDispatcher: optimize 32-bit blsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
f5ad7682c3
OpcodeDispatcher: optimize 32-bit pdep
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:06:56 -04:00
Ryan Houdek
56d8080ec9
Merge pull request #3345 from Sonicadvance1/fix_syscall_registers
...
OpcodeDispatcher: Fixes syscall rcx/r11 generation
2024-01-22 15:21:13 -08:00
Billy Laws
e323938173
FEXCore: Fix RCL/RCR shift wraparound behaviour
...
This ends up being cleaner to handle outside of
CalculateFlags_ShiftVariable as constant masking is only needed for
RCL/RCR.
2024-01-21 18:15:50 +00:00
Ryan Houdek
1f7a619c79
OpcodeDispatcher: Fixes syscall rcx/r11 generation
...
Noticed this while writing #3342 .
Fixes #3343
The syscall instruction is defined in the documentation that it will set
RCX to the next instruction's RIP and R11 to be RFLAGS. We entirely
skipped this which I noticed while writing unit tests.
Adds unittests to test both 32-bit and 64-bit behaviour because our
helper shares code with both.
I don't know if anything actually relied on this behaviour but we should
definitely support it.
2024-01-12 19:14:30 -08:00
Alyssa Rosenzweig
58127bd0e8
OpcodeDispatcher: optimize trivial cmpxchgs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-12 12:23:34 -04:00