Alyssa Rosenzweig
d3f1397325
OpcodeDispatcher: eliminate constants in RCR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
0a164428fa
OpcodeDispatcher: eliminate select in RCR
...
the nzcv clobber I actually came ofr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
7496175100
OpcodeDispatcher: optimize 32-bit rcl/rcr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
0616a9cef1
OpcodeDispatcher: eliminate move in rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
97f8775354
OpcodeDispatcher: optimize <32-bit rcr op1
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
c92099aa98
OpcodeDispatcher: fuse orlshl in rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
7c288b09f1
OpcodeDispatcher: rmif mask rcl smaller OF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
680af7b1b0
OpcodeDispatcher: rcr op 8x1 cleanup
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
349bc9efab
OpcodeDispatcher: unify rcr op 1bit codepaths
...
get additional opt for <32-bit
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
ad5c3cb268
OpcodeDispatcher: rmif mask for OF in rcr smaller
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
be8d37ef3d
OpcodeDispatcher: optimize 32-bit rol/ror imm
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
6ad2514bfe
OpcodeDispatcher: rmif mask rcl smaller cf
...
better on flagm. extra moves on non-flagm but, meh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
3fa6129a14
OpcodeDispatcher: rmif mask rcr smaller cf
...
and do some constant folding to do so more.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
a57cebaf58
OpcodeDispatcher: skip OF calc for constant rotate >= 2
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
34fdb14da1
OpcodeDispatcher: add and use AndConst
...
this skips the constant folding, which saves the branching in the rotate
immediate implementations.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
974baca09c
OpcodeDispatcher: allow upper garbage with rcl/rcr smaller
...
we're masking immediately to something smaller
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
f22094a493
OpcodeDispatcher: use a branch for 8/16-bit rotate flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
d979b3a1da
OpcodeDispatcher: note idea to further optimize rcl
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig
6d82c957fa
OpcodeDispatcher: fuse orlshl in rcl
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:22:57 -04:00
Mai
fa3352004e
Merge pull request #3381 from alyssarosenzweig/opt/masking
...
Allow upper garbage on a bunch of instructions
2024-01-30 10:07:53 -05:00
Mai
58f3d3caf5
Merge pull request #3380 from alyssarosenzweig/opt/pdep
...
Optimize PDEP
2024-01-29 13:27:15 -05:00
Alyssa Rosenzweig
16a54742e6
OpcodeDispatcher: optimize 32-bit tzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
ad9aa0bc87
OpcodeDispatcher: optimize 32-bit lzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
bd2b3f35a3
OpcodeDispatcher: optimize 32-bit popcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
50169ce640
OpcodeDispatcher: optimize 32-bit pext
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
baae2d68f9
OpcodeDispatcher: optimize 32-bit bextr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
ad8d038b8a
OpcodeDispatcher: optimize 32-bit blsi
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
820932e3c7
OpcodeDispatcher: optimize 32-bit blsmsk
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
6f11f2e6f4
OpcodeDispatcher: optimize 32-bit blsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig
f5ad7682c3
OpcodeDispatcher: optimize 32-bit pdep
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-29 13:06:56 -04:00
Ryan Houdek
56d8080ec9
Merge pull request #3345 from Sonicadvance1/fix_syscall_registers
...
OpcodeDispatcher: Fixes syscall rcx/r11 generation
2024-01-22 15:21:13 -08:00
Billy Laws
e323938173
FEXCore: Fix RCL/RCR shift wraparound behaviour
...
This ends up being cleaner to handle outside of
CalculateFlags_ShiftVariable as constant masking is only needed for
RCL/RCR.
2024-01-21 18:15:50 +00:00
Ryan Houdek
1f7a619c79
OpcodeDispatcher: Fixes syscall rcx/r11 generation
...
Noticed this while writing #3342 .
Fixes #3343
The syscall instruction is defined in the documentation that it will set
RCX to the next instruction's RIP and R11 to be RFLAGS. We entirely
skipped this which I noticed while writing unit tests.
Adds unittests to test both 32-bit and 64-bit behaviour because our
helper shares code with both.
I don't know if anything actually relied on this behaviour but we should
definitely support it.
2024-01-12 19:14:30 -08:00
Alyssa Rosenzweig
58127bd0e8
OpcodeDispatcher: optimize trivial cmpxchgs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-12 12:23:34 -04:00
Alyssa Rosenzweig
e8945dfb6d
OpcodeDispatcher: optimize gpr cmpxchg
...
NZCV stuff.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-12 12:03:28 -04:00
Ryan Houdek
4b3792196f
Merge pull request #3303 from Sonicadvance1/initial_runtime_longmode_switch
...
OpcodeDispatcher: Initial support for runtime long-mode switch
2024-01-04 18:17:54 -08:00
Alyssa Rosenzweig
04a88ed3ab
Merge pull request #3353 from Sonicadvance1/public_interface_cleaning
...
FEXCore interface cleaning
2024-01-03 15:14:54 -04:00
Ryan Houdek
d8f20751fe
FEXCore: Moves IREmitter from the public API to backend
...
No functional change
2023-12-25 07:00:29 -08:00
Ryan Houdek
bce694ebb5
FEXCore: Moves BitUtils to FHU
...
No functional change
2023-12-25 06:38:51 -08:00
Ryan Houdek
9e5d7aa5fe
OpcodeDispatcher: Initial support for runtime long-mode switch
...
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2023-12-21 01:54:19 -08:00
Mai
3d2cbc5d08
Merge pull request #3317 from Sonicadvance1/fix_imul_flags2
...
OpcodeDispatcher: Fixes flags generation in imul
2023-12-19 11:43:21 -05:00
Ryan Houdek
1a2f41922c
OpcodeDispatcher: Optimize SIB addr calculation
...
When the address calculation for SIB has both index and base then we can
optimize this to an add with a shifted register. This will convert a
three instruction sequence in to one instruction in most cases.
2023-12-15 13:08:46 -08:00
Ryan Houdek
5c6f229e76
OpcodeDispatcher: Fixes flags generation imul
...
On overflow with 32-bit we weren't setting the flags correctly.
2023-12-07 01:08:02 -08:00
Alyssa Rosenzweig
2a2c389be6
OpcodeDispatcher: rm masking in 32-bit andn
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-12-05 09:04:32 -04:00
Alyssa Rosenzweig
9417c93110
OpcodeDispatcher: remove outdated comment
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:23:41 -04:00
Alyssa Rosenzweig
068599b1ec
OpcodeDispatcher: use size-appropriate alu in bt*
...
saves zero-extending move for 32-bit ops.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
7216415bfc
OpcodeDispatcher: reorder flag calcs in bt*
...
saves big moves.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
3020626506
OpcodeDispatcher: unify bt/btc/bts/btr impls
...
they're all copypastes of each other, unify into one general "bit test & perform
action" template. this means most of the wins from the previous commits now
apply for bt* without more copypaste.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
0a79fa8d5d
OpcodeDispatcher: remove masking for 32/64-bit bt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:40:18 -04:00
Alyssa Rosenzweig
f8380b9adb
OpcodeDispatcher: use smaller shifts for BT
...
if the shift is < N, and we grab bit 0 after, we only need to consider <=N
bits of the source. this lets us use 32-bit lsr for 32-bit bt, which will
reduce masking in the next commit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:40:18 -04:00