Ryan Houdek
ab8ee64352
Merge pull request #3497 from Sonicadvance1/movmaskb_constant
...
JIT: Optimize pmovmaskb with a named vector constant
2024-03-18 16:08:40 -07:00
Alyssa Rosenzweig
2a9fcc6a66
Merge pull request #3492 from Sonicadvance1/implement_prefetch
...
OpcodeDispatcher: Implement support for the various prefetch instructions
2024-03-18 07:49:47 -04:00
Ryan Houdek
fd391b1b18
JIT: Optimize pmovmaskb with a named vector constant
...
I was looking at some other JIT overheads and this cropped up as some
overhead. Instead of materializing a constant using mov+movk+movk+movk,
load it from the named vector constant array.
In a micro-benchmark this improved performance by 34%.
In bytemark this improved on subbench by 0.82%
2024-03-17 18:40:46 -07:00
Ryan Houdek
f79991a9d8
OpcodeDispatcher: Implement rdpid
...
Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
2024-03-14 20:07:58 -07:00
Ryan Houdek
ca6b2e43e6
Merge pull request #3491 from alyssarosenzweig/rclse/waw
...
RCLSE: Optimize store-after-store
2024-03-14 03:23:05 -07:00
Ryan Houdek
8056bee82b
OpcodeDispatcher: Implement support for the various prefetch instructions
...
x86 has a few prefetch instructions.
- prefetch - One of two classic 3DNow! instructions
- Prefetch in to L1 data cache
- prefetchw - One of two classic 3DNow! instructions
- Implies prefetch in to L1 data cache
- Prefetch cacheline with intent to write and exclusive ownership
- prefetchnta
- Prefetch non-temporal data in respect to /all/ cache levels
- Assumes inclusive caches?
- prefetch{t0,t1,t2}
- Prefetch data with respect to each cache level
- T0 = L1 and higher
- T1 = L2 and higher
- T2 = L3 and higher
**Some silly duplicates**
- prefetchwt1
- Duplicate of prefetchw but explicitly L1 data cache
- prefetch_exclusive
- Duplicate of prefetch
God Of War 2018 uses prefetchw as a hint for exclusive ownership of the
cacheline in some very aggressive spin-loops. Let's implement the
operations to help it along.
2024-03-12 21:37:31 -07:00
Ryan Houdek
cc635a54f8
IR: Implements support for prefetch operation
2024-03-12 21:19:50 -07:00
Ryan Houdek
217d9d8c50
ARMEmitter: Fixes prfm with negative or unaligned offsets
2024-03-12 21:18:23 -07:00
Alyssa Rosenzweig
7629007cfa
OpcodeDispatcher: allow upper garbage on STOS
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
03c6abdad4
OpcodeDispatcher: optimize DF add
...
fuse the shift the right way
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
c99cbe6d0a
JIT: switch DF representation
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
e3ee65e491
OpcodeDispatcher: use transformed DF for memset/memcpy
...
Use the 1/-1 representation instead of 0/1. This will be better by the end of
the series.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
aee00f524c
OpcodeDispatcher: use DF retrieval helpers
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
a76321c6c1
OpcodeDispatcher: add DF retrieval helpers
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
85f8ad3842
JIT: fix sha256msg1 encoding
...
botched move in the !tied reg case.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:41:23 -04:00
Alyssa Rosenzweig
11880459a5
OpcodeDispatcher: use SETF for DEC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig
0ef0bb2c97
OpcodeDispatcher: use SETF for INC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig
72edee7c6f
IR: add SETF8/SETF16 ir ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-01 19:40:53 -04:00
Ryan Houdek
009ae55ff0
Merge pull request #3475 from alyssarosenzweig/opt/lock-dec
...
Optimize lock dec
2024-02-29 08:44:24 -08:00
Ryan Houdek
98572b9e23
Merge pull request #3473 from Sonicadvance1/remove_mov_swap
...
Arm64: Stop moving source in atomic swap
2024-02-29 08:44:16 -08:00
Alyssa Rosenzweig
fed5e6d546
OpcodeDispatcher: use fetchadd for atomic DEC
...
Avoids a NEG.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-29 09:28:21 -04:00
Ryan Houdek
c318947695
Arm64: Stop moving source in atomic swap
...
ldswpal doesn't overwrite the source register and only reads the bits
required for the sized operation.
Not sure exactly why we were doing a copy here.
Removing it means improving Skyrim's hottest code block, as seen in #3472
2024-02-29 03:07:05 -08:00
Alyssa Rosenzweig
811487ad98
OpcodeDispatcher: use real branch for INT
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig
4f4e38ace2
OpcodeDispatcher: use real branch for rep cmps
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig
edd6becc56
OpcodeDispatcher: use real branch for rep scas
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig
e47a94cae7
OpcodeDispatcher: skip mask with shld
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:18:29 -04:00
Alyssa Rosenzweig
cc82dba1ca
OpcodeDispatcher: use mvn for AF with constants
...
This reduces pointless constant usage. For now, it's no net change to
instcountci, but it should make it easier to get wins later. Hopefully.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:18:28 -04:00
Alyssa Rosenzweig
8232669b22
OpcodeDispatcher: simplify
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:00:13 -04:00
Alyssa Rosenzweig
28936073c4
OpcodeDispatcher: allow upper garbage on NEG
...
like SUB.
due to RA silliness, this is a loss for inst count but a win for cycles.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 09:59:57 -04:00
Alyssa Rosenzweig
ef2559d911
OpcodeDispatcher: allow garbage with SCAS
...
it's just feeding SUB flags which allow it
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 09:16:08 -04:00
Ryan Houdek
f346f89678
OpcodeDispatcher: Don't use AddShift with no shift
...
This accidentally removed optimizations elsewhere that was only checking
for Add.
2024-02-27 19:56:12 -08:00
Alyssa Rosenzweig
49e798ab2b
Merge pull request #3461 from alyssarosenzweig/opt/sbc
...
Optimize SBC
2024-02-27 11:29:45 -04:00
Ryan Houdek
946c805d84
Merge pull request #3459 from Sonicadvance1/fix_591
...
Capture a 64-bit process trying to jump to 32-bit syscall handler
2024-02-26 21:57:03 -08:00
Alyssa Rosenzweig
12cc980603
OpcodeDispatcher: shuffle adc flag order
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
f3d55dd721
OpcodeDispatcher: shuffle SBC flag order
...
avoids clobbering nzcv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
91cef6b76f
OpcodeDispatcher: use native ADC even for 8/16-bit
...
we mask off the upper bits, and they agree in the lower bits.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
270cbf39b5
OpcodeDispatcher: specialize SALC
...
this gets rid of the awkward non-flag SBB case, which streamlines SBB. while
getting better codegen for the demon opcode (-:
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:33:41 -04:00
Alyssa Rosenzweig
2e0be0a5e7
OpcodeDispatcher: allow more upper garbage with adc
...
missed this last series.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:42:13 -04:00
Alyssa Rosenzweig
d60c089697
OpcodeDispatcher: allow upper garbage with sbb
...
for the usual reasons
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig
e76ebeab58
OpcodeDispatcher: use 1-op "src + CF"
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig
333271d490
OpcodeDispatcher: fuse sbb when flags calculated
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig
a750870abf
OpcodeDispatcher: use fused sbcs calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:25:03 -04:00
Alyssa Rosenzweig
15db72ef60
IR: add Sbb, SbbWithFlags ops
...
For fusing sbc+sbcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:25:03 -04:00
Ryan Houdek
4f028b8614
Capture a 64-bit process trying to jump to 32-bit syscall handler
...
Fixes #591
Adds a simple unittest
2024-02-26 05:37:29 -08:00
Alyssa Rosenzweig
1e153e0c81
OpcodeDispatcher: allow garbage with adcs
...
for the usual reasons
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-25 10:50:02 -04:00
Alyssa Rosenzweig
6994fc3a01
IR,OpcodeDispatcher,JIT: fuse adcs flags
...
The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!
Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-25 10:49:32 -04:00
Ryan Houdek
d703f3ccee
Fixes zero register flag generation
...
Fixes 140976d322
Adds a unit test to ensure it keeps working.
2024-02-24 16:32:25 -08:00
Alyssa Rosenzweig
80e632db8a
OpcodeDispatcher: garbage collect
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
852e3c4e93
OpcodeDispatcher: fuse XADD
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
e86547bbcb
OpcodeDispatcher: fuse INC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00