Alyssa Rosenzweig
ce8e6e5c0c
OpcodeDispatcher: optimize adc 0
...
clang generates this.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-22 10:51:06 -04:00
Alyssa Rosenzweig
6951284924
OpcodeDispatcher: optimize ADCX/ADOX
...
rewrite to use native adcs with flag fixups.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
b03c613fbf
OpcodeDispatcher: optimize JP/JNP
...
fuse and+cbnz into tbz/tbnz.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
b893bdb8df
OpcodeDispatcher: optimize fcmovu
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
db45f6eec8
OpcodeDispatcher: optimize test with small immediate
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
d53e689e22
IR: model tbz/tbnz
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
5d613e8716
OpcodeDispatcher: optimize test x, x
...
fewer uops now that we invert carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 21:17:08 -04:00
Ryan Houdek
877b2f4fef
Merge pull request #3955 from alyssarosenzweig/opt/pop-final
...
Rearrange SRA to let us coalesce cmpxchg moves
2024-08-20 17:36:11 -07:00
Alyssa Rosenzweig
ffb85e6305
ArchHelpers: rearrange SRA layout to coalesce cmpxchg
...
linux only for now, arm64ec should do something similar.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 20:14:08 -04:00
Alyssa Rosenzweig
d7a20fa28f
OpcodeDispatcher: fix tso checks
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 20:11:06 -04:00
Alyssa Rosenzweig
4c4c6e7807
OpcodeDispatcher: pair loads on cortex
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:36:50 -04:00
Alyssa Rosenzweig
a8c9c71ce3
OpcodeDispatcher: use loadcontextpair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8a4bd5f22c
OpcodeDispatcher: pair avx128 save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
138a36c69a
OpcodeDispatcher: pair avx128 restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
3400ca5d42
OpcodeDispatcher: pair avx restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
5d8164da4e
OpcodeDispatcher: pair restore x87
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8b89b30a6f
OpcodeDispatcher: pair restore SSE
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
e66d7cfd6c
OpcodeDispatcher: pair mm save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
63afa29dae
OpcodeDispatcher: pair save AVX
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
926a9b40e2
OpcodeDispatcher: pair save mxcsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
9bb43264b5
OpcodeDispatcher: pair SSE store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig
32d6daf558
OpcodeDispatcher: use ldp/stp for AVX load/store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
bebcb73c68
OpcodeDispatcher: pair AVX high writes
...
reduces instr count with AVX-128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
d0e040514f
IR: add LoadContextPair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
850c027d52
IR: add AllocateFPR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
5df90563e2
IR: add LoadMemPair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
7a489d18c6
IR: add StoreMemPair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
1db092e96f
IR: add StoreContextPair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Ryan Houdek
689b461d7b
Merge pull request #3984 from Sonicadvance1/atomic_tso_subchecks
...
FEXCore: Splits up atomic enablement checks
2024-08-20 15:16:54 -07:00
Ryan Houdek
9c8438f264
FEXCore: Splits up atomic enablement checks
...
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.
Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.
This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.
Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
2024-08-20 14:34:34 -07:00
Ryan Houdek
caf7ad53e6
Arm64: Allow directly correlating an ARM register back to an x86 register
...
Fixes a bug that is getting introduced in to #3955 when it rearranged
register allocation.
2024-08-20 14:17:13 -07:00
Alyssa Rosenzweig
84e4960e52
OpcodeDispatcher: optimize btc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 09:24:02 -04:00
Alyssa Rosenzweig
e5cb583cc0
OpcodeDispatcher: optimize 8-bit rol
...
rotate right by less than 8:
ror(________________7654321076543120, ...)
rotate left by less than 8:
rol(76543210________________76543210, ...)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig
3ef83a9e23
OpcodeDispatcher: optimize ror al, cl
...
rely on the masking.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig
7de6a5cb07
OpcodeDispatcher: optimize rotate flags
...
introduce an IR op for it so we can reason about flags across rotate
instructions easily.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig
c495f82f4f
OpcodeDispatcher: optimize 8/16-bit bitwise with constants
...
avoid masking. apparently even modern compilers will do cute tricks with 8-bit
math in hot loops ...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-19 08:54:47 -04:00
Ryan Houdek
ef4c4f6e9b
FEXCore: Disable vixl linking if vixl disasm or simulator is disabled
...
This was mostly there, just needed to remove some extraneous headers and
only insert vixl in to the library list if the options were enabled.
2024-08-16 07:29:41 -07:00
Ryan Houdek
2eb7a9ff28
OpcodeDispatcher: Remove old bad assumption in INC/DEC
...
Somewhere there was an assumption made that INC and DEC supported the
repeat prefix. This isn't actually the case, while the prefix can be
encoded, it is a nop and should only expect to be used for padding.
Adds a unittest to ensure that behaviour is as expected.
2024-08-15 10:22:31 -07:00
Ryan Houdek
933c65d805
Merge pull request #3956 from alyssarosenzweig/opt/pop-return
...
small optimizations for returns
2024-08-15 03:30:23 -07:00
Ryan Houdek
df0ecad15b
Merge pull request #3933 from bylaws/arm64-suspend
...
Support cooperative suspend on ARM64EC
2024-08-15 01:22:28 -07:00
Alyssa Rosenzweig
2dc92c122e
BranchOps: micro-optimize ExitFunction
...
this should be slightly faster on Firestorm and no worse on recent cortex
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 16:59:34 -04:00
Alyssa Rosenzweig
0db17bd58a
OpcodeDispatcher: optimize IRET
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 16:59:34 -04:00
Alyssa Rosenzweig
aac16493f1
OpcodeDispatcher: optimize RET
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 16:59:34 -04:00
Ryan Houdek
aa5d2ff31c
Merge pull request #3951 from alyssarosenzweig/opt/pops
...
Add a hack for multiple destinations & make good use of it
2024-08-14 12:03:00 -07:00
Alyssa Rosenzweig
c70e44cf05
OpcodeDispatcher: extract Push helper
...
Mirrors the Pop helper we added. This cleans up a bunch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
ede13e37d9
OpcodeDispatcher: optimize Thunk
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
a7dada457a
OpcodeDispatcher: optimize POPF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
6367554d30
OpcodeDispatcher: optimize LEAVE
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
81969a684e
OpcodeDispatcher: optimize pop segment
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
ee339b5960
OpcodeDispatcher: optimize POPA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00