Commit Graph
1791 Commits
Author SHA1 Message Date
Ryan Houdek 23572539f8 Merge pull request #3991 from Sonicadvance1/fix_feat_lrcpc_sigbus
Arm64: Fixes SIGBUS handler for FEAT_LRCPC
2024-08-22 13:47:39 -07:00
Alyssa Rosenzweig ce8e6e5c0c OpcodeDispatcher: optimize adc 0
clang generates this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-22 10:51:06 -04:00
Ryan Houdek be9c5678c0 Arm64: Fixes SIGBUS handler for FEAT_LRCPC
When #3899 refactored some of this code, it had split an if-else chain
in to two independent if-else chains. Turns out there was a minor
dependency between them. Add back the few checks necessary to ensure
that FEAT_LRCPC falls through before hitting the FEAT_LSE checks.

This is necessary because their instructions share an encoding so the
masks between the two instruction classes overlap slightly.
2024-08-21 18:01:55 -07:00
Alyssa Rosenzweig 6951284924 OpcodeDispatcher: optimize ADCX/ADOX
rewrite to use native adcs with flag fixups.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig b03c613fbf OpcodeDispatcher: optimize JP/JNP
fuse and+cbnz into tbz/tbnz.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig b893bdb8df OpcodeDispatcher: optimize fcmovu
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig db45f6eec8 OpcodeDispatcher: optimize test with small immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig d53e689e22 IR: model tbz/tbnz
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig 5d613e8716 OpcodeDispatcher: optimize test x, x
fewer uops now that we invert carry

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 21:17:08 -04:00
Alyssa Rosenzweig 894aaa980f Merge pull request #3975 from Sonicadvance1/update_wfe_comment
SpinWaitLock: Update comment about WFE spurious wakeups
2024-08-20 20:55:30 -04:00
Ryan Houdek 877b2f4fef Merge pull request #3955 from alyssarosenzweig/opt/pop-final
Rearrange SRA to let us coalesce cmpxchg moves
2024-08-20 17:36:11 -07:00
Alyssa Rosenzweig ffb85e6305 ArchHelpers: rearrange SRA layout to coalesce cmpxchg
linux only for now, arm64ec should do something similar.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 20:14:08 -04:00
Alyssa Rosenzweig d7a20fa28f OpcodeDispatcher: fix tso checks
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 20:11:06 -04:00
Alyssa Rosenzweig 4c4c6e7807 OpcodeDispatcher: pair loads on cortex
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:36:50 -04:00
Alyssa Rosenzweig a8c9c71ce3 OpcodeDispatcher: use loadcontextpair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8a4bd5f22c OpcodeDispatcher: pair avx128 save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 138a36c69a OpcodeDispatcher: pair avx128 restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 3400ca5d42 OpcodeDispatcher: pair avx restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 5d8164da4e OpcodeDispatcher: pair restore x87
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8b89b30a6f OpcodeDispatcher: pair restore SSE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig e66d7cfd6c OpcodeDispatcher: pair mm save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 63afa29dae OpcodeDispatcher: pair save AVX
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 926a9b40e2 OpcodeDispatcher: pair save mxcsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 9bb43264b5 OpcodeDispatcher: pair SSE store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig 32d6daf558 OpcodeDispatcher: use ldp/stp for AVX load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig bebcb73c68 OpcodeDispatcher: pair AVX high writes
reduces instr count with AVX-128

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig d0e040514f IR: add LoadContextPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 850c027d52 IR: add AllocateFPR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 5df90563e2 IR: add LoadMemPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 7a489d18c6 IR: add StoreMemPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 1db092e96f IR: add StoreContextPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Ryan Houdek 689b461d7b Merge pull request #3984 from Sonicadvance1/atomic_tso_subchecks
FEXCore: Splits up atomic enablement checks
2024-08-20 15:16:54 -07:00
Ryan Houdek 9c8438f264 FEXCore: Splits up atomic enablement checks
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.

Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.

This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.

Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
2024-08-20 14:34:34 -07:00
Ryan Houdek caf7ad53e6 Arm64: Allow directly correlating an ARM register back to an x86 register
Fixes a bug that is getting introduced in to #3955 when it rearranged
register allocation.
2024-08-20 14:17:13 -07:00
Alyssa Rosenzweig 84e4960e52 OpcodeDispatcher: optimize btc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 09:24:02 -04:00
Ryan Houdek 99afd876ba Merge pull request #3976 from alyssarosenzweig/opt/huffman-
Improvements from bytemark "huffman"
2024-08-19 16:10:29 -07:00
Alyssa Rosenzweig e5cb583cc0 OpcodeDispatcher: optimize 8-bit rol
rotate right by less than 8:

  ror(________________7654321076543120, ...)

rotate left by less than 8:

  rol(76543210________________76543210, ...)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig 3ef83a9e23 OpcodeDispatcher: optimize ror al, cl
rely on the masking.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig 7de6a5cb07 OpcodeDispatcher: optimize rotate flags
introduce an IR op for it so we can reason about flags across rotate
instructions easily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 10:40:53 -04:00
Alyssa Rosenzweig e06b3a4186 ConstProp: optimize and x, -1
mitigates regression from previous patch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 08:54:47 -04:00
Alyssa Rosenzweig c495f82f4f OpcodeDispatcher: optimize 8/16-bit bitwise with constants
avoid masking. apparently even modern compilers will do cute tricks with 8-bit
math in hot loops ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 08:54:47 -04:00
Alyssa Rosenzweig 49e4426ed5 RedundantFlagCalculationElimination: drop register writes
this allows us to form real cmp instructions when we know PF is dead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 08:54:47 -04:00
Alyssa Rosenzweig c4e2436885 ConstProp: optimize add eax, -1 and friends
need to sign extend to get the constant inlined.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-19 08:54:47 -04:00
Tony Wasserka e5149fba57 Merge pull request #3965 from Sonicadvance1/followup_3964
Library Forwarding: Follow up from #3964
2024-08-19 11:22:43 +02:00
Ryan Houdek 0ea3de95e0 SpinWaitLock: Update comment about WFE spurious wakeups
With recent bug fixes, WFE now can sleep for roughly as long as the
programmed architecture timer of 100 microseconds. Still nowhere near as
close as what x86 CPUs can get with waitx and waitpkg, because those
don't get spuriously woken up by an architecture timer.

100 microseconds is a significantly improvement over 52ns although.
2024-08-18 17:50:20 -07:00
Ryan Houdek d1ec242e4f FEXCore: Stop installing static library
This hasn't ever mattered. Our API isn't even stable enough to really
install the shared library.
2024-08-16 19:55:00 -07:00
Ryan Houdek a82fcdecd7 PR review comments 2024-08-16 10:21:18 -07:00
Ryan Houdek 27acbe305d ArchHelpers: Adjust ClearICache for its usage
Instead of clearing a hardcoded 16 bytes, adjust for the actual number
of instructions modified. The implementation will still only clear a
single cacheline so it doesn't change behaviour.
2024-08-16 10:21:18 -07:00
Ryan Houdek cd0739a534 Arm64Helpers: Moves instruction definitions to implementation
This used to exist in the FEXCore header since the unaligned handler was
done in the frontend. Once it got moved in to FEXCore it had stayed
there. Move it over now.
2024-08-16 10:21:18 -07:00
Ryan Houdek f1055d0713 Arm64: On backpatch ensure DMB instructions get patched in first
In the case of a visibility tear when one thread is backpatching while
another is executing. The executing thread can /potentially/ see the
writing of instructions depending on coherency rules or filling of
cachelines.

By ensuring the DMB instructions are backpatched over the NOP
instructions first, this ensures correct atomic visibility even on tear.
2024-08-16 10:21:18 -07:00