When #3899 refactored some of this code, it had split an if-else chain
in to two independent if-else chains. Turns out there was a minor
dependency between them. Add back the few checks necessary to ensure
that FEAT_LRCPC falls through before hitting the FEAT_LSE checks.
This is necessary because their instructions share an encoding so the
masks between the two instruction classes overlap slightly.
PR #3980 is adding a feature to merge loadstores in to paired
loadstores, but it was using the incorrect atomic check to determine if
it can safely merge them or not. It was using the GPR atomic check
instead of the vector atomic check. While this would improve performance
on Apple Silicon with its hardware TSO implementation, it would have had
zero impact on Cortex and Oryon.
Instead split out the three config options to live as a boolean check in
the ContextImpl similar to how we disable "AtomicTSOEmulation". Removing
the various configs in the JIT and CPUID so that it queries from the
same context. This makes it clearer that if you are wanting the current
active configuration for memcpy, vector, or general atomic TSO
emulation, you should query one of those three getters.
This also fixes a weird edge case bug in the arm64 JIT where you could
have TSO emulation disable, but still have vector TSO enabled partially.
Just because half a config wasn't checked in {Load,Store}MemTSO for
vectors. If the global "TSOEnabled" option is disabled then TSO should
always be disabled.
Alyssa will be able to pull this in to #3980 once merged and get the
performance uplift on Cortex and Oryon, since our default configuration
is to have vector and memcpy TSO emulation disabled.
rotate right by less than 8:
ror(________________7654321076543120, ...)
rotate left by less than 8:
rol(76543210________________76543210, ...)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
avoid masking. apparently even modern compilers will do cute tricks with 8-bit
math in hot loops ...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
With recent bug fixes, WFE now can sleep for roughly as long as the
programmed architecture timer of 100 microseconds. Still nowhere near as
close as what x86 CPUs can get with waitx and waitpkg, because those
don't get spuriously woken up by an architecture timer.
100 microseconds is a significantly improvement over 52ns although.
Instead of clearing a hardcoded 16 bytes, adjust for the actual number
of instructions modified. The implementation will still only clear a
single cacheline so it doesn't change behaviour.
This used to exist in the FEXCore header since the unaligned handler was
done in the frontend. Once it got moved in to FEXCore it had stayed
there. Move it over now.
In the case of a visibility tear when one thread is backpatching while
another is executing. The executing thread can /potentially/ see the
writing of instructions depending on coherency rules or filling of
cachelines.
By ensuring the DMB instructions are backpatched over the NOP
instructions first, this ensures correct atomic visibility even on tear.