If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.
Difference at 95.0% confidence
-0.00138911 +/- 0.00104724
-0.418608% +/- 0.315587%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.
Difference at 95.0% confidence
-0.00474273 +/- 0.00119189
-1.40908% +/- 0.354114%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.
This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!
But now as we are no longer utilizing it, it is time to remove it.
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.
To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset
If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.
Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.
(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.
(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).
(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).
(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).
(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).
(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
The host stack pointer can't be reused since explicit bounds checks
would be far too expensive, and on-stack signals prevent implicit ones
using guard pages from working.
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.
This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.
The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.
Been a while since I last looked at this, added a new instcountci file
to show the improvement.
Unifies the interface, so that there's no need for a stark difference.
Previously, calling the single/double variant didn't require explicit template
arguments, but the half-float version did, which is inconsistent.
Technically it also made the interface more cumbersome to use in the event the
element size isn't able to be determined as a constant ahead of time.