In order to reuse code buffers, when invalidating we will no longer be
able to invalidate code while inside the code we are invalidating.
Changes the IR operation to be a block ender itself and control jumping
to the new RIP. This way it can invalidate from the dispatcher and just
restart JIT execution.
Have it treated as a block terminator IR and instead of doing
syscall+exitfunction, just do syscall and jump in to the dispatcher.
This is necessary for code invalidation and reuse for code that is stuck
in long running syscalls. We'll need to do something similar for thunks
later, although the recursive nature of thunk+callback can make that a
little squirrely.
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
All buffers should be disowned leaving their respective compilation
sites, and reowning a buffer should never have the flag already be
owned.
Throw an assert in both cases because that would be a programming error
and result in some squirrely buffer handling
UpdateTopForPop_Slow() now invalidates ST(0)'s tag by default,
so every slow-path StackPop() does this consistently instead of
the previous special case only in OP_POPSTACKDESTROY.
FINCSTP is an exception as it only moves the stack pointer without
invalidating the tag.
PR #5902 technically introduced a bug where we would read past the end
of bounds for thunk instructions when full smc was enabled. Luckily this
never occurs in practice as the Mono hacks never are on VDSO boundaries,
and no one is expected to enable full smc detection really.
Switch this path over to using crc32 unconditionally. This raises our
minspec technically to armv8-a+crc, but nothing that matters shipped
without crc so it's fine.
This also is a minor speed and JIT size reduction due less branches
polluting the BTB. But really only for mono/unity games.
Requires revving the DiskCache version again.
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.
Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.
This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.
Bumps the DiskCache version again because it causes codegen to change.
Gets rid of potential extraneous copies. We also add handling for cases
where two passes with the same name are unintentionally added.
Previously we'd blindly overwrite the mapping.
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
Previously this wouldn't have worked, since .first isn't a valid member.
The only reason it wasn't caught is because the function is never
instantiated.
This needs to divide by 8 to get a proper byte size for all type sizes.
The only usage of this is currently a uint64_t, so it worked by
coincidence, since sizeof(uint64_t) == 8.