In order to reuse code buffers, when invalidating we will no longer be
able to invalidate code while inside the code we are invalidating.
Changes the IR operation to be a block ender itself and control jumping
to the new RIP. This way it can invalidate from the dispatcher and just
restart JIT execution.
Have it treated as a block terminator IR and instead of doing
syscall+exitfunction, just do syscall and jump in to the dispatcher.
This is necessary for code invalidation and reuse for code that is stuck
in long running syscalls. We'll need to do something similar for thunks
later, although the recursive nature of thunk+callback can make that a
little squirrely.
I forgot on Linux by default we didn't have the syscall instructions
count as block end. Change this so that it counts as block end now.
This has the additional benefit that now the frontend needs to modify
the RIP manually as well which is fine as it's what arm64ec and wow64
does.
Also add back the UnimplementedOp in RDPID that accidentally got caught
up. Also increment DiskCache version as both changes will change
codegen.
Fixes#5942
CHPE_V2_CPU_AREA_INFO is the mechanism used on ARM64EC to coordinate
thread state, suspension and exception handling with the kernel.
It's typically set up by ntdll for ARM64EC processes, but Wine uses
the TEB pointer without restricting where it's coming from. Set
ChpeV2CpuAreaInfo in FEX WOW64 thread initialization.
Fixes 32-bit debug events and a number of CPU context handling tests.
Reduces the number of wineserver round-trips required for suspend and
avoids the need for remote thread creation for inter-process suspends.
We would check the FileID of mapped sections, but BSS is an anonymous mapping.
Grab the ELF image extents when we parse the file, and add an additional check
to the relocation filter to bail out additional relocs if we know that size.
The LookupKey based on data available at Lookup time gives a list of possible
candidate entries, which may have different guest code sizes/footprints.
If only one candidate, it's stored in-line in the map like before - if we grow
past that, an additional (multi)map is allocated, sorted by footprint.
The footprint sorting lets us reduce the amount of hashes performed at Lookup
to the strict minimum.
Anon entries keys are based on a limited size decoded prefix, so this is an
important part of geting the best hit rate possible out of all the other work
with anon keys, as there's often multiple candidates with the same prefix.
Cap the max amount of entries per bucket to limit growth.
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
Otherwise GuestSize ends up different across compiles of the same valid code
extents, including the copy of it baked in JITCodeTail of the host code itself.
All buffers should be disowned leaving their respective compilation
sites, and reowning a buffer should never have the flag already be
owned.
Throw an assert in both cases because that would be a programming error
and result in some squirrely buffer handling
Decode a few bytes in advance to get a hashable prefix to use as key.
Generate touched pages dynamically since they can be misaligned now, as the
cached-hit guest code isn't necessarily in the same spot as the store was.
PR #5902 technically introduced a bug where we would read past the end
of bounds for thunk instructions when full smc was enabled. Luckily this
never occurs in practice as the Mono hacks never are on VDSO boundaries,
and no one is expected to enable full smc detection really.
Switch this path over to using crc32 unconditionally. This raises our
minspec technically to armv8-a+crc, but nothing that matters shipped
without crc so it's fine.
This also is a minor speed and JIT size reduction due less branches
polluting the BTB. But really only for mono/unity games.
Requires revving the DiskCache version again.
Ensures that whenever a file handle is transferred anywhere, that the
moved from instance won't end up closing the file handle when the
destructor runs.
Does less hashing, improves hit rate when there's data adjacent to code,
and/or when the SMC check makes us rebuild code that hasn't actually been
changed.
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.
Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.
This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.
Bumps the DiskCache version again because it causes codegen to change.
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.
Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.
Use printable characters in FOZ key names as intended.
Turns out packed enum classes without specifying an underlying type
causes problems. Declare its underlying type as uint32_t to match
everything else here.