Implements an offline JIT compiler backend that compiles x86 code blocks
from a given code map into an ARM64 code cache on Windows. Separate
binaries are built for WOW64 (32-bit) and ARM64EC (64-bit) targets.
NOTE: This patch originally added a separate binary; instead the new
functionality is added to the existing FEXOfflineCompiler and will
be properly integrated in the next patches.
This operation will take care of any pending code cache operations:
* import new code maps from `CACHE_DIR/codemap/new` and process them to `CACHE_DIR/codemap/ready`
* generate caches for updated code maps with new blocks
* ensure caches already exist for all other code maps (and generate them if needed)
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.
It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
This allows the context and parent thread objects to be created earlier,
allowing the VDSO and ELFCodeLoader mapping functions to have a thread
object for tracking memory mappings through the regular guest routines.
This means that we no longer need to do any form of deferred handling
for code caching as all the state is ready early in the initialization
process.
A little bit of care needed to be taken to ensure we still close the
ELFCodeLoader's FDs later and that VDSO unmapping happens before tearing
down the parent thread, but overall this is mostly just passing the
InternalThreadState object around as normal.
I couldn't find any functional regression from this change alongside
code caching, but it would be good for @neobrain to double check this.