Adds a config option to disable pruning if desired.
Had to change FEXCore::HostFeatures slightly to remove the HostType from
the Hash for caching, as we don't want the bucket to inadvertently
delete wow64,arm64ec,linux buckets when executing in the same namespace.
Minor changes to the DiskCache::Init code to handle returning two
values, moving the HostType to be a part of the "ProcessBucketHash"
instead of the "MachineBucketHash".
Then uses the directory walking function added in the previous commit to
find mismatched/old "MachineBucketHash" directories and remove them.
The base directory is now the machine bucket, which is everything that affects
codegen that you can ask FEX about offline: FEX compiler version, host features.
The DB files now carry a process bucket, which is everything else that can
affect codegen, on the user/process side: 32-bitness and config.
Both together are still the bucket hash that gets laced into every lookup key,
so we still won't load mismatching code because a file is in the wrong spot.
By going through relocations before computing/hashing extents, we can simply
undo datamasks that haven't been claimed by the backend instead of failing
the whole Store. They fall back to being part of the normal hashed extents.
We would check the FileID of mapped sections, but BSS is an anonymous mapping.
Grab the ELF image extents when we parse the file, and add an additional check
to the relocation filter to bail out additional relocs if we know that size.
The LookupKey based on data available at Lookup time gives a list of possible
candidate entries, which may have different guest code sizes/footprints.
If only one candidate, it's stored in-line in the map like before - if we grow
past that, an additional (multi)map is allocated, sorted by footprint.
The footprint sorting lets us reduce the amount of hashes performed at Lookup
to the strict minimum.
Anon entries keys are based on a limited size decoded prefix, so this is an
important part of geting the best hit rate possible out of all the other work
with anon keys, as there's often multiple candidates with the same prefix.
Cap the max amount of entries per bucket to limit growth.
Decode a few bytes in advance to get a hashable prefix to use as key.
Generate touched pages dynamically since they can be misaligned now, as the
cached-hit guest code isn't necessarily in the same spot as the store was.
Does less hashing, improves hit rate when there's data adjacent to code,
and/or when the SMC check makes us rebuild code that hasn't actually been
changed.
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.
Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.
Use printable characters in FOZ key names as intended.
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.
We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.
So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.
Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.
Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.
All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.
Add some bounds checking to index population to protect against corruption.
Serializes code blocks to disk - only blocks coming from known regions, for now
Disabled by default, key and versioning still needs work, but works for testing