The LookupKey based on data available at Lookup time gives a list of possible
candidate entries, which may have different guest code sizes/footprints.
If only one candidate, it's stored in-line in the map like before - if we grow
past that, an additional (multi)map is allocated, sorted by footprint.
The footprint sorting lets us reduce the amount of hashes performed at Lookup
to the strict minimum.
Anon entries keys are based on a limited size decoded prefix, so this is an
important part of geting the best hit rate possible out of all the other work
with anon keys, as there's often multiple candidates with the same prefix.
Cap the max amount of entries per bucket to limit growth.
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
Otherwise GuestSize ends up different across compiles of the same valid code
extents, including the copy of it baked in JITCodeTail of the host code itself.
Return the same as the NonExecutableRange for now. I have a unittest
that is mostly correct with this change, just need to fixup some of the
supplementary data.
For now push this change to get a bad assert removed that can happen
during code discovery. I'll fix up the remaining details afterwards.
Just be very explicit about this rather than questioning why MMX moves
can't be a nop in the regular vector move code.
Doesn't change codegen so binarycacheversion doesn't need to change.
All buffers should be disowned leaving their respective compilation
sites, and reowning a buffer should never have the flag already be
owned.
Throw an assert in both cases because that would be a programming error
and result in some squirrely buffer handling
UpdateTopForPop_Slow() now invalidates ST(0)'s tag by default,
so every slow-path StackPop() does this consistently instead of
the previous special case only in OP_POPSTACKDESTROY.
FINCSTP is an exception as it only moves the stack pointer without
invalidating the tag.
Decode a few bytes in advance to get a hashable prefix to use as key.
Generate touched pages dynamically since they can be misaligned now, as the
cached-hit guest code isn't necessarily in the same spot as the store was.
PR #5902 technically introduced a bug where we would read past the end
of bounds for thunk instructions when full smc was enabled. Luckily this
never occurs in practice as the Mono hacks never are on VDSO boundaries,
and no one is expected to enable full smc detection really.
Switch this path over to using crc32 unconditionally. This raises our
minspec technically to armv8-a+crc, but nothing that matters shipped
without crc so it's fine.
This also is a minor speed and JIT size reduction due less branches
polluting the BTB. But really only for mono/unity games.
Requires revving the DiskCache version again.
Ensures that whenever a file handle is transferred anywhere, that the
moved from instance won't end up closing the file handle when the
destructor runs.
Does less hashing, improves hit rate when there's data adjacent to code,
and/or when the SMC check makes us rebuild code that hasn't actually been
changed.