This is unfiltered data that ends up in HostFeatures. If it was
serialized when set then it would effectively never allow the code
serialization to be used.
If we're shrinking the L1 cache then we just deleted the entry that we
just looked up. Add it back to ensure we don't get yet another lookup
for this entry.
CMAKE_CXX_FLAGS is a space-separated string rather than a semicolon-separated
list. Passing ${CMAKE_CXX_FLAGS} directly to execute_process(COMMAND ...)
passes the entire multi-flag string as a single argv argument to the compiler,
causing option parsing to fail when multiple flags are present (such as
flags configured via the CXXFLAGS environment variable).
Use separate_arguments() to convert CMAKE_CXX_FLAGS into a list so each flag
is passed as an individual argument.
Growing widens L1PointerMask without touching the table, so an entry whose
address has the new mask bit set sits where InvalidateCache no longer looks,
and a later shrink hands it back to the JIT. Wipe [0, old) on grow, the way
the shrink wipes [new, MAX); the still-valid entries go with it.
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.
Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.
Use printable characters in FOZ key names as intended.
Turns out packed enum classes without specifying an underlying type
causes problems. Declare its underlying type as uint32_t to match
everything else here.
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!
Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.
Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
This was using a one second alarm which could catch the JIT while it is
still busy. Instead wait for it to let an alarm thread that it is ready,
then tgkill the thread. This removes the race that this was hitting.
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.
Just a smidge of cleanup, NFC.
Serializes every option, even ones that are set to default to ensure
validation that if any config option value is added or changed that they
are captured.
Skips a handful of options that are either meta options, environment
options that don't matter, or HostFeatures which is handled elsewhere.
(HostFeatures will be controlling bucketing rather than the remaining
options).
This serialization is currently 1413 bytes and generates in 17580ns on
my A1A. So it's not the fastest, definitely don't want to be generating
it constantly per process.
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.
We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.
So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.
Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.
Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.
All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.