All buffers should be disowned leaving their respective compilation
sites, and reowning a buffer should never have the flag already be
owned.
Throw an assert in both cases because that would be a programming error
and result in some squirrely buffer handling
Decode a few bytes in advance to get a hashable prefix to use as key.
Generate touched pages dynamically since they can be misaligned now, as the
cached-hit guest code isn't necessarily in the same spot as the store was.
PR #5902 technically introduced a bug where we would read past the end
of bounds for thunk instructions when full smc was enabled. Luckily this
never occurs in practice as the Mono hacks never are on VDSO boundaries,
and no one is expected to enable full smc detection really.
Switch this path over to using crc32 unconditionally. This raises our
minspec technically to armv8-a+crc, but nothing that matters shipped
without crc so it's fine.
This also is a minor speed and JIT size reduction due less branches
polluting the BTB. But really only for mono/unity games.
Requires revving the DiskCache version again.
Ensures that whenever a file handle is transferred anywhere, that the
moved from instance won't end up closing the file handle when the
destructor runs.
Does less hashing, improves hit rate when there's data adjacent to code,
and/or when the SMC check makes us rebuild code that hasn't actually been
changed.
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.
Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.
This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.
Bumps the DiskCache version again because it causes codegen to change.
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.
Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.
Use printable characters in FOZ key names as intended.
Turns out packed enum classes without specifying an underlying type
causes problems. Declare its underlying type as uint32_t to match
everything else here.
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.
Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.
Just a smidge of cleanup, NFC.
Serializes every option, even ones that are set to default to ensure
validation that if any config option value is added or changed that they
are captured.
Skips a handful of options that are either meta options, environment
options that don't matter, or HostFeatures which is handled elsewhere.
(HostFeatures will be controlling bucketing rather than the remaining
options).
This serialization is currently 1413 bytes and generates in 17580ns on
my A1A. So it's not the fastest, definitely don't want to be generating
it constantly per process.
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.
Add some bounds checking to index population to protect against corruption.
Serializes code blocks to disk - only blocks coming from known regions, for now
Disabled by default, key and versioning still needs work, but works for testing
Useful for removing integer division instructions when we know the
source value is aligned to be power of two. As integer division is quite
slow, we want to use this when possible.
ShouldClose was never being set in the event we opened a regular file.
The only time it was set (to false) is when it's used to encapsulate
stderr and stdout.
So anything opened by a File instance was essentially held open.
According to POSIX open docs, this is a completely separate flag that
isn't a combination of O_RDONLY and O_WRONLY, so we need to handle this
separately.
Makes the codepath behaviorally symmetric with the Windows one.
Instead of burning roughly a million watts, put this spinloop on a WFE.
This tends to occur on a crash during shutdown that isn't fully able to
be avoided. The least we can do is not consume all the power in the
world.
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.
So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.
microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```
microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```