NFC
- Renames CodeBufferManager to SharedCodeBufferManager to be more
explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
of the CPUBackend code
Makes it easier to parse ownership and lifetime semantics of these
buffers.
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.
NFC
STLXR cannot use the same register as both the status register and the
value register, otherwise it's architecturally unpredictable
behavior.
Only applies to hardware without FEAT_LSE, so this only meaningfully
affects hardware using the v8.0 spec, since FEAT_LSE becomes mandatory
in v8.1 and newer.
Lets us shave off an instruction and also avoid using a temporary
register in some cases. We can also tweak our worst case that requires a
predicate to eliminate the temporary as well.
We can also expand our cmpps cases, so that we can reflect the
BSL2N usages in instcountci.
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.
Tiny saving, but reduces overall register use in some cases.
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.
So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.
microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```
microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```