The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
- stlxp does monitor check before alignment check, use loads for all
for consistency
- Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
ARMv9.0-a hardware
- Adds product name to CPUID
- Oryon-3 being CPU PartID 2 isn't a mistake.
- No distinction between Oryon-1 and Oryon-2, both are partid 1.
cpuinfo:
```
processor : 0
BogoMIPS : 38.40
Features : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part : 0x002
CPU revision : 1
```
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.
Fixes `Devil May Cry 4`
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.
With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
Disables THP on some key locations that are fairly sparse
- rpmalloc
- This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
- These get in the hundreds of megabytes, while not being sparse they
trend towards only using a handful of pages and ballooning to 2MB
per thread is quite heavy.
- Lookup cache
- L1 specifically gets hit here which adds a decent chunk of overhead
due to sparsity.
Win32 for all of these also aren't handled, but that will need to be a
followup.
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.
Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.
This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.
Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.
We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.