This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.
So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.
microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```
microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.
Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:
```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
88 | struct prctl_mm_map {
| ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
134 | struct prctl_mm_map {
| ^
1 error generated.
```
prefer <sys/prctl.h> and do not include <linux/prctl.h>
fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.
It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.
Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
- stlxp does monitor check before alignment check, use loads for all
for consistency
- Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
ARMv9.0-a hardware
- Adds product name to CPUID
- Oryon-3 being CPU PartID 2 isn't a mistake.
- No distinction between Oryon-1 and Oryon-2, both are partid 1.
cpuinfo:
```
processor : 0
BogoMIPS : 38.40
Features : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part : 0x002
CPU revision : 1
```
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.
Fixes `Devil May Cry 4`
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.