fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.
Misunderstood the assignment before.
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
- stlxp does monitor check before alignment check, use loads for all
for consistency
- Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
ARMv9.0-a hardware
- Adds product name to CPUID
- Oryon-3 being CPU PartID 2 isn't a mistake.
- No distinction between Oryon-1 and Oryon-2, both are partid 1.
cpuinfo:
```
processor : 0
BogoMIPS : 38.40
Features : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part : 0x002
CPU revision : 1
```
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.
Fixes `Devil May Cry 4`
Docker's seccomp filter fails to follow AAPCS64 and SysV zero-extension
rules. For values smaller than 64-bit they were required in their
seccomp filters to truncate the value to the specific size but do not.
Instead they do a 64-bit comparison operation against smaller arguments
(in this case 32-bit). This means 64-bit -1 and 32-bit -1 passed through
have different values for this `personality` syscall.
The real fix would be for Docker to audit their seccomp filter rules and
ensure they zero-extend every argument that is smaller than 64-bit, but
we don't control that. So there is likely to be more bugs in their
filter that we encounter, this is just an easy one to resolve.
This was causing a surprisingly high amount of branch mispredicts in
Death Stranding. Suspend doorbell is fairly rare so just invert the
check and move the target down out of the hot path. Then the doorbell
handling code will trampoline to the correct location still.
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.
Instead compare it to the appname that now gets passed over to it for
matching.
This allows us to do the following:
```
{
"Config": {
"ProfileStats": "1",
"X87ReducedPrecision": "1",
"TSOEnabled": "1",
"VectorTSOEnabled": "0",
"MemcpySetTSOEnabled": "0",
"HalfBarrierTSOEnabled":"1",
"MaxInst": "500",
"Multiblock": "1"
},
"AppOverrides" : {
"setup*" : {
"Comment": [
"292030 - The Witcher 3: Wild Hunt"
],
"X87ReducedPrecision": "0"
}
}
}
```
Based on #121 which needs to get merged first.
Code Review
Code Review: Class deletion
Add support for question mark and plus mark in regex, supply testing for star
Added more characters to the regex alphabets, add more test case
Added support for regex matching of configs, awaiting reviews
Rename variable to CamelCase
Addresses PR reviews
Remove unnecessary features and test cases
Rewrite to naive regex with dp
Addresses PR reviews
Build fixes
Code Review
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.
With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.