A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.
Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.
Additionally some minorly incorrect assertions.
Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.
Fixes a spurious crash in Ender Magnolia.
Just helps when an unhandled ESR occurs, it was always a case of needing
to go in to the ARM ARM to decode it which was a bit of a pain. Add a
textual representation of it.
Fixes crash that occurs in applications that use both GL and Vulkan,
Like UE5 Vulkan native games. Fixes Ender Magnolia.
The issue here is that UE5 loads libGL first, which initializes our
libGL thunks, setting its X11Manager's functions.
It then loads libvulkan, which calls our oninit constructor, which
because of the symbol conflict, calls in to the libGL thunk's host
functions to reinitialize its function pointers, never initializing the
Vulkan X11Manager's functions. It would then crash as soon as an X11
function was used.
Give them unique symbol names so we don't accidentally look up the
incorrect symbol.
QEmu 10.2 is going to expose MIDR with Apple's vendor ID with variant 0.
That's the best they can do because they don't can't pin threads to
particular cores. So give a string for it, and detect it in the fit
script.
This reverts commit e1a45a2720, reversing
changes made to bd7edd8651.
The change rendered pressure-vessel non-functional on muvm-based setups
like Fedora Asahi Remix.
This was missed before, where the non-repeating strings instructions
were still using TSO even when the memcpy/set config option was
disabled. Make sure it listens to the config option and disable TSO in
those instances.
Noticed this while profiling Dishonored, and WINE's `sse2_memmove`
function was showing up as a high amount of CPU time. This is due to
them using non-repeating string operations on the header and tail of
their memmove to align to 16-byte.
With this fixed, it causes the game to go from ~62FPS to ~67FPS,
becoming bottlenecked by x87 emulation instead of memmove. Doing about
23 million soft-float operations per second, because it needs full
precision to remove some flickering artifacts.
- Cache miss counts
- Useful for determining if L2 cache or dynamic cache could help
- Cache read/write lock contention times
- Useful to see if threads are blocking each other on contention
- Read lock is the case where a read-lock is beneficial, even if we
currently use a write lock.
- JIT count
- Useful to see if any new JIT blocks are generating
On top of #4951 because it fiddles with the cache stuff.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
With our flags being optimized, this does even less than when it was
introduced. It's a hack, people are tinkering with it thinking it'll do
something. Get rid of it.
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.
Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)
The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.
The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.
Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.
Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.
Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.
These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.
Based on #4940 as a base which needs to be merged first.
Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
JIT Time: 0.486630 ms/second (0.00 percent)
Signal Time: 0.065880 ms/second (0.00 percent)
SIGBUS Cnt: 38 (38.160780 per second)
SMC Cnt: 0
Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
JIT resident: 95 mB
OpDispatcher resident: 38 mB
Frontend resident: 8 mB
CPUBackend resident: 624 kB
Lookup cache resident: 0 (null)
Lookup L1 cache resident: 7 mB
ThreadStates resident: 460 kB
Unaccounted resident: 217 mB
```
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.