This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.
Not that big of a deal, it's only a minor optimization anyway.
Fixes#5227
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.
NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
This wasn't handling negatives correctly which was causing xalia.exe to
assert. Disable for now rather than further changing logic, with a TODO that it
should get fixed in the future.
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.
The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Instead of just a trivial pad being on or off, support a tri-state
on/off/auto where on will always pad, off will never pad, and auto will
pad only if code caching is enabled.
Further augment this by allowing a byte-width to be passed in, which can
be used with pointers to force a 48-bit VA width to only ever pad to
three instructions, reducing the common worst-case situation from 4
instructions to 3. This works because we're not going to expose a VA
width larger than 47-bit to the guest.
Fixes the handful of use-cases that explicitly chose their NOP padding,
and a bug in Arm64Relocations.cpp where it was incorrectly asking to not
receive padding even though it requires it.
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
While this worked great for the singular unit test. I remembered thatour
pool allocator returns the minimum working size asked for but will
return larger sizes if exact fitment couldn't occur.
Because we are dealing with guard pages, we need to return the full
buffer size to the "client" so they can tell the frontend where the
guard page actually lives. Otherwise the JIT will tell the frontend the
guard page is at the end of the requested size, blow past the limit,
and fault in a completely different location.
With a bit of logging I saw in a multithreaded environment that we were
basically always getting a larger requested buffer while Steam was
starting up.
When the JIT CodeBuffer overflows, we will now catch accesses to the
guard page and longjump while restarting the JIT with a larger buffer
request.
Fixes#4877
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.
We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.
Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.
This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).
But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
L1 cache residency can get quite large. Solution, start out small and
scale quickly on L1 cache misses but L2/L3 cache hits.
Some stats on L1 cache residency change:
- Teardown: 40MB -> 16MB (40%)
- Ender Lilies: 79MB -> 32MB (40.5%)
- Death Stranding: 186MB -> 93MB (50%)
- Steam: 75MB -> 7MB (9.3%)
The cost of this option is effectively free in our JIT. It changes a
single LDR to be a single LDP, which on Cortex CPUs cost the same. We do
this by moving the L1 pointer mask in to the CPUState object, making it
dynamic so it lives next to the L1 pointer. We then use that directly
rather than having the hardcoded value.
The lookup cache does a little bit of additional tracking and heuristics
to determine when the current L1 cache should increase or decrease in
size. From 128KB to 16MB per thread, allocating the full VA range as
previously.
Once the heuristic determines that L1 should be increased, it simply
changes the max and the L1 pointer size to compensate, the kernel will
fault in whichever pages are necessary.
Decreasing the size is a little bit more complex, as we want to madvise
the resulting L1 range to ensure we don't have that memory as resident
anymore. Same heuristic but going in the opposite direction otherwise.
Tends to be the case that L1 cache increases a bit on loading screens
then backs down once in-game.
These heuristic values are exposed for increasing and decreasing because
while I think I've picked reasonable values, we will likely need some
more fine tuning over time. Kind of expert user toggles at that point.
Based on #4940 as a base which needs to be merged first.
Full tracked stats from steam as an example of where we are:
```
Total (1000 millisecond sample period):
JIT Time: 0.486630 ms/second (0.00 percent)
Signal Time: 0.065880 ms/second (0.00 percent)
SIGBUS Cnt: 38 (38.160780 per second)
SMC Cnt: 0
Softfloat Cnt: 0
FEX JIT Load: 0.004585 (cycles: 552510)
Total FEX Anon memory resident: 368 mB
JIT resident: 95 mB
OpDispatcher resident: 38 mB
Frontend resident: 8 mB
CPUBackend resident: 624 kB
Lookup cache resident: 0 (null)
Lookup L1 cache resident: 7 mB
ThreadStates resident: 460 kB
Unaccounted resident: 217 mB
```
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.
The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.
Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.
I didn't do that exercise since that can be followed up in a subsequent
PR.