ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.
We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.
Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.
This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).
But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
When MAP_FIXED is used, if it was larger than the VMA region it was
trying to fit in to, then it would overallocate, corruption memory
adjacent to the VMA region. This was due to a typo in the LiveRegion
range checking.
Fix the typo, add a unittest that tries to overallocate space. Would
assert out without this bug fix.
In the case that overlapping `MAP_FIXED` mmap functions were used, we
were incorrectly tracking the full mapped regions size as new
allocation. We instead need to track which pages have already been
previously allocated and only track those. Would behave like FEX was
running out of memory, but we were just mapping the same location many
times.
Adds a unittest to track this.
Now that our Lookup cache mutex is no longer recursive, we can safely
use a shared_mutex instead. The problem with a c++ std::shared_mutex is
that it doesn't guarantee any form of priority, so tens of thousands of
read-locks per second can cause a writer to never acquire the lock, or
take too much time.
The bad news is that C++ doesn't provide us a primitive with
write-priority, so we need to construct our own that is still compatible
with Linux futex. So this is what we do.
- Windows: Uses an SRWLock instead.
- Only way for WINE to provide us a futex fallback that priorities
write-priority without stampeding.
A couple things here, we were never returning the last searched element,
either the last or first depending on search direction.
Also the backward scan would return incorrect indexes in some cases.
Also scanning beyond its page bounds.
Additionally some minorly incorrect assertions.
Adds a new unit test that ensures that we can allocate in to every
location, and that we get the correct indexes back. Also allocated
within guarded pages to ensure it doesn't read outside the bounds.
Fixes a spurious crash in Ender Magnolia.
Just helps when an unhandled ESR occurs, it was always a case of needing
to go in to the ARM ARM to decode it which was a bit of a pain. Add a
textual representation of it.
Fixes crash that occurs in applications that use both GL and Vulkan,
Like UE5 Vulkan native games. Fixes Ender Magnolia.
The issue here is that UE5 loads libGL first, which initializes our
libGL thunks, setting its X11Manager's functions.
It then loads libvulkan, which calls our oninit constructor, which
because of the symbol conflict, calls in to the libGL thunk's host
functions to reinitialize its function pointers, never initializing the
Vulkan X11Manager's functions. It would then crash as soon as an X11
function was used.
Give them unique symbol names so we don't accidentally look up the
incorrect symbol.