In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)
However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.
One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.
Additionally some of the metadata generated around the signal delegation
wasn't correct.
With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.
I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.
With this change, Elden Ring is now stabilized and works under FEX.
Allows for a smaller struct size.
This movement of the struct members also requires us to modify the
DeadContextStorePass to take the new locations into account.
On the plus side, we get to remove all of the padding bits in the
CPUState struct, so we can remove the handling for them.
The recursive algorithm used here previously led to deeply nested function
calls, which eventually exhausted the available stack space. The simple
non-recursive algorithm used now avoids this problem at the expense of
small overhead.
Gets rid of some magic numbers and reduces the number of things that
need to manually change (e.g. when supporting AVX and needing to
increase the xmm size).
This is an IR op that produces nor consumes any SSA values, but has side
effects.
Turns in to the pause instruction on x86 and yield instruction on
AArch64.
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.
This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.
Apple M1 supports this extension, didn't test with Cortex-X2/A710.
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.
A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.
Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
these stale allocations and reclaim it from the idling thread. Saving
further memory.
This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.
Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.
This makes the compile service never be invoked so we can just remove
it.
We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.
This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.
This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.
Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
This matches the AArch64 implementation fairly well.
Bundles RDRAND and RDSEED together for simplification, both instructions
are a single flag on AArch64.
This greatly simplifies the IR format by using string parsing for
gathering the information.
Tons of redundant information removed.
Significantly more difficult to mess up adding a new IR op.
Significantly improves the generator functions in IREmitter
In some cases we can generate more optimal code if we have more
information about a syscall which number gets const-propagated.
In particular optimizing through syscalls, not synchronizing state, and
never returning.
- Noreturn is used by a syscall that never returns, like exit.
This means that it never needs to try and synchronize state coming back
- Not synchronizing state and optimizing through syscalls
Useful for syscalls that don't read the state past arguments and only
returns a value.