In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)
However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
In most cases we aren't using JIT symbols but still creating the perf
map file.
Early check if we should generate the file or not, this way we stop
polluting the /tmp folder.
glibc lazily initializes the SETXID signal handler until first thread
creation.
Once we create our first pthread, steal it back from GLIBC after the
fact.
Fixes SOMA again.
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.
One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.
Additionally some of the metadata generated around the signal delegation
wasn't correct.
With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.
I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.
With this change, Elden Ring is now stabilized and works under FEX.
Now that we have the AVX config option in place, we can use it to
determine how to set up the stack for storing and loading all of the
necessary state we need.
If we have SVE2 support and a vector length of at least 256, then we can
adequately handle AVX.
However we also add in the ability for application profiles to disable
AVX if necessary for any reason.
Not currently used, but will be in subsequent changes.
On Linux the bottom 12 bytes of the legacy FXSAVE context are used to
encode information about any extension blocks that may follow it.
We can use the same approach to encode our high lanes of the AVX
YMM registers into memory and vice-versa.
If an application is doing long jumps on exceptions then we need to
ensure that RIP is synchronized at least to block entry for some amount
of safety.
Due to our block linking which doesn't ensure that RIP is synchronized
on block entry, our exception handling wouldn't see a RIP change, thus
not going down the path that we jump to Dispatcher loop top on RIP
change.
Burn a couple of instructions on block entry to ensure that RIP
synchronized but only on config.
When some SRA code was being shuffled around, this config variable was
never being set. This caused the guest signal handlers to never expect
SRA to be enabled.
Moves the Dispatcher config object to the actual dispatcher rather than
having it in the arch specific dispatcher class. Then have the signal
delegator check the config directly instead of having this secondary
config value.
Fixes vc_redist_x64 and vc_redist_x86. Probably also fixes a bunch of
other random Proton/Wine things that rely on signal long jump.