This is necessary for the fexserver to function correctly when chrooting
in to our rootfs and doing things.
Requires independent rootfs script modifications which will come with
the next rootfs update.
Problem comes down to a chroot supporting multiple users, where our
typical use case is only one user. Bind the server file to a single
server for the entire chroot session regardless of users, solving this
problem inside the chroot.
Fixes apt-get inside of chroot, which runs as user _apt.
This changes how host trampolines for guest functions are created. Instead
of doing this purely on the guest-side, it's either the host-side that
creates them in a single step *or* a cooperative two-step initialization
process must be used. In the latter, trampolines are allocated and partially
initialized on the guest and must be finalized on the host before use.
This function must be able to handle both guest heap pointers *and* host heap
pointers, so it only forwards to the native host library for the latter.
This is because Xlibint users allocate memory using internal macros aliasing
to libc's malloc but then they free using the function XFree. For libX11,
this is not a problem since the allocation happens in a thunked API function
(and hence on the host heap), but if a function from an unthunked library
accesses Xlibint, it will allocate on the guest heap.
One notable example where this was encountered is XF86VidModeGetAllModeLines.
In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)
However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
In most cases we aren't using JIT symbols but still creating the perf
map file.
Early check if we should generate the file or not, this way we stop
polluting the /tmp folder.
glibc lazily initializes the SETXID signal handler until first thread
creation.
Once we create our first pthread, steal it back from GLIBC after the
fact.
Fixes SOMA again.
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.
One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.
Additionally some of the metadata generated around the signal delegation
wasn't correct.
With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.
I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.
With this change, Elden Ring is now stabilized and works under FEX.
Allows for a smaller struct size.
This movement of the struct members also requires us to modify the
DeadContextStorePass to take the new locations into account.
On the plus side, we get to remove all of the padding bits in the
CPUState struct, so we can remove the handling for them.
Now that we have the AVX config option in place, we can use it to
determine how to set up the stack for storing and loading all of the
necessary state we need.
If we have SVE2 support and a vector length of at least 256, then we can
adequately handle AVX.
However we also add in the ability for application profiles to disable
AVX if necessary for any reason.
Not currently used, but will be in subsequent changes.
On Linux the bottom 12 bytes of the legacy FXSAVE context are used to
encode information about any extension blocks that may follow it.
We can use the same approach to encode our high lanes of the AVX
YMM registers into memory and vice-versa.
If an application is doing long jumps on exceptions then we need to
ensure that RIP is synchronized at least to block entry for some amount
of safety.
Due to our block linking which doesn't ensure that RIP is synchronized
on block entry, our exception handling wouldn't see a RIP change, thus
not going down the path that we jump to Dispatcher loop top on RIP
change.
Burn a couple of instructions on block entry to ensure that RIP
synchronized but only on config.
When some SRA code was being shuffled around, this config variable was
never being set. This caused the guest signal handlers to never expect
SRA to be enabled.
Moves the Dispatcher config object to the actual dispatcher rather than
having it in the arch specific dispatcher class. Then have the signal
delegator check the config directly instead of having this secondary
config value.
Fixes vc_redist_x64 and vc_redist_x86. Probably also fixes a bunch of
other random Proton/Wine things that rely on signal long jump.