This matches the systemd path, no more _32 and _64 versions, just
binfmt_misc.
Having a mixed install where one program does one architecture and
another is weird and unsupported anyway.
There is a ELF note for this but currently clang doesn't support a
`no-gcs` flag. The best we can do is check if the kernel has GCS enabled
for the current process and early exit.
Then continue to use the kernel's locking functionality to disable it if
the guest happens to try, ensuring safety.
Avoids the non-fatal error: "[ERROR] Close closing FEX FD 2"
that happens when the guest program executes a syscall to close fd 2,
and it's not in fex tracked set.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.
We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
If any `sysconf(_SC_PAGESIZE);` errors then we can get bad values, make
sure to at minimum use the x86 page size.
Also changes a hardcoded page size to use the FEX pagesize define.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
This was causing FEXServer to look in to global installed paths and
local paths for things when FEXServer was started.
Ensure it listens to FEX_PORTABLE so this doesn't occur.
This also requires us to scan both data directories and config
directories to find them.
For some reason steamwebhelper is setting a zero length environment
variable. This was causing an assert to be raised early as the web
helper was starting up.
Just stop trying to memcpy the zero length string, gets steamwebhelper
working in the steam beta client again
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.
This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.
Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.
This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.
On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.
Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.
This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.
Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.
The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.
Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.
This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
Some early FEXServer startup log failures weren't getting printed
correctly. They were going through the LogManager but before FEXServer
setup, or even stderr/stdout logman setup. So they were just getting
written to -1 and failing.
Fixes#4155
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.
- Instead of using RunUntilExit, all threads can use `ExecuteThread`
directly, since there's nothing special about the primary thread now.
- This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
`ExecuteThread` has returned.
- Just make gdbserver is cleaned up early if it exists since it may
want to send some things to the connected gdb instance before
threads are exited.
FEXCore hasn't been returning anything other than EXIT_SHUTDOWN for a
long time, so this ended up just moving data around for no reason.
This isn't going to be used for further GdbServer work anyway, so just
completely remove it.
This is a Linux construct, move it to the frontend.
This is going to need some changes in the future since exit_group and
exit syscalls are supposed to behave differently than how FEX implements
it. For now just move it to the frontend.
This was brought up by #3831 but I finally got the courage to look at
the hard problem.
Although I'm only tackling half of the problem with this PR, which is
that FEXLoader needs to strip the rootfs path from the executed path if
it begins with the rootfs, plus some changes to the surrounding code.
The primary concern here is that when an application has been executed
under FEX, specifically through binfmt_misc, then FEX needs to prepend
the full rootfs path otherwise Linux can't find the program.
Additionally execveat with an FD will resolve a full path to the rootfs.
So past FEX's initial setup, we need to strip off the rootfs path to
provide an "absolute" path that is visible to the guest application
later. Which is kind of funny since we have a `RootFSRedirect` function
which did the exact opposite. This was due to legacy problems in the
original ELFLoader that couldn't handle symlinks correctly, which has
since been resolved, so that no longer needs to exist.
There was also some weirdness in `GetApplicationNames` where the passed
in argument list was modifying Args[0] and then saving the Program as
well. Which I just got rid of. Also stopped passing in the arguments by
value because....why did I write it like that?
In InterpreterHandler we now need to check if we can open the path
inside the rootfs or fallback without it. Plus I had to change the
shebang handling so it stopped prefixing the rootfs AGAIN. Took the time
to change the shebang handling there so it stops creating string copies
and instead just generates views.
Overall this fixes a fairly major flaw with how we were representing
`/proc/self` to the application, which was breaking wine since it would
prefix the rootfs multiple times, which was weird.
It doesn't address the remaining problem in #3831, which is that
applications can still see some of the leaky abstractions with symlinks
through the rootfs, but I want to get at least this step in.
The size of the input_event struct differs between 32 bits applications
and 64 bits applications. To deal with this, the kernel implements a
compat variant for the input syscalls, but it's only enabled for 32 bit
processes.
In libkrunfw we're introducing a prctl that enables a 64 bit process to
request the kernel to enable the compat variant for the input syscalls.
This commit makes use of that interface for enabling/disabling the
compat input variant as required.
The visible effect is that input devices such as gamepads work properly
on emulated 32 bit applications.
Signed-off-by: Sergio Lopez <slp@redhat.com>
This is an interesting vdso implementation because it only exists in
x86-64 with kernel v6.11. For Arm64 the implementation is likely to land
in a couple kernel versions from now.
Some differences with vdso_getrandom versus the regular getrandom
syscall
- Has two additional arguments, opaque_state and opaque_len
- Expects the userspace to mmap/munmap this opaque state structure
- Lets userspace query information about this structure upfront
- If the opaque data structure isn't provided then it falls back to
regular SYS_getrandom
With the "glibc" implementation, we can tell the interface to allocate a
single page that gets unused (otherwise glibc ends up not using the
interface), and fallback to the regular SYS_getrandom.
In the case of the vdso interface, currently the arguments just get
passed through.
Keeping this as a WIP until the ARM64 kernel patches land and I can
actually test the things. Running the "glibc" path works today with the
selftest, but the vdso path is currently completely untested.
From prep commit 511103ee5dcc9474b0b7468c05f61ce10fea4393.
Now that the frontend is setup, we can remove the temporary TLS
variables and use the ThreadObject directly.
With the new systemd support, we had forgotten to wrap this check in an
arm64 check. I had done an install and broke my x86 machine.
Add the check back so we stop breaking x86 machines...again. Took me
only about three hours to find the issue this time.
Since the frontend has changed to informing the backend if AVX is
supported, there is no reason to feed that configuration back in to
SignalDelegator from the backend.
Instead inform the SignalDelegator directly in the frontend instead of
this now weird round-about path.
Upstream has rejected this flag which would have let FEX close the gap
in functionality between binfmt_misc interpreters and PT_INTERP
interpreters for how `/proc/exe` is handled. Since we are unable to
change their opinions, just remove the code from FEX since it's never
going to happen.
Removes a little bit of confusing behaviour in execveat and
filemanagement handling. Leaving us with only the regular confusing
behaviour of trying to track accesses to `/proc/exe` instead.