When FEXServer is daemonizing through an instance of FEXLoader or
FEXInterpreter, it would leave a zombie process which was waiting for us
to read the process status.
Since we don't care about the child status and don't want to get blocked
by waitpid, just ignore the signal.
This tells the kernel that we don't care about the signal and will kill
the zombie process immediately.
Didn't notice this before since FEXServer started failing to daemonize.
Fixes#2443
I found out with some profiling that this we were spending a decent
amount of time with the `openat` syscall in heavily utilized situations.
While not super common in active gameplay situations, it matters
significantly in loading screens that this is fairly optimal.
The bulk of the time is spent in the emulated files handler to ensure
that whatever path we are given, we can capture file paths that we need
to emulate. The largest contributor being the std::filesystem::canonical
function call.
A couple of optimizations in place here.
1) Do a quick hashmap check right at the start to see if we exactly fit
2) Change from `std::fs::canonical` to `realpath`
3) Switch `GetEmulatedFDPath` to not use optional so it stops building
on the stack
I'm still not super happy with the performance of `realpath` and also
not happy that we still need to use `lexically_normal` in one code path.
But short of writing a super hand-optimized `realpath` that fits our
constraints, I don't think we can do better.
Micro benchmark needs to test four different situations due to this
optimization.
1) Non-EmuFD path
2) Non-EmuFD path with dirfs
3) EmuFD path
4) EmuFD path with dirfs
And the performance improvement for each situation respectively
1) 12% performance improvement
- 213413 openat syscalls/s -> 238999 syscalls/s
2) 17% performance improvement
- 202085 openat syscalls/s -> 237309 syscalls/s
3) 17% performance improvement (/proc/cpuinfo)
- 56616 openat syscalls/s -> 66231 syscalls/s
- Includes overhead of generating temp FD and close syscall
4) 5% performance improvement (/proc/cpuinfo)
- 51080 openat syscalls/s -> 53956 syscalls/s
- Includes overhead of generating temp FD and close syscall
And for sake of comparison to the non-emulated system; My test system
can hit around 1-1.1 million openat syscalls per second in the same
microbench.
Nice little performance uplift.
I made the assumption from some bad historical knowledge that the kernel
will canonicalize relative filenames and symlinks for applications that
execute through execve.
This turns out to not be true. In fact it passes pathname untouched to
the interpreter. So we need to do an additional fix up on relative paths
to ensure glibc doesn't break.
Fixes a major bug that breaks a bunch of games.
Fixes#2136
This is a fairly tricky edge case to support with FEX.
If execveat is used with AT_EMPTY_PATH then the application can pass an
FD to execve instead of a filename. This includes FDs that have been
deleted from the disk so the child process can't open it by filename
anymore.
To work around this limitation, we need to pass the FD to the new FEX
process and open it directly, similar to how binfmt_misc works with FDs.
The FD will get passed through environment variables, which the new
process will check for and then remove the variable from the
environment.
Lots of prickly edge cases to support here.
Without binfmt_misc:
- Passes the FD to FEXLoader directly.
- Requires duplicating the FD if it has O_CLOEXEC on the FD.
With binfmt_misc:
- Shebang file, pass directly to FEXLoader, just like without binfmt.
- x86 ELF Files, rely on the kernel's binfmt_misc support here.
- Unsupported ELF files, let kernel handle it through binfmt_misc
Argument handling:
- The application can pass in no arguments.
- Means our application configurations were failing to find a config
- Also various checks in the frontend were failing.
- If opened through an FD, find the symlink for that FD for the
application configuration instead.
Side note:
Fixed a performance issue in execve where when we were checking for file
format support. Either ELF or Shebang files, we were reading the /whole/
file upfront. We only need to read a header worth of ELF files, and only
257 bytes if it is potentially a shebang file. Should dramatically
reduce some application's execve times.
On first FEXInterpreter execution, it is expected that `ConnectToServer`
will fail with `ECONNREFUSED` because FEXServer won't be running.
Skip printing this first messaage to stderr if configured.
If it is some other error message then ensure it is still printed.
Two changes here.
- Make the mount path follow server temp folder requirements.
- Will be mounted in `/tmp/` or `$XDG_RUNTIME_DIR/` now
- Switch the FEXServer socket to an "abstract" AF_UNIX socket.
- If the socket is in `/tmp/` then systemd will put the service in a
private `/tmp` folder that only exists for the service.
- If the socket is in `$XDG_RUNTIME_DIR` then pressure-vessel can't
chroot anymore since they make their own runtime directory.
- If it is in `$HOME/.fex-emu/` then it breaks usage where the
filesystem is a mount that doesn't support AF_UNIX like sshfs.
The only reasonable thing to do is to switch over to `abstract` sockets
which will work in all cases.
Tested with pressure-vessel and systemd and now it works in all
situations.
This will be useful for keying specific executables to steamids.
This is sadly required because a bunch of games end up naming themselves
"game.exe" so we can't safely enable thunks for all things shipping a
generic name.
In the case of a platform enabling PrivateTmp then the FEXServer and
FEXInterpreter won't have a tmp folder that shares the socket location.
The runtime directory is a perfect place to share these across
processes.
Noticed recently that `FEXServer -w` was broken and couldn't understand
why. Turns out that FHU syscall handling was /always/ falling down the
`#else` path in the handlers since cmake `add_definitions` follows
folder scoping rules.
This means it was always returning -1, which was causing FEXServer's
pidfd_open usage to always receive -1, which meant the sendmsg with FD
was always failing, which meant the `FEXServer -w` would forever wait
for a message that was never sent.
Converting the utility over to a target not only fixes definition
scoping problems, but also makes the other paths actually work.
This found some compiling bugs and instead lets us define SYS_pidfd_open
if it doesn't exist. Letting the kernel return the ENOSYS if it doesn't
exist on that platform.
Main thing, fixes FEXServer -w hanging forever.
This is necessary for the fexserver to function correctly when chrooting
in to our rootfs and doing things.
Requires independent rootfs script modifications which will come with
the next rootfs update.
Problem comes down to a chroot supporting multiple users, where our
typical use case is only one user. Bind the server file to a single
server for the entire chroot session regardless of users, solving this
problem inside the chroot.
Fixes apt-get inside of chroot, which runs as user _apt.
While we were getting the application name for the application layer, we
were failing to store the filename for telemetry.
Save the filename we get for application layers and store it for the
telemetry file.
Otherwise these were just alway ending up as wine or wine-preloader.
Fixes a bug in `ExecAndWaitForResponse` where results > 1024 bytes would
overwrite data.
Switches from a custom format txt file to a json file.
JSON file now has a "Type" field to specify squashfs versus erofs.
JSON is now versioned so we don't need to move the file around, just
append to a new versioned segment.
Only shows erofs files if you have the bleeding edge `erofsfuse`
application.
This application was available starting with erofs-utils v1.5 which was
released on 2022-06-13, so it isn't available pretty much everywhere.