Fixes#2136
This is a fairly tricky edge case to support with FEX.
If execveat is used with AT_EMPTY_PATH then the application can pass an
FD to execve instead of a filename. This includes FDs that have been
deleted from the disk so the child process can't open it by filename
anymore.
To work around this limitation, we need to pass the FD to the new FEX
process and open it directly, similar to how binfmt_misc works with FDs.
The FD will get passed through environment variables, which the new
process will check for and then remove the variable from the
environment.
Lots of prickly edge cases to support here.
Without binfmt_misc:
- Passes the FD to FEXLoader directly.
- Requires duplicating the FD if it has O_CLOEXEC on the FD.
With binfmt_misc:
- Shebang file, pass directly to FEXLoader, just like without binfmt.
- x86 ELF Files, rely on the kernel's binfmt_misc support here.
- Unsupported ELF files, let kernel handle it through binfmt_misc
Argument handling:
- The application can pass in no arguments.
- Means our application configurations were failing to find a config
- Also various checks in the frontend were failing.
- If opened through an FD, find the symlink for that FD for the
application configuration instead.
Side note:
Fixed a performance issue in execve where when we were checking for file
format support. Either ELF or Shebang files, we were reading the /whole/
file upfront. We only need to read a header worth of ELF files, and only
257 bytes if it is potentially a shebang file. Should dramatically
reduce some application's execve times.
Segment registers are indexed significantly more than they are changed.
Pay the cost of indexing during the set and store rather than the per
register index.
Should be a fairly significant performance improvement for 32-bit
applications. At least on hardware that doesn't have a data dependent
prefetcher.
Breaks Steam atm and isn't clean.
We weren't adjusting the guest stack size when using clone3.
If the clone comes from CLONE3 then we need to offset the RSP by the
provided stack size.
This also translates to fork/vfork through clone, so make sure to adjust
stack in that case as well.
Fixes Ender Lilies crashing with Ubuntu 22.04 rootfs
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.
So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.
Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
Syscall entry points still have different argument orders,
Moves the arguments to the clone3 argument structure and passes to generic handler.
Also implements clone3 while doing this
Notably this allows applications to work that don't require the namespace but check up front if they are able to clone with it.
Civ 6's launcher checks this as an example
The vast majoirty of syscalls don't need anything in thread or frame.
So lets save an indirection for all those syscalls.
Most of the syscalls which do need Thread (or CTX via
Thread are in Thread.cpp or Memory.cpp
These have all been modifiy to fetch Thread from Frame
We need to check to see if the application wanting to be executed is in
the rootfs first. Otherwise we end up in a case where we either lose
tracking of applications in FEX, or applications fail to launch since
they don't exist in the global filesystem
Sadly these things can't be split without breaking functionality so it
turns in to a bit of a mess.
SyscallHandler is very much something that is a Linux only construct and
shouldn't be in FEXCore itself. Lets the frontend register a
Syscallhandler with FEXCore. FEXCore itself is then aware of the current
syscall ABI and handles the ABI in an optimal fashion.
So it is not a 100% clean break otherwise we would lose performance.
The SignalDelegator then needs to move to the frontend since the
SyscallHandler requires it for signal based syscalls.
The CPU backend signal handling still needs to happen in FEXCore because
it is a very tight coupling with the CPU backend.
Once we need to support more Signal handling we can give the backends
cleaner support to select which specific OS handler to handle.