Apparently Chromium/CEF can chroot or otherwise sandbox the filesystem
away before forking and checking for directory FDs, making /proc
inaccessible, which means we can't stat it for our inode check, breaking
the hiding.
So, double down on things and do what Chromium does: open an fd to /proc
ahead of time, so that continues to work. Then we use it to update the
inode of our RootFS fd instead, and finally, also do the /proc fd itself
to hide that one too.
We also don't need to check the st_dev of /proc more than once, since
that's not expected to change anyway.
Fixes cefsimple.
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.
To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.
As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
From prep commit 511103ee5dcc9474b0b7468c05f61ce10fea4393.
Now that the frontend is setup, we can remove the temporary TLS
variables and use the ThreadObject directly.
We were creating a copy of the FEXCore::Core::CPUState object when we
didn't need to. We can pass the host thread's CPUState frame through to
the creation handlers since it's read-only (so modify it to be const).
We then just move the RAX and RSP setting to /after/ the CreateThread
handling instead of before.
This reduces stack usage from ~1392 bytes to ~80 bytes.
Since the frontend has changed to informing the backend if AVX is
supported, there is no reason to feed that configuration back in to
SignalDelegator from the backend.
Instead inform the SignalDelegator directly in the frontend instead of
this now weird round-about path.
Upstream has rejected this flag which would have let FEX close the gap
in functionality between binfmt_misc interpreters and PT_INTERP
interpreters for how `/proc/exe` is handled. Since we are unable to
change their opinions, just remove the code from FEX since it's never
going to happen.
Removes a little bit of confusing behaviour in execveat and
filemanagement handling. Leaving us with only the regular confusing
behaviour of trying to track accesses to `/proc/exe` instead.
No functional change here.
- CoreRunningMode enum and variable wasn't used anymore.
- Code was moved to the frontend
- CustomCPUFactory wasn't used anymore
- All special signal handling and various features were moved to
TestHarnessRunner
- We also don't want to support actual custom CPU cores.
- TestHarnessRunner just runs as a host runner if compiled on an
x86-64 device if vixl sim isn't enabled now.
- Removes the Core config option entirely.
- Moves VDSOPointers struct to the frontend
- Every use of this lives in the Linux frontend instead now
The pointer is allowed to be null or garbage if the number of nfds
passed in is zero.
Unify the x86 and x86-64 implementations to ensure consistency.
Fixes Warhammer 40k: Relics of War
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.
The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.
There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.
FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.
While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.
**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour
**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.
**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this
Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
- This means we serialize and deserialize the calling thread's
seccomp filters
- An execve that escapes FEX will also escape seccomp. Not much we
can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
synchronize seccomp filters after the fact
Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
- Just like our syscall handler, we are hardcoded to the arch that the
application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
- Currently we must ensure all syscalls go through the frontend
syscall handler
- Runtime invalidation of code cache with inline syscalls will get
fixed in the future.
This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.
Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
Two primary things here:
- Remove the static `GlobalDelegator`
- Move the thread_local SignalDelegator::ThreadState information
directly in to ThreadStateObject
Having the ThreadStateObject and the SignalDelegator information
disjoint was confusing but was required when we didn't have any object
in the frontend that could have its own independent data. Since we fixed
this with the `ThreadStateObject` type we can now move this over.
The `GlobalDelegator` object is now instead stored in
`ThreadStateObject` instead.
Instead of using a thread_local variable, we now just consume 8-bytes of
the signal alt-stack since the kernel gives us that information about
where it lives. This then converts all the thread_local usage to use
either the passed in CPU state if it exists, or fetching it from the
alt-stack offset.
Very minor changes in behaviour here, will help when trying to improve
FEX's behaviour around signals.
These are a Linux construct and should live here. Removes a weird
passthrough API from FEXCore and keeps it in the frontend instead.
This isn't even typically allocated in a real setup, as it's only a
fallback for if VDSO isn't loaded.
The CallbackReturn function stays in FEXCore because it would have
caused an API in the other direction instead.