This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.
In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.
This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.
With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.
Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
This patch does cover up all v4l2 syscalls in i386. It just works
well with the poor v4l2 driver of wine32.
Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.
In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.
Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.
This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.
The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>) = 0
<...>
41227 <... mmap resumed>) = 0xba87b000
```
While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```
The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.
This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.
The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.
Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
`llseek` returns only ever 0 or errno in the return register. This is in
contrast to `lseek` which returns the result (or errno) in the return
register.
We were accidentally returning the result on non-error conditions which
could freak out some software. Thanks to
[OFFTKP](https://github.com/OFFTKP) for pointing out this issue
Since this syscall doesn't exist, we need to convert it to the
equivalent utimensat like the kernel does internally.
This is fairly trivial but there are some safety nets in place.
Just four new *at variants of the xattr syscalls.
This will also let us use the *at variants for the non-at versions but I
didn't implement that optimization because this is brand new.
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.
The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.
Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.
This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.
To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.
As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
The pointer is allowed to be null or garbage if the number of nfds
passed in is zero.
Unify the x86 and x86-64 implementations to ensure consistency.
Fixes Warhammer 40k: Relics of War
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.
The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.
There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.
FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.
While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.
**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour
**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.
**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this
Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
- This means we serialize and deserialize the calling thread's
seccomp filters
- An execve that escapes FEX will also escape seccomp. Not much we
can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
synchronize seccomp filters after the fact
Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
- Just like our syscall handler, we are hardcoded to the arch that the
application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
- Currently we must ensure all syscalls go through the frontend
syscall handler
- Runtime invalidation of code cache with inline syscalls will get
fixed in the future.
This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.
Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.