I forgot on Linux by default we didn't have the syscall instructions
count as block end. Change this so that it counts as block end now.
This has the additional benefit that now the frontend needs to modify
the RIP manually as well which is fine as it's what arm64ec and wow64
does.
Also add back the UnimplementedOp in RDPID that accidentally got caught
up. Also increment DiskCache version as both changes will change
codegen.
Fixes#5942
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.
Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.
This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.
Bumps the DiskCache version again because it causes codegen to change.
building using musl fails due to missing defintions:
```
Source/Tools/LinuxEmulation/LinuxSyscalls/Syscalls.cpp:913:23: error: no member named 'sleep_for' in namespace 'std::this_thread'
913 | std::this_thread::sleep_for(std::chrono::milliseconds {10});
| ^~~~~~~~~
1 error generated.
```
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
When an anonymous FD is passed to execveat that has the CLOEXEC flag
set, then binfmt_misc fails with ENOENT.
Workaround this limitation by duplicating the FD, stripping its CLOEXEC
flag in the process.
For #5234
FEX_CONFIG_OPT can only be used as a standalone statement, which is
inconvenient for config values that are only used once. The new functions
(e.g. Get_DUMPIR()) can be used in conditions or other expressions.
Allows applications running entirely under Linux to use the same
extended volatile metadata as Windows.
For example `FEX_EXTENDEDVOLATILEMETADATA=iw4sp.exe\;0xe9da0-0xe9ec7`
this configuration works for both wow64 and Linux to disable the TSO
emulation on the memcpy routine in that game that consumes around 80% of
CPU time in TSO emulation.
Works with Linux native games as well of course.
Still creates a copy of FEXInterpreter from FEX for downstream projects
to have some time to get off the old name. Creating a symlink is kind of
a pain in cmake so just doing an install copy is easy.
This matches kernel behaviour for how brk works. I initially implemented
this as a way to speed up some early applications that relied heavily on
brk. We're long past that time and we should match Linux behaviour
instead.
This fixes issues where applications (FEXLoader really) can see we've
allocated a larger granule and breaks self-hosting. There's no reason to
be smarter around this, if applications want more perf then they'll
switch to a better allocator than brk.
Also need to make sure that after the brk region has been reserved, to
unmap its initial mapping to allow future mmap syscalls to overwrite it.
We need to reserve it early to ensure the region initially exists and we
don't accidentally map other things in that space. Then it is up to the
guest if they don't want to overwrite it.
Implied the maximum size that brk could grow to. Which isn't correct, it
is the current maximum size mapped which is page size, versus the
current brk offset which is byte ranged.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
Make sure to pass the clone3 arguments all the way to the fork handler
so it can check the flags. Currently nothing I know of uses fork plus
the new clone3 flags, but it would be hard to see without any logging.
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.
The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.
Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.
This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
- Parse the shebang line properly (use FHU::ParseArgumentsFromString
which is the same code the loader uses)
- Make native-interpreter shebang files work by deferring to the kernel
in that case (previously, they'd get executed through the loader and
it would choke on the architecture of the interpreter)
- Do not use the RootFS-prepended path when executing shebang files. The
loader will prepend that anyway when looking it up, but it needs the
bare guest path so it can pass it as an argument to the interpreter,
which (since it's emulated) will do the lookup through the RootFS.
With a merged RootFS, all binaries are executed through the RootFS. When
executing a binary that is actually a native binary, we want to do so
outside the RootFS. Handle this by stripping the RootFS prefix in that
case.
clone3 was added in Linux 5.3 but our minimum spec is 5.0. Additionally
the Raspberry Pi 5 kernel seems to complain about clone3 for some
reason?
Just use clone instead of clone3
We were using this variable for two things, letting the frontend signal
to the backend that it wants to start executing once the thread is
created, and also for handling thread pausing. These two features are
conflated with one another and actually makes things more confusing.
- Move StartRunning/StartPaused to the frontend, because its a construct
that only needs to exist in the frontend
- Adds a FEX::HLE::ThreadStateObject CV for handling pausing, which only
needs to exist for gdbserver
Needed by Discord, part of the Chromium sandbox code. The warning still
triggers because Chromium asks for CLONE_VM on x86_64, but that can be
safely ignored (CLONE_FS is the one that matters).
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.
To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.
As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
From prep commit 511103ee5dcc9474b0b7468c05f61ce10fea4393.
Now that the frontend is setup, we can remove the temporary TLS
variables and use the ThreadObject directly.
Upstream has rejected this flag which would have let FEX close the gap
in functionality between binfmt_misc interpreters and PT_INTERP
interpreters for how `/proc/exe` is handled. Since we are unable to
change their opinions, just remove the code from FEX since it's never
going to happen.
Removes a little bit of confusing behaviour in execveat and
filemanagement handling. Leaving us with only the regular confusing
behaviour of trying to track accesses to `/proc/exe` instead.
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.
The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.
There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.
FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.
While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.
**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour
**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.
**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this
Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
- This means we serialize and deserialize the calling thread's
seccomp filters
- An execve that escapes FEX will also escape seccomp. Not much we
can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
synchronize seccomp filters after the fact
Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
- Just like our syscall handler, we are hardcoded to the arch that the
application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
- Currently we must ensure all syscalls go through the frontend
syscall handler
- Runtime invalidation of code cache with inline syscalls will get
fixed in the future.
This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.
Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.
Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
Pulled from the seccomp WIP PR where it pulls this object more frequently.
Since it is an opaque frontend pointer it needs to be cast and we
already have a few locations that use it.
No functional change.