Commit Graph
57 Commits
Author SHA1 Message Date
Ryan Houdek 66cad978c3 FEXCore: Removes syscall optimization
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.

Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.

This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.

Bumps the DiskCache version again because it causes codegen to change.
2026-08-31 19:23:18 -07:00
LC a3d0b67777 LinuxEmulation: Resolve missing prototype warnings
Ensures all functions are marked whether they're intended to be
internally linked or not.
2026-07-17 00:12:28 -04:00
LC 2ad6254894 x64/Thread: Amend faulting copy handling related to LDTs
CopyToUser doesn't return the number of bytes copied, but rather returns
0 to indicate success, otherwise a fault has occurred (and the SIGSEGV
handler has set X0 to EFAULT)
2026-07-11 11:57:38 -04:00
Ryan Houdek d2f096187c Merge pull request #5705 from lioncash/thread3 2026-07-10 22:34:04 -07:00
LC 8b0f07b2dc Syscalls: Remove unnecessary usages of namespace FEXCore::IR
These aren't necessary.
2026-07-10 21:02:12 -04:00
LC 017c898ed0 x64/Signals: Amend set size in rt_sigtimedwait
We should be checking the size passed in, not the sizeof of it.
2026-07-10 19:40:05 -04:00
Ryan Houdek 7e8aa711ef Linux: Pass gettimeofday through glibc
This ensures that it hits the VDSO path if possible.
2026-06-03 17:45:41 -07:00
Gio 486c8805ed Syscalls: fix TraceFormatString for DEBUG_STRACE
Replaces the % strings with {}
2025-12-06 14:00:44 +08:00
Tony Wasserka 6403da3715 LinuxSyscalls: Shield code map FD from guest access
This prevents chromium/CEF from closing the FD.
2025-11-20 19:13:18 +01:00
Ryan Houdek 474f2dc267 FEX: Name remaining allocations as "Misc"
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.

With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
        Misc resident:        54 MiB
    JEMalloc resident:        208 MiB
```

So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
2025-10-10 17:34:03 -07:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Lioncache 10c3b2fb5d SyscallHandler: Shrink definition struct from 24 to 16 bytes
Just a minor space saving, given how many syscall definitions we'll
have.
2025-10-08 16:40:20 -04:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Ryan Houdek 4516e4e9f6 LinuxSyscalls: Support modify_ldt on x64
This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.

In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.

This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.

With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
2025-08-20 13:22:56 -07:00
Tony Wasserka fbef5c6a6a Syscalls: Fix DEBUG_STRACE build 2025-08-05 15:51:34 +02:00
Tony Wasserka df461546c5 LinuxSyscalls: Make error return values consistent 2025-06-03 11:10:08 +02:00
Tony Wasserka 1c1c43cd86 LinuxSyscalls: Fix formatting 2025-06-03 11:10:08 +02:00
Tony Wasserka 3eac9f937e LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:08 +02:00
Tony Wasserka f13a0d8e84 LinuxSyscalls: Unify shmdt implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ca12dc9213 LinuxSyscalls: Unify shmat implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 365ed2cd70 LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 7efbfed0bf LinuxSyscalls: Unify mmap and munmap implementations 2025-06-03 11:10:07 +02:00
Ryan Houdek ad132267ec Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.

This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.

The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>)             = 0
<...>
41227 <... mmap resumed>)               = 0xba87b000
```

While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```

The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.

This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.

The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.

Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
2025-05-29 12:15:26 -07:00
Ryan Houdek 68939c5a5c LinuxSyscalls: Emulate futimesat syscalls
Since this syscall doesn't exist, we need to convert it to the
equivalent utimensat like the kernel does internally.

This is fairly trivial but there are some safety nets in place.
2025-03-29 11:06:51 -07:00
Lioncache 218b0d491a LinuxSyscalls: Replace use of deprecated std::is_trivial template
This type trait is deprecated in C++26, so we can just be more specific.
2025-03-29 00:53:30 -04:00
Ryan Houdek b0b41d00ee Various: More static analysis warnings cleanup
NFC
2025-03-12 17:27:41 -07:00
Ryan Houdek 031afbfb18 Various: Adds missing SPDX file headers
NFC
2025-03-07 11:34:31 -08:00
Ryan Houdek fca4c7e6bf LinuxSyscalls: Update for new v6.13 syscalls
Just four new *at variants of the xattr syscalls.
This will also let us use the *at variants for the non-at versions but I
didn't implement that optimization because this is brand new.
2025-01-19 18:41:30 -08:00
Ryan Houdek 4cfb81156f FEXLoader: Increase minimum kernel requirement from 5.0 to 5.15
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.

The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.

Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.

This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
2025-01-01 11:22:54 -08:00
Ryan Houdek dd8a3a9aea LinuxEmulation: Don't use clone3 for fork
clone3 was added in Linux 5.3 but our minimum spec is 5.0. Additionally
the Raspberry Pi 5 kernel seems to complain about clone3 for some
reason?

Just use clone instead of clone3
2024-12-05 15:14:37 -08:00
Asahi Lina 3d701f5fcf FileManagement: Hide the FEX RootFS fd from /proc/self/fd
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.

To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.

As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
2024-10-27 07:05:00 +09:00
Ryan Houdek eeff198ff1 LinuxEmulation: Update syscalls for v6.11
Only thing added was uretprobe on x86-64. We can't support this so just
return ENOSYS.
2024-10-02 14:40:11 -07:00
Ryan Houdek 611aa0a5b9 Thunks: Wire up TLS handling in the frontend
Because it was moved from the backend to the frontend, re-wire up the
TLS handling which requires going through our syscall handler.
2024-09-27 15:50:21 -07:00
Ryan Houdek f6a8e595df Linux/Syscalls: Add todo for futimesat 2024-09-08 19:34:38 -07:00
Ryan Houdek 8ea4e4f663 Linux: Adds some missing brace initializers on conversion operators 2024-09-08 18:03:20 -07:00
Ryan Houdek 64c5362580 LinuxSyscalls: With poll syscall, ensure fds is writable only if nfds is not zero
The pointer is allowed to be null or garbage if the number of nfds
passed in is zero.
Unify the x86 and x86-64 implementations to ensure consistency.

Fixes Warhammer 40k: Relics of War
2024-09-05 14:33:52 -07:00
Ryan Houdek ac32876e4e LinuxEmulation: Implement support for seccomp
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.

The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.

There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.

FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.

While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.

**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour

**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.

**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this

Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
   - This means we serialize and deserialize the calling thread's
     seccomp filters
   - An execve that escapes FEX will also escape seccomp. Not much we
     can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
  synchronize seccomp filters after the fact

Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
  - Just like our syscall handler, we are hardcoded to the arch that the
    application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
  - Currently we must ensure all syscalls go through the frontend
    syscall handler
  - Runtime invalidation of code cache with inline syscalls will get
    fixed in the future.

This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.

Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
2024-09-02 14:07:53 -07:00
Ryan Houdek 9056d9b9de SignalDelegator: Refactor how thread local data is stored
Two primary things here:
- Remove the static `GlobalDelegator`
- Move the thread_local SignalDelegator::ThreadState information
  directly in to ThreadStateObject

Having the ThreadStateObject and the SignalDelegator information
disjoint was confusing but was required when we didn't have any object
in the frontend that could have its own independent data. Since we fixed
this with the `ThreadStateObject` type we can now move this over.

The `GlobalDelegator` object is now instead stored in
`ThreadStateObject` instead.

Instead of using a thread_local variable, we now just consume 8-bytes of
the signal alt-stack since the kernel gives us that information about
where it lives. This then converts all the thread_local usage to use
either the passed in CPU state if it exists, or fetching it from the
alt-stack offset.

Very minor changes in behaviour here, will help when trying to improve
FEX's behaviour around signals.
2024-09-02 06:46:19 -07:00
Ryan Houdek 3a79b61f58 x64/Thread: Adds EFAULT handlers for two syscalls
For the remaining syscalls we need to introduce some new concepts around
checking if pointers and string arrays are readable and I don't want to
do that yet.
2024-08-28 04:14:16 -07:00
Ryan Houdek 129ec63e92 x64/Time: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 3459369c6e x64/Signals: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 7a490a3811 x64/Semaphore: Add EFAULT checks 2024-08-27 03:04:23 -07:00
Ryan Houdek ee17fe239a x64/EPoll: Add EFAULT checks 2024-08-27 03:04:23 -07:00
Ryan Houdek ce88f5f948 Merge pull request #3941 from Sonicadvance1/fault_assertion_checking
LinuxSyscalls: Implements less invasive assertion only EFAULT handlers
2024-08-15 01:14:15 -07:00
Ryan Houdek 2b26d0fff5 LinuxSyscalls: Implements less invasive assertion only EFAULT handlers
With the previous Copy{To,From}User helpers we need to actually
implement the handlers correctly. We want something that is a bit
lighter so we don't need to implement the faulting path in the syscall
handlers.

Implements a handful of helpers that just check for readable and
writable capability which can be thrown in to an assertion handler that
is zero cost in release mode.

Readable is checked by just attempting to read all bytes.
Writable is checked by attempting to read each byte and writing it back
to the same location.

Uses these helpers in x64/FD.cpp to showcase how they will be used to
detect EFAULT. Tested locally that they work correctly by writing some
small tests for the syscalls that expect EFAULT.
2024-08-15 00:48:57 -07:00
Ryan Houdek 924723d433 LinuxSyscalls: Some minor cleanups
- We can have the SyscallFunctionDefinitions be the correct size out of
  the gate. Both tables are always 512 entries in size.
- In the RegisterSyscall_{32,64} handlers, just get the reference using
  operator[]. We always know we will be under the size of the array, add
  a an assert to check. Removes a bit of vector range checking overhead.
- Namespace 32-bit syscalls like 64-bit syscalls and include in the
  regular header like 64-bit. This was just an oversight
- Use std::fill for the syscall gap for the invalid syscall, just a
  minor cleanup.

No functional change.
2024-08-13 14:06:16 -07:00
Ryan Houdek 67663b812e Linux: Update syscall defines for v6.10 2024-07-30 19:01:52 -07:00
Ryan Houdek 3b2e657fd4 FEXCore: Removes ThreadManager
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.

Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
2024-07-25 14:54:10 -07:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00