Commit Graph
62 Commits
Author SHA1 Message Date
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Lioncache b2407352a9 IR: Remove unnecessary forward declarations/headers
Pares the base IR header down to a modest size and also reveals a few indirect
reliances on the allocator facilities.
2025-09-08 22:03:25 -04:00
Ryan Houdek e9f8c01e04 Merge pull request #4848 from lioncash/thread
InternalThreadState: Remove unused includes
2025-09-08 16:59:15 -07:00
Lioncache 1e03219ca3 Syscalls/Thread: Add missing forward declarations
Currently we were lucky that these were being seen already when being included.
2025-09-08 14:58:30 -04:00
Lioncache 6b2363d3f6 InternalThreadState: Remove unused headers
Removes some includes in the core header and resolves indirect includes elsewhere.
2025-09-08 14:13:01 -04:00
Lioncache fd10ce0600 CodeLoader: Return auxv results by value
Same behavior, but without the need to care about out values.
2025-09-08 12:07:14 -04:00
Ryan Houdek 4516e4e9f6 LinuxSyscalls: Support modify_ldt on x64
This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.

In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.

This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.

With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
2025-08-20 13:22:56 -07:00
Ryan Houdek 7ed9bea16b Linux: Implement enough of ptrace to allow Ubisoft launcher
Wine does some minimal attaching, peeking, poking, and detaching to have
a different process inspect another's TEB region. Ubisoft's launcher
does this to check if a debugger is attached and rejects it in the case
that it is.

While this implementation isn't all-encompassing, it is good enough for
this family of games.
2025-06-09 16:32:42 -07:00
offtkp c02b88baff Fix 32-bit fadvise64 2025-05-03 03:56:16 +03:00
Ryan Houdek c44d8eed35 LinuxEmulation: Fixes remaining memory leak on pthread teardown
With the previous stack leak fix, RUINER reduced its memory leaking down
to around 50MB/s. The remaining memory leaks come from the pthread stack
that we are required to allocate (128KB per thread) and some internal
DTV tracking structures.

The problems come in the fact that glibc/pthread only tears down its
internal state for these if the pthread function actually returns! We
can **technically** switch the initial stack over to a "user" stack but
that introduces more problems around internal dtv tracking that we
already fixed months ago, so we can't actually do that in practice.

This leaves us no choice, we effectively are mandated to return from the
pthread function in order to free the memory from glibc. The only way we
can safely do this is with a long jump and deferring some data structure
management until that case.

This is all incredibly sucky but it's necessary to work. With these
changes, RUINER is no longer leaking memory (Hovering at around 3GB used
while in-game) and even Steam is consuming less memory.

It doesn't solve the problem that thread creation and teardown could
likely be faster, but not many applications are creating 720
threads/second.
2025-04-17 18:16:32 -07:00
Ryan Houdek 86c492da13 LinuxSyscalls: Fixes a major stack memory leak
Fixes an issue where a thread that exits with the `exit` syscall never
actually frees its pivot stack. This is common practice and it was
missed when I was fixing the previous stack pivot leak.

This was uncovered when looking at the game
[RUINER](https://store.steampowered.com/agecheck/app/464060/?curator_clanid=4777282)
for timing bugs. Turns out the Linux build of the game creates and
destroys a VLC object every tick of the engine. Creating this VLC
object creates six threads behind the scenes. At what I assume the
default tick of the engine is of 120Hz(?) this would mean it is
attempting to create and destroy 720 threads per second, quickly
leading to memory exhaustion under FEX.

While this hits the biggest memory leak we have around thread creation,
this game is still hitting thread creation so hard that I can see other
leaks that I need to track down still.
2025-04-16 19:47:56 -07:00
Ryan Houdek 2ba0b66426 LinuxSyscalls: Implements support for Linux side profile stats
This is fairly straightforward. It creates the shared memory region in
/dev/shm/fex-<pid>-stats so that Mangohud can sample it.
2025-02-11 12:56:11 -08:00
Ryan Houdek 9858ab7388 LinuxEmulation: Ensure syscall wrapper declaration has CpuStateFrame as the first argument
Otherwise crashes occur.
2025-01-23 12:26:10 -08:00
Ryan Houdek fca4c7e6bf LinuxSyscalls: Update for new v6.13 syscalls
Just four new *at variants of the xattr syscalls.
This will also let us use the *at variants for the non-at versions but I
didn't implement that optimization because this is brand new.
2025-01-19 18:41:30 -08:00
Ryan Houdek 8efa5febd0 LinuxSyscalls: Log unhandled clone3 fork flags
Make sure to pass the clone3 arguments all the way to the fork handler
so it can check the flags. Currently nothing I know of uses fork plus
the new clone3 flags, but it would be hard to see without any logging.
2025-01-03 09:03:47 -08:00
Ryan Houdek 4cfb81156f FEXLoader: Increase minimum kernel requirement from 5.0 to 5.15
Brought up in #4225 where it had issues with Openat2 which was added in
5.8.

The main driving force around minimum kernel version requirement is that
the lowest kernel version in our CI is 5.15. A benefit to this choice is
that this is an LTS release, which is also what Ubuntu 22.04 is
shipping.

Once the single CI machine is fixed to ship something newer then the
next logical choice would be kernel 6.1 which is also LTS, but until
then just lift it to 5.15. This version was released in October 2021,
and is supported by the kernel developers until 2026. Our previous
minimum of 5.0 was released in March 2019, so a two year leap here.

This removes the openat2 workaround that was necessary to pass our CI
since it is no longer necessary.
2025-01-01 11:22:54 -08:00
Ryan Houdek beec203f56 LinuxSyscalls: Fixes exit syscall
if an application is using `exit` then it is usually a faulting
condition rather than cleanly exiting. When cleanly exiting
applications will typically use `exit_group` instead.

`exit` is useful to quickly cause a single thread to exit in a
multi-threaded environment as well, where `exit_group` will take down
the entire process group.

FEX had implemented this in a way that would do a double Stop signal,
cascading to a crash. When tied in to a crash handler, this could get
caught in a weird way.

This /should/ fix #4198, but I can't confirm locally. It looks like in
that issue that the steam install is slightly buggered (as evident by
missing srt-logger and steam-runtime-identify-library-abi).

This is a bug regardless so fix it and create a unittest. If it doesn't
fix the user's bug, then we have another workaround that will definitely
solve it.
2024-12-08 05:14:19 -08:00
Ryan Houdek dd8a3a9aea LinuxEmulation: Don't use clone3 for fork
clone3 was added in Linux 5.3 but our minimum spec is 5.0. Additionally
the Raspberry Pi 5 kernel seems to complain about clone3 for some
reason?

Just use clone instead of clone3
2024-12-05 15:14:37 -08:00
Ryan Houdek efb276f489 FEXCore: Removes ExitHandler and RunUntilExit
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.

- Instead of using RunUntilExit, all threads can use `ExecuteThread`
  directly, since there's nothing special about the primary thread now.
  - This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
  `ExecuteThread` has returned.
  - Just make gdbserver is cleaned up early if it exists since it may
    want to send some things to the connected gdb instance before
    threads are exited.
2024-12-01 10:45:38 -08:00
Ryan Houdek 802eaee9c8 FEXCore: Moves InternalThreadState ExecutionThread to the frontend
Once again this is another frontend construct, so move it to
ThreadStateObject
2024-11-29 13:33:56 -08:00
Ryan Houdek e771e25632 LinuxSyscalls/Thread: Build child thread arguments on parent stack
Now that most of the thread tracking is in the frontend, change this
over to building the thread execution handler on the parent thread.

Removes a memory allocation/free pair, and removes the copy of each
variable in the child thread.
2024-11-29 09:38:36 -08:00
Ryan Houdek 25c202575e FEXCore: Move InternalThreadState StartRunning to frontend
We were using this variable for two things, letting the frontend signal
to the backend that it wants to start executing once the thread is
created, and also for handling thread pausing. These two features are
conflated with one another and actually makes things more confusing.

- Move StartRunning/StartPaused to the frontend, because its a construct
  that only needs to exist in the frontend
- Adds a FEX::HLE::ThreadStateObject CV for handling pausing, which only
  needs to exist for gdbserver
2024-11-29 09:38:24 -08:00
Ryan Houdek 9f681f9e41 FEXCore: Moves ThreadWaiting to the frontend
Only in one location does the frontend actually care about this, the
backend doesn't care at all.
2024-11-29 09:10:36 -08:00
Ryan Houdek 1bf7e2544a FEXCore: Moves StatusCode to the frontend
This is a Linux construct, move it to the frontend.

This is going to need some changes in the future since exit_group and
exit syscalls are supposed to behave differently than how FEX implements
it. For now just move it to the frontend.
2024-11-28 15:55:46 -08:00
Asahi Lina bfed21870f Support CLONE_FS and CLONE_FILES with fork() semantics
Needed by Discord, part of the Chromium sandbox code. The warning still
triggers because Chromium asks for CLONE_VM on x86_64, but that can be
safely ignored (CLONE_FS is the one that matters).
2024-11-18 22:09:33 +09:00
Asahi Lina 3d701f5fcf FileManagement: Hide the FEX RootFS fd from /proc/self/fd
Chromium/CEF has code that iterates through all open FDs and bails if
any are directories (apparently a sandboxing sanity check). To avoid
this check, we need to hide the RootFS FD. This requires hooking all the
getdents variants to skip that entry.

To keep the runtime cost low, we keep track of the inode of
/proc/self/fd/<rootfs fd> (note: not the RootFS inode, the inode of the
magic symlink in /proc), and first do a quick check on that. If it
matches, then we stat the dirfd we are reading and check against the
procfs device, to complete the inode equality check.

As an extra benefit, this also fixes code that tries to iterate and
close all/extra FDs and ends up closing the RootFS fd.
2024-10-27 07:05:00 +09:00
Ryan Houdek 7f17519fbf LinuxEmulation/personality: Support PER_LINUX32 2024-10-11 04:52:23 -07:00
Ryan Houdek a65884f9ae LinuxEmulation/personality: Support UNAME26 2024-10-11 04:52:21 -07:00
Ryan Houdek d70766f4c8 LinuxEmulation: Support personality tracking
Doesn't handle the emulation of it, but handle passing it to the host
kernel, tracking the value, and inheriting it through new threads.
2024-10-11 04:52:21 -07:00
Ryan Houdek 611aa0a5b9 Thunks: Wire up TLS handling in the frontend
Because it was moved from the backend to the frontend, re-wire up the
TLS handling which requires going through our syscall handler.
2024-09-27 15:50:21 -07:00
Ryan Houdek f5def7ae1c Linux: Also optimize HandleNewClone stack usage
Drops from ~1392 bytes of stack usage to ~80 bytes
2024-09-22 10:51:25 -07:00
Ryan Houdek 58bcf691d6 Linux: Optimize CreateNewThread stack usage
We were creating a copy of the FEXCore::Core::CPUState object when we
didn't need to. We can pass the host thread's CPUState frame through to
the creation handlers since it's read-only (so modify it to be const).

We then just move the RAX and RSP setting to /after/ the CreateThread
handling instead of before.
This reduces stack usage from ~1392 bytes to ~80 bytes.
2024-09-22 10:45:21 -07:00
Ryan Houdek 64c5362580 LinuxSyscalls: With poll syscall, ensure fds is writable only if nfds is not zero
The pointer is allowed to be null or garbage if the number of nfds
passed in is zero.
Unify the x86 and x86-64 implementations to ensure consistency.

Fixes Warhammer 40k: Relics of War
2024-09-05 14:33:52 -07:00
Ryan Houdek ac32876e4e LinuxEmulation: Implement support for seccomp
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.

The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.

There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.

FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.

While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.

**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour

**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.

**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this

Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
   - This means we serialize and deserialize the calling thread's
     seccomp filters
   - An execve that escapes FEX will also escape seccomp. Not much we
     can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
  synchronize seccomp filters after the fact

Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
  - Just like our syscall handler, we are hardcoded to the arch that the
    application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
  - Currently we must ensure all syscalls go through the frontend
    syscall handler
  - Runtime invalidation of code cache with inline syscalls will get
    fixed in the future.

This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.

Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
2024-09-02 14:07:53 -07:00
Ryan Houdek 9056d9b9de SignalDelegator: Refactor how thread local data is stored
Two primary things here:
- Remove the static `GlobalDelegator`
- Move the thread_local SignalDelegator::ThreadState information
  directly in to ThreadStateObject

Having the ThreadStateObject and the SignalDelegator information
disjoint was confusing but was required when we didn't have any object
in the frontend that could have its own independent data. Since we fixed
this with the `ThreadStateObject` type we can now move this over.

The `GlobalDelegator` object is now instead stored in
`ThreadStateObject` instead.

Instead of using a thread_local variable, we now just consume 8-bytes of
the signal alt-stack since the kernel gives us that information about
where it lives. This then converts all the thread_local usage to use
either the passed in CPU state if it exists, or fetching it from the
alt-stack offset.

Very minor changes in behaviour here, will help when trying to improve
FEX's behaviour around signals.
2024-09-02 06:46:19 -07:00
Ryan Houdek 40bbcb7061 Syscalls: Updates for v6.10
Only mseal was added and can be a simple passthrough for us. Which is
nice.
2024-07-30 19:08:27 -07:00
Ryan Houdek 3b2e657fd4 FEXCore: Removes ThreadManager
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.

Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
2024-07-25 14:54:10 -07:00
Ryan Houdek 2fa6c3c918 LinuxEmulation: Add a helper for getting the ThreadStateObject from CPU frame
Pulled from the seccomp WIP PR where it pulls this object more frequently.
Since it is an opaque frontend pointer it needs to be cast and we
already have a few locations that use it.

No functional change.
2024-06-15 18:31:37 -07:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Ryan Houdek d372552593 FEXLoader: Changes frontend thread management to wrap FEXCore thread objects
A bit of refactoring necessary before we can move the remaining Linux
specific code to the frontend.

Most of this taken from #3535 but attempting to be NFC as much as
possible.
2024-05-05 07:43:09 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek 20eb338644 FEXCore: Moves CodeLoader to frontend
FEXCore no longer has a need for this since a bunch of related code was
already moved to the frontend. Move the CodeLoader now.
2024-03-29 02:24:53 -07:00
Ryan Houdek 5a35e119fe Telemetry: Adds tracker for non-canonical memory access crash
This may be useful for tracking TSO faulting when it manages to fetch
stale data. While most TSO crashes are due to nullptr dereferences, this
can still check for the corruption case.
2024-03-21 20:47:36 -07:00
Ryan Houdek 45ea0cd782 Removes false termux support
This was a funny joke that this was here, but it is fundamentally
incompatible with what we're doing. All those users are running proot
anyway because of how broken running under termux directly is.

Just remove this from here.
2024-03-20 22:04:32 -07:00
Ryan Houdek 8a607135fd Linux: Update syscalls for v6.8 2024-03-10 15:22:51 -07:00
Ryan Houdek 151e2279af Linux: Converts passthrough syscalls to direct passthrough handlers
Reimagining of #3355 without any json generators or new concepts.

Fixes some mislabeling of system calls. Some getting inlined when they
shouldn't be, a lot not getting inlined when they can be.

This really cleans up the syscall implementation, all syscalls that can
be passthrough implementations require a very small two line
declaration.
Additionally cleans up a bit of implementation cruft where some
passthrough syscalls were using the glibc syscall handler, and some were
using the glibc implementation. We have had multiple issues in the past
where the glibc implementation does something subtly different than the
raw syscall and breaks things. Now all passthrough handlers do a system
call directly, removing at least one indirection and some ambiguity.

This makes it significantly easier to add new passthrough syscalls as
well. Only need to do a version check and add the three lines per
syscall. Which there are new syscalls incoming that we will want to add.

Tangible improvements:
- Syscalls are lower overhead than ever.
- When I'm adding more syscalls I have less chance of mucking it up.
2024-02-27 02:40:53 -08:00
Ryan Houdek 93ada89708 Linux: Move unimplement ustat and sysfs
AArch64 doesn't implement these and will return ENOSYS.
Moving them to NotImplemented so we can get a log if an application
tries to use these.
2024-02-27 02:39:36 -08:00
Ryan Houdek 0a64f8a9c5 Moves SignalDelegator TLS tracking to the frontend
FEXCore doesn't need track the TLS state of the SignalDelegator, this is
a frontend concept.

Removes the tracking from the backend and keeps it in the frontend.
2024-02-24 01:07:29 -08:00
Ryan Houdek 80f20ad121 Linux: Make sure to destroy thread object when thread shuts down
This fixes a fairly large memor leak.
2024-02-16 15:08:29 -08:00
Ryan Houdek 577372c203 Linux: Consolidate LockBeforeFork usage
Moves the CTX LockBeforeFork in to the Syscallhandler's LockBeforeFork.

This lets the syscall handler just call its own LockBeforeFork and
UnlockAfterFork functions rather than two on each call site.

Also moves the CTX->UnlockAfterFork in to the SyscallHandler's to be
consistent with the LockBeforeFork half.

No functional change.
2024-02-09 05:55:23 -08:00