Commit Graph
124 Commits
Author SHA1 Message Date
Ryan Houdek 5f9a30af7a FileManagement: Use const ref for RootFSPath
No need to make a copy
2024-09-08 15:20:04 -07:00
Ryan Houdek 453ef0b94c Linux/ExecveHandler: Move a path rather than copy
We can do move assignment instead of copy assignment here.
2024-09-08 15:17:04 -07:00
Ryan Houdek 4779e64cf5 Linux/EmulatedFiles: Check return for seal
This shouldn't be possible to occur, but have a log message just in
case.
2024-09-08 15:15:38 -07:00
Ryan Houdek 64c5362580 LinuxSyscalls: With poll syscall, ensure fds is writable only if nfds is not zero
The pointer is allowed to be null or garbage if the number of nfds
passed in is zero.
Unify the x86 and x86-64 implementations to ensure consistency.

Fixes Warhammer 40k: Relics of War
2024-09-05 14:33:52 -07:00
Alyssa Rosenzweig b368223d50 Merge pull request #3628 from Sonicadvance1/seccomp
LinuxEmulation: Implement support for seccomp
2024-09-03 08:49:16 -04:00
Ryan Houdek ac32876e4e LinuxEmulation: Implement support for seccomp
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.

The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.

There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.

FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.

While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.

**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour

**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.

**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this

Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
   - This means we serialize and deserialize the calling thread's
     seccomp filters
   - An execve that escapes FEX will also escape seccomp. Not much we
     can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
  synchronize seccomp filters after the fact

Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
  - Just like our syscall handler, we are hardcoded to the arch that the
    application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
  - Currently we must ensure all syscalls go through the frontend
    syscall handler
  - Runtime invalidation of code cache with inline syscalls will get
    fixed in the future.

This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.

Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
2024-09-02 14:07:53 -07:00
Ryan Houdek f8eaf9c14f VDSO: Fixes a pretty nasty bug where we were never using the host VDSO
This must have happened during a refactor or something, but since we're
making a copy of the function pointers, it would have only gotten the
version /prior/ to loading host VDSO symbols.

Moves the VDSO thunk definition setting to the end after VDSO symbol
definitions in order to get host VDSO symbols working again.
2024-09-02 09:18:37 -07:00
Ryan Houdek 9056d9b9de SignalDelegator: Refactor how thread local data is stored
Two primary things here:
- Remove the static `GlobalDelegator`
- Move the thread_local SignalDelegator::ThreadState information
  directly in to ThreadStateObject

Having the ThreadStateObject and the SignalDelegator information
disjoint was confusing but was required when we didn't have any object
in the frontend that could have its own independent data. Since we fixed
this with the `ThreadStateObject` type we can now move this over.

The `GlobalDelegator` object is now instead stored in
`ThreadStateObject` instead.

Instead of using a thread_local variable, we now just consume 8-bytes of
the signal alt-stack since the kernel gives us that information about
where it lives. This then converts all the thread_local usage to use
either the passed in CPU state if it exists, or fetching it from the
alt-stack offset.

Very minor changes in behaviour here, will help when trying to improve
FEX's behaviour around signals.
2024-09-02 06:46:19 -07:00
Alyssa Rosenzweig d9544e7e02 Merge pull request #4018 from Sonicadvance1/move_guest_frame_manage
LinuxEmulation: Moves guest signal frame generation to its own file
2024-09-02 09:21:06 -04:00
Alyssa Rosenzweig 62e1767ee0 Merge pull request #4019 from Sonicadvance1/more_efault_handlers
More EFAULT handlers
2024-09-02 09:20:55 -04:00
Ryan Houdek c748dbf0e3 FEX: Moves sigreturn symbols to frontend
These are a Linux construct and should live here. Removes a weird
passthrough API from FEXCore and keeps it in the frontend instead.
This isn't even typically allocated in a real setup, as it's only a
fallback for if VDSO isn't loaded.

The CallbackReturn function stays in FEXCore because it would have
caused an API in the other direction instead.
2024-08-31 07:43:04 -07:00
Ryan Houdek cf7ee98831 x32/Thread: Adds EFAULT handlers 2024-08-28 04:31:06 -07:00
Ryan Houdek 5529948479 x32/Time: Adds EFAULT handlers 2024-08-28 04:25:13 -07:00
Ryan Houdek 91b6aeffe1 x32/Timer: Adds EFAULT handlers 2024-08-28 04:19:45 -07:00
Ryan Houdek 3a79b61f58 x64/Thread: Adds EFAULT handlers for two syscalls
For the remaining syscalls we need to introduce some new concepts around
checking if pointers and string arrays are readable and I don't want to
do that yet.
2024-08-28 04:14:16 -07:00
Ryan Houdek e92b24302f LinuxEmulation: Moves guest signal frame generation to its own file
No functional change, just moving the code.

This is entirely freestanding from the rest of the signal delegator
handling and mostly gets in the way when I'm working with the rest of
the signal handling. Separate it out to improve readability.
2024-08-28 03:23:00 -07:00
Ryan Houdek 9c7f44f8d4 x32/Signals: Add EFAULT checks 2024-08-27 03:04:42 -07:00
Ryan Houdek 5117ba351e x32/Semaphore: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 441fdb689d x32/Sched: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek e786dfc998 x32/Msg: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek a1d3183c14 x32/Info: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek c5ef0910c5 x32/IO: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek affa1d0efc x32/FD: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek a83dc27a42 x32/EPoll: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 129ec63e92 x64/Time: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 3459369c6e x64/Signals: Add EFAULT checks 2024-08-27 03:04:24 -07:00
Ryan Houdek 7a490a3811 x64/Semaphore: Add EFAULT checks 2024-08-27 03:04:23 -07:00
Ryan Houdek ee17fe239a x64/EPoll: Add EFAULT checks 2024-08-27 03:04:23 -07:00
Ryan Houdek 6e46383cdb x32/Signals: Fixes bug in the sigqueue syscalls
In rt_sigqueueinfo and rt_tgsigqueueinfo we were failiny to pass along
the provided siginfo_t information to the syscall.

Fix that
2024-08-27 02:54:27 -07:00
Ryan Houdek f6cb914a2a SignalDelegator: Changes AFmt to ERROR_AND_DIE_FMT
To ensure that it actually dies after logging.
2024-08-25 04:27:59 -07:00
Ryan Houdek 0dc0117e86 SignalDelegator: Make sure to defer a signal if the guest signal mask desires
Fairly trivial because we already supported deferring these signals. We
had just failed the final step of blocking the signal if we can't block
the signal (like with SIGSEGV, SIGBUS, etc).

Once the guest unblocks the signal mask with sigprocmask, the signal
will fire again.

Apparently older glibc relied on this behaviour for signal raising which
this fixes.
2024-08-22 19:48:23 -07:00
Ryan Houdek 5edc69b692 LinuxSyscalls: Adds missing header 2024-08-22 14:02:17 -07:00
Ryan Houdek ce88f5f948 Merge pull request #3941 from Sonicadvance1/fault_assertion_checking
LinuxSyscalls: Implements less invasive assertion only EFAULT handlers
2024-08-15 01:14:15 -07:00
Ryan Houdek 2b26d0fff5 LinuxSyscalls: Implements less invasive assertion only EFAULT handlers
With the previous Copy{To,From}User helpers we need to actually
implement the handlers correctly. We want something that is a bit
lighter so we don't need to implement the faulting path in the syscall
handlers.

Implements a handful of helpers that just check for readable and
writable capability which can be thrown in to an assertion handler that
is zero cost in release mode.

Readable is checked by just attempting to read all bytes.
Writable is checked by attempting to read each byte and writing it back
to the same location.

Uses these helpers in x64/FD.cpp to showcase how they will be used to
detect EFAULT. Tested locally that they work correctly by writing some
small tests for the syscalls that expect EFAULT.
2024-08-15 00:48:57 -07:00
Ryan Houdek 924723d433 LinuxSyscalls: Some minor cleanups
- We can have the SyscallFunctionDefinitions be the correct size out of
  the gate. Both tables are always 512 entries in size.
- In the RegisterSyscall_{32,64} handlers, just get the reference using
  operator[]. We always know we will be under the size of the array, add
  a an assert to check. Removes a bit of vector range checking overhead.
- Namespace 32-bit syscalls like 64-bit syscalls and include in the
  regular header like 64-bit. This was just an oversight
- Use std::fill for the syscall gap for the invalid syscall, just a
  minor cleanup.

No functional change.
2024-08-13 14:06:16 -07:00
Ryan Houdek 40bbcb7061 Syscalls: Updates for v6.10
Only mseal was added and can be a simple passthrough for us. Which is
nice.
2024-07-30 19:08:27 -07:00
Ryan Houdek 67663b812e Linux: Update syscall defines for v6.10 2024-07-30 19:01:52 -07:00
Ryan Houdek 403fd62b34 Merge pull request #3890 from Sonicadvance1/refactor_frontend_threadmanager
FEXCore: Removes ThreadManager
2024-07-26 13:27:43 -07:00
Mai 93eead243f Merge pull request #3864 from Sonicadvance1/threads_atexit_remove
Threads: Setup the stack tracker to not need global initialization
2024-07-26 06:38:29 -04:00
Ryan Houdek 3b2e657fd4 FEXCore: Removes ThreadManager
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.

Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
2024-07-25 14:54:10 -07:00
Ryan Houdek ce8bc9d25c FEXCore: Refactor ExitHandler slightly
Instead of passing the TID back to the exit handler, just pass the whole
thread object. This will allow some cleanups with the frontend thread
tracking soon

NFC
2024-07-24 14:39:56 -07:00
Ryan Houdek d2f903ae55 EmulatedFiles: Adds a few leaf CPUID flags
We support leaf functions now, so add the few that were calling for it.
We will be gaining support for the xsave ones relatively soon, so its
good to have them supported.

Also deletes a couple of cdt/cqm things that aren't exposed and we won't
be supporting.
2024-07-18 07:02:49 -07:00
Tony Wasserka bf9a6d763c EmulatedFiles: Fix bad formatting 2024-07-18 15:06:57 +02:00
Ryan Houdek d1b5dfd4b1 Threads: Setup the stack tracker to not need global initialization
Also removes the atexit handler installation
This now gets tracked by an object owned by FEXLoader (and shared with
the pthreads interface)
2024-07-12 03:00:20 -07:00
Ryan Houdek 6cdaea680d Ioctl32: Removes static fextl::vector in ioctlemulation
Removes a global static initializer for the vector and its atexit
handler.

This handler array can be consteval similar to the x86 tables so it can
be generated entirely at compile time.
2024-07-12 02:05:32 -07:00
Ryan Houdek 5ef0db994d VDSO: Stop using a vector for a static
This causes a global initializer that registers an atexit handler.

Be smarter, use an std::array and pass its data around using a span
instead.

Removes the global initializer and removes the atexit installation
2024-07-11 23:53:57 -07:00
Tony Wasserka 4dec8f22f8 Fix packed-non-pod warnings 2024-07-11 09:54:30 +02:00
Paulo Matos ad52514b97 Use number of jobs as defined by TEST_JOB_COUNT
At the moment we always run ctest with max number of cpus. If
undefined, it will keep current behaviour, otherwise it will
honour TEST_JOB_COUNT.

Therefore to run ctest one test at a time, use
`cmake ... -DTEST_JOB_COUNT=1`
2024-07-03 14:09:39 +02:00
Ryan Houdek be6ff52709 Linux: Calculate cycle counter frequency for cpuinfo
Some applications don't measure rdtsc correctly and instead use cpuinfo
to get the CPU core's base clock speed. Which for most x86 CPUs is the
base clock speed which also matches their cycle counter speed.

Did this as a quick test to see if this would help `Unbound: Worlds
Apart` stuttering while BinaryNinja was disassembling the binary.

Turns out the game doesn't use cpuinfo for its cycle counter speed
determination, but it is good to implement this regardless.
2024-06-28 16:38:49 -07:00
Ryan Houdek f5fea8af96 SignalDelegator: Use new YMM register reconstruction helpers
Otherwise we would be setting up signal handlers with incorrect register
state.
2024-06-21 17:13:56 -04:00