This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.
L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.
One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.
Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!
Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.
before:
```
Total FEX Anon memory resident: 1132 mB
JIT resident: 60 mB
OpDispatcher resident: 97 mB
Frontend resident: 37 mB
CPUBackend resident: 500 kB
Lookup cache resident: 629 mB
Lookup L1 cache resident: 108 mB
ThreadStates resident: 436 kB
```
after:
```
Total FEX Anon memory resident: 438 mB
JIT resident: 62 mB
OpDispatcher resident: 56 mB
Frontend resident: 22 mB
CPUBackend resident: 496 kB
Lookup cache resident: 0 (null)
Lookup L1 cache resident: 109 mB
ThreadStates resident: 436 kB
```
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.
Fix for __builtin_issignaling() test of SPEC2017 classify test.
Only wired up for wow64 and arm64ec. Gives more granular control over
TSO enabling and disabling. Matches arm64ec volatile metadata except
with one more additional feature that whole modules can be disabled at a
time.
Once we know the mapped size of files in Linux then we'll be able to do
the same thing there, but there's not a full mechanism wired up for that
yet.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.
We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
Not wired up, just the definitions so it lives in the
InternalThreadState.
We want this accessible from both FEXCore and the frontends so it needs
to live there.
Two types of events supported. Scoped cyclecounts and instant
increments.
This gives us JIT time and Signal handling time, plus events for number
of SIGBUS and number of SMC events.
All useful statistics for seeing stutter live.
This was causing FEXServer to look in to global installed paths and
local paths for things when FEXServer was started.
Ensure it listens to FEX_PORTABLE so this doesn't occur.
This also requires us to scan both data directories and config
directories to find them.
No functional change here.
- CoreRunningMode enum and variable wasn't used anymore.
- Code was moved to the frontend
- CustomCPUFactory wasn't used anymore
- All special signal handling and various features were moved to
TestHarnessRunner
- We also don't want to support actual custom CPU cores.
- TestHarnessRunner just runs as a host runner if compiled on an
x86-64 device if vixl sim isn't enabled now.
- Removes the Core config option entirely.
- Moves VDSOPointers struct to the frontend
- Every use of this lives in the Linux frontend instead now
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.
The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.
There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.
FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.
While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.
**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour
**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.
**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this
Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
- This means we serialize and deserialize the calling thread's
seccomp filters
- An execve that escapes FEX will also escape seccomp. Not much we
can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
synchronize seccomp filters after the fact
Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
- Just like our syscall handler, we are hardcoded to the arch that the
application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
- Currently we must ensure all syscalls go through the frontend
syscall handler
- Runtime invalidation of code cache with inline syscalls will get
fixed in the future.
This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.
Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.
I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.
One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.
This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
After two months of testing I finally have enough confidence that these
default options getting changed is safe enough that most games won't
notice the difference. But the performance differences can be wild.
Might want to come back and reenable this on platforms that support
hardware TSO and LRCPC3 (for the vector feature), but we can care once
those platforms come up to speed.
A feature of FEX's JIT is that when an unaligned atomic load/store
operation occurs, the instructions will be backpatched in to a barrier
plus a non-atomic memory instruction. This is the half-barrier technique
that still ensures correct visibility of loadstores in an unaligned
context.
The problem with this approach is that the dmb instructions are HEAVY,
because they effectively stop the world until all memory operations in
flight are visible. But it is a necessary evil since unaligned atomics
aren't a thing on ARM processors. FEAT_LSE only gives you unaligned
atomics inside of a 16-byte granularity, which doesn't match x86
behaviour of cacheline size (effectively always 64B).
This adds a new TSO option to disable the half-barrier on unaligned
atomic and instead only convert it to a regular loadstore instruction,
ommiting the half-barrier. This gives more insight in to how well a
CPU's LRCPC implementation is by not stalling on DMB instructions when
possible.
Originally implemented as a test to see if this makes Sonic Adventure 2
run full speed with TSO enabled (but all available TSO options disabled)
on NVIDIA Orin. Unfortunately this basically makes the code no longer
stall on dmb instructions and instead just showing how bad the LRCPC
implementation is, since the stalls show up on `ldapur` instructions
instead.
Tested Sonic Adventure 2 on X13s and it ran at 60FPS there without the
hack anyway.
This environment variable had an incorrect priority on the configuration
system. The expectation was higher priority than most other layers.
Now the only layer that has higher priority is the environment
variables.
This is used for instcountci to ensure instruction counts don't change
when a compiler supports this feature or not. Always runtime disable
when running in instcountci.
CMake option from #3394 can still be useful so leaving that in place.