This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
the modifying thread. Old versions of the CodeBuffer persist as read-only
data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
decrease the reference count and eventually trigger deallocation of the
old version
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.
This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.
The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>) = 0
<...>
41227 <... mmap resumed>) = 0xba87b000
```
While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```
The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.
This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.
The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.
Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.
NFC, just helps my brain wrap this more easily.
With fortifications enabled, glibc long jump has some additional checks
in place that break because we do a stack pivot. The only way around
this is to do our own long jumps. Luckily this is trivial.
Fixes#4558
I noticed this cascade of mprotects when poking at Crypt of the
Necrodancer, since it consistently is invalidating code. I saw us
calling mprotect on the same page 32 times in a tight loop and thought
surely this isn't FEX doing this.
Turns out we were calling the callback after invalidating each thread.
It should instead be done once at the end of invalidating the thread's
caches while still holding the locks.
Fixes this weird cascade of mprotects that equal the number of FEX
threads.
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.
In addition, there a couple of clang-tidy fixes which should be NFC.
With the previous stack leak fix, RUINER reduced its memory leaking down
to around 50MB/s. The remaining memory leaks come from the pthread stack
that we are required to allocate (128KB per thread) and some internal
DTV tracking structures.
The problems come in the fact that glibc/pthread only tears down its
internal state for these if the pthread function actually returns! We
can **technically** switch the initial stack over to a "user" stack but
that introduces more problems around internal dtv tracking that we
already fixed months ago, so we can't actually do that in practice.
This leaves us no choice, we effectively are mandated to return from the
pthread function in order to free the memory from glibc. The only way we
can safely do this is with a long jump and deferring some data structure
management until that case.
This is all incredibly sucky but it's necessary to work. With these
changes, RUINER is no longer leaking memory (Hovering at around 3GB used
while in-game) and even Steam is consuming less memory.
It doesn't solve the problem that thread creation and teardown could
likely be faster, but not many applications are creating 720
threads/second.
Fixes an issue where a thread that exits with the `exit` syscall never
actually frees its pivot stack. This is common practice and it was
missed when I was fixing the previous stack pivot leak.
This was uncovered when looking at the game
[RUINER](https://store.steampowered.com/agecheck/app/464060/?curator_clanid=4777282)
for timing bugs. Turns out the Linux build of the game creates and
destroys a VLC object every tick of the engine. Creating this VLC
object creates six threads behind the scenes. At what I assume the
default tick of the engine is of 120Hz(?) this would mean it is
attempting to create and destroy 720 threads per second, quickly
leading to memory exhaustion under FEX.
While this hits the biggest memory leak we have around thread creation,
this game is still hitting thread creation so hard that I can see other
leaks that I need to track down still.
In regular x86 programs, when a signal occurs, the signal will not be handled within the signal handler. However, under FEX's defer signal mechanism, the signal is not immediately masked when it is deferred. When returning to the location that receives the signal and continues processing, the signal might be received again, causing inconsistency between the emulation and the actual program.
Here is an unit test for this patch from ltp:
https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/syscalls/timer_settime/timer_settime03.c
Instead of hardcoding the bogomips calculation, more closely match what
the Linux kernel does for bogomips. There are some applications out
there that use bogomips for silly timing calculations, so this is a
better version.
`llseek` returns only ever 0 or errno in the return register. This is in
contrast to `lseek` which returns the result (or errno) in the return
register.
We were accidentally returning the result on non-error conditions which
could freak out some software. Thanks to
[OFFTKP](https://github.com/OFFTKP) for pointing out this issue
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
Since this syscall doesn't exist, we need to convert it to the
equivalent utimensat like the kernel does internally.
This is fairly trivial but there are some safety nets in place.
Previously, new VMA entries were always prepended to the list of the
associated MappedResource. This usually made FirstVMA erroneously point to
the *highest* VMA instead of the lowest.
The Linux kernel clears these flags on signal, DF is particularly
dangerous because it would break ABI if a signal happened to occur in
the middle of a memory operation that changed the direction of copy.
Because of how frequently wine uses signals, this is actually fairly
likely to occur inside of a memcpy/memset function.
Shout out to BlinkDagger on Discord who found that we forgot to do this.