This matches behaviour with the x86 host runner that RSP isn't
guaranteed to be set to a valid memory location. Make sure when running
tests on ARM that it gets the same behaviour.
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.
NFC, just helps my brain wrap this more easily.
With fortifications enabled, glibc long jump has some additional checks
in place that break because we do a stack pivot. The only way around
this is to do our own long jumps. Luckily this is trivial.
Fixes#4558
I noticed this cascade of mprotects when poking at Crypt of the
Necrodancer, since it consistently is invalidating code. I saw us
calling mprotect on the same page 32 times in a tight loop and thought
surely this isn't FEX doing this.
Turns out we were calling the callback after invalidating each thread.
It should instead be done once at the end of invalidating the thread's
caches while still holding the locks.
Fixes this weird cascade of mprotects that equal the number of FEX
threads.
Newer fuse releases changed how they are waiting on child processes to
exit. Setting the signal action to SIG_IGN would cause
erofsfuse/squashfuse to inherit the ignored action and cause their
internal `wait4` syscalls to fail with ECHLD.
Set the action back to default inside the FEXServer because our original
reasoning for setting the ignoring is no longer valid. FEXInterpreter
still ignores SIGCHLD while launching FEXServer.
Maybe fixes the muvm thing people have been complaining about.
Also fixes accidental comma delimiter usage.
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.
In addition, there a couple of clang-tidy fixes which should be NFC.
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.
Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.
So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.
This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.
This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.
This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
With the previous stack leak fix, RUINER reduced its memory leaking down
to around 50MB/s. The remaining memory leaks come from the pthread stack
that we are required to allocate (128KB per thread) and some internal
DTV tracking structures.
The problems come in the fact that glibc/pthread only tears down its
internal state for these if the pthread function actually returns! We
can **technically** switch the initial stack over to a "user" stack but
that introduces more problems around internal dtv tracking that we
already fixed months ago, so we can't actually do that in practice.
This leaves us no choice, we effectively are mandated to return from the
pthread function in order to free the memory from glibc. The only way we
can safely do this is with a long jump and deferring some data structure
management until that case.
This is all incredibly sucky but it's necessary to work. With these
changes, RUINER is no longer leaking memory (Hovering at around 3GB used
while in-game) and even Steam is consuming less memory.
It doesn't solve the problem that thread creation and teardown could
likely be faster, but not many applications are creating 720
threads/second.
Fixes an issue where a thread that exits with the `exit` syscall never
actually frees its pivot stack. This is common practice and it was
missed when I was fixing the previous stack pivot leak.
This was uncovered when looking at the game
[RUINER](https://store.steampowered.com/agecheck/app/464060/?curator_clanid=4777282)
for timing bugs. Turns out the Linux build of the game creates and
destroys a VLC object every tick of the engine. Creating this VLC
object creates six threads behind the scenes. At what I assume the
default tick of the engine is of 120Hz(?) this would mean it is
attempting to create and destroy 720 threads per second, quickly
leading to memory exhaustion under FEX.
While this hits the biggest memory leak we have around thread creation,
this game is still hitting thread creation so hard that I can see other
leaks that I need to track down still.
In regular x86 programs, when a signal occurs, the signal will not be handled within the signal handler. However, under FEX's defer signal mechanism, the signal is not immediately masked when it is deferred. When returning to the location that receives the signal and continues processing, the signal might be received again, causing inconsistency between the emulation and the actual program.
Here is an unit test for this patch from ltp:
https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/syscalls/timer_settime/timer_settime03.c
Instead of hardcoding the bogomips calculation, more closely match what
the Linux kernel does for bogomips. There are some applications out
there that use bogomips for silly timing calculations, so this is a
better version.
`llseek` returns only ever 0 or errno in the return register. This is in
contrast to `lseek` which returns the result (or errno) in the return
register.
We were accidentally returning the result on non-error conditions which
could freak out some software. Thanks to
[OFFTKP](https://github.com/OFFTKP) for pointing out this issue
This changes the instcountCI code to consistently load test data in to
RIP 0x1'0000 so we don't have any spurious changes due to json changes.
This has been a minor annoyance where if a test was added, it had the
potential to shift the rest of the data in the tests. This now ensures
it is consistent.
Avoids the non-fatal error: "[ERROR] Close closing FEX FD 2"
that happens when the guest program executes a syscall to close fd 2,
and it's not in fex tracked set.
These are completely unused since this ELFContainer is significantly
less utilized than original expectations.
More of this code is dead and can be removed in the future.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
Since this syscall doesn't exist, we need to convert it to the
equivalent utimensat like the kernel does internally.
This is fairly trivial but there are some safety nets in place.