A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.
In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.
This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.
With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
Effectively NFC today while we don't allow installing LDT entries, but
is required to have signal handlers execute in the correct bitness if
32-bit code segments are enabled.
Nothing too crazy, just that four segments are stuffed in to the
`REG_CSGSFS` data value, and then CS/SS gets restored on restore.
FS and GS are ignored in this instance because edge case behaviour where
those don't get reset on signal entry, so if they get changed then it is
on the userspace application to fix it.
Gets another change out of my stashes.
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.
Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
This patch does cover up all v4l2 syscalls in i386. It just works
well with the poor v4l2 driver of wine32.
Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.
In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.
Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.
The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
FEX should only set fpx_sw_bytes.magic1 with FP_XSTATE_MAGIC when
avx is enabled. Otherwise it may cause segfault in ntdll::save_context,
which requires access to extended xstate info if magic1 equals FP_XSTATE_MAGIC.
Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
WINE uses this to determine TSC frequency and because it doesn't say
`tsc` on ARM devices, it was ignoring TSC and instead using CPU maximum
frequency.
This was causing Horizon to think the TSC ran at whatever the max
frequency of a core was (1.8Ghz to 2.6Ghz depending?) This was causing
all of Horizon Zero Dawn's physics to run at slower than real time
speeds because our 1Ghz (on Orion) TSC is significantly lower than the
max clock speeds of the cores.
This is still a bug in Wine that it is using the maximum CPU clock speed
in the case of current_clocksource not being TSC, but that's a battle
for a different time.
Wine does some minimal attaching, peeking, poking, and detaching to have
a different process inspect another's TEB region. Ubisoft's launcher
does this to check if a debugger is attached and rejects it in the case
that it is.
While this implementation isn't all-encompassing, it is good enough for
this family of games.
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
the modifying thread. Old versions of the CodeBuffer persist as read-only
data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
decrease the reference count and eventually trigger deallocation of the
old version
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.
This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.
The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>) = 0
<...>
41227 <... mmap resumed>) = 0xba87b000
```
While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```
The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.
This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.
The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.
Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.
NFC, just helps my brain wrap this more easily.