Consider a page with two blocks in it, A and B. Block A performs SMC on B then A.
With the previous logic, the SMC write of B would unprotect the page and then
the inline SMC touching A would not be detected and a single-step would not be forced.
Rather than relying on wine-specific alphabetical behaviour, that
has since been changed. Rely on their addresses being sorted which
is more stable behaviour also present in Windows.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
This is a little trickier, we actually open the
`/dev/shm/fex-<pid>-stats` file directly using Windows APIs that way
Mangohud (which is going to be on the Linux side, or potentially even
embedded in to Gamescope) can safely pick up the stats.
A little quirky plus doesn't support expanding its size since WINE
doesn't support NtExtendSection, but that's fine.
The thread termination callback can be called for other threads in the
process, not just the current one, in which case we cannot call DeinitCRT.
Deinitializing the CRT of another thread would be awkward so just skip that
and accept the small leak for now.
When an SMC trap happens: reconstruct the context before the SMC write
then compile the write as a single instruction block to reduce it to
regular SMC. SMC where the writing instruction is the instruction being
patched will hit the signal handler at most twice: the 1st will trigger
the write to be compiled as a single instuction block, the 2nd will
detect inline SMC of a single instruction block and then just take the
usual invalidate+reprotect+continue step, avoiding a potential infinite
loop of recompilation.
As section permissions are set on the unix side we don't get a
protection callback for them, workaround this by iterating over
the sections of all executables after mapping and tracking the RWX
ones.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
CEF on Windows patches these before FEX is even loaded, so this needs to
be setup as early as possible using some direct syscalls which skip any
such patches.
The previous approach assumed that ntdll wouldn't have been patched
before FEX was loaded, but that isn't necessarily the case thanks to
CEF. This takes the same approach used by xtajit which wine recently gained
support for wine. The overhead to this isn't ideal but most syscalls
aren't directly done through x86 code anyway and in the future FEX could
forward unpatched thunks to their ARM variants at JIT time.