Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.
We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.
Performance-wise this change falls within noise of my x87 microbench.
This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
Not wired up, just the definitions so it lives in the
InternalThreadState.
We want this accessible from both FEXCore and the frontends so it needs
to live there.
Two types of events supported. Scoped cyclecounts and instant
increments.
This gives us JIT time and Signal handling time, plus events for number
of SIGBUS and number of SMC events.
All useful statistics for seeing stutter live.
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.
This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.
Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.
This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.
On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.
Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.
This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.
Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
Placing optimization hints everywhere interferes with debugging of
RelWithDebInfo builds, since the debugger won't be able to reliably
inspect variables or control flow. These hints are better placed on an
individual basis after identifying bottlenecks in a profiler.
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
guest instruction
Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.
- Instead of using RunUntilExit, all threads can use `ExecuteThread`
directly, since there's nothing special about the primary thread now.
- This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
`ExecuteThread` has returned.
- Just make gdbserver is cleaned up early if it exists since it may
want to send some things to the connected gdb instance before
threads are exited.
These are all frontend constructs with mostly deprecated constraints.
WaitingToStart isn't used anymore, Running is effectively always true
(and behaviour has changed that if a thread is alive, it's running).
The only one that remains is `ThreadSleeping` which is only handled in
the frontend, and there was some conflation between ThreadSleeping and
Running which was hard to gauge. So delete `Running` and
`WaitingToStart`, but move `ThreadSleeping` to the frontend.
We were using this variable for two things, letting the frontend signal
to the backend that it wants to start executing once the thread is
created, and also for handling thread pausing. These two features are
conflated with one another and actually makes things more confusing.
- Move StartRunning/StartPaused to the frontend, because its a construct
that only needs to exist in the frontend
- Adds a FEX::HLE::ThreadStateObject CV for handling pausing, which only
needs to exist for gdbserver
FEXCore hasn't been returning anything other than EXIT_SHUTDOWN for a
long time, so this ended up just moving data around for no reason.
This isn't going to be used for further GdbServer work anyway, so just
completely remove it.
This is a Linux construct, move it to the frontend.
This is going to need some changes in the future since exit_group and
exit syscalls are supposed to behave differently than how FEX implements
it. For now just move it to the frontend.
Alloc::OSAllocator uses a TLS variable of the thread object so it can
use a forkable mutex plus a deferring signal section. This was setup
when the FEXCore "ExecutionThread" function is called, which is a bit
awkward and is an artifact from when the thread creation was mixed
between the frontend and the backend.
Instead let the frontend inform the backend when to install the TLS
variable.
This is one step required to make GdbServer work correctly again since
the thread initialization and pausing is awkward today.
This was working around an edge case in the GdbServer where a thread was
getting created while the process was shutting down. This edge case is
getting removed so get rid of it.