This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
Fixes#4535
A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher
This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.
This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.
In addition, there a couple of clang-tidy fixes which should be NFC.
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.
Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.
So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.
This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.
This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.
This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.
We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.
Performance-wise this change falls within noise of my x87 microbench.
This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
Not wired up, just the definitions so it lives in the
InternalThreadState.
We want this accessible from both FEXCore and the frontends so it needs
to live there.
Two types of events supported. Scoped cyclecounts and instant
increments.
This gives us JIT time and Signal handling time, plus events for number
of SIGBUS and number of SMC events.
All useful statistics for seeing stutter live.
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.
This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.
Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.
This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.
On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.
Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.
This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.
Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
Placing optimization hints everywhere interferes with debugging of
RelWithDebInfo builds, since the debugger won't be able to reliably
inspect variables or control flow. These hints are better placed on an
individual basis after identifying bottlenecks in a profiler.
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
guest instruction
Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.
- Instead of using RunUntilExit, all threads can use `ExecuteThread`
directly, since there's nothing special about the primary thread now.
- This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
`ExecuteThread` has returned.
- Just make gdbserver is cleaned up early if it exists since it may
want to send some things to the connected gdb instance before
threads are exited.