Commit Graph
321 Commits
Author SHA1 Message Date
Tony Wasserka 2a1d29d2df Config: Clean up use of templates 2025-05-29 18:38:35 +02:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
Tony Wasserka 2e24ee7a5f LogManager: Print source location when failing assertions 2025-05-24 09:35:11 +02:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Ryan Houdek 6b85fa5611 SignalScopeGuards: Review 2025-05-14 10:56:44 -07:00
Ryan Houdek 395d870814 SignalScopeGuards: Add checks for locks being held by the calling thread
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
2025-05-14 10:56:44 -07:00
Tony Wasserka d4fdb28e72 PoolBufferWithTimedRetirement: Add support for updating the buffer size
This should be done at low frequency since it may unclaim the buffer.
2025-05-14 13:37:46 +02:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
Tony Wasserka f41501444d ThreadPoolAllocator: Add a dedicated interface to try reowning a buffer without fallback 2025-05-14 13:29:44 +02:00
Tony Wasserka 86b26b80ce FixedSizePooledAllocation: Clean up documentation 2025-05-14 13:29:44 +02:00
Tony Wasserka c19119bcd6 ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
Default-constructed iterators can't be copied.
2025-05-09 15:02:49 +02:00
Ryan Houdek 92a82c3134 FEXCore: Removes unused argument on CreateThread
ParentTID is purely a Linux construct and has been moved entirely to the
frontend at this point. Remove this argument which is now unused.
2025-05-05 11:17:02 -07:00
Ryan Houdek 7ed17f68c5 FEXCore: Remove unused InvalidateGuestCodeRange with callback 2025-05-02 01:15:50 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Ryan Houdek bfee39ae70 HostFeatures: Passthrough if the host supports ECV 2025-04-15 15:38:00 -07:00
Billy Laws 039aaae041 FEXCore: Add a pre-compilation frontend callback to SyscallHandler 2025-04-11 12:07:16 +01:00
Billy Laws 4fca8fb8e6 AllocatorHooks: Correctly restore the region protection in VirtualDontNeed 2025-04-11 12:07:16 +01:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek 8aecdc536c Merge pull request #4471 from alyssarosenzweig/opt/cvtss2si
Optimize float->integer conversions with Feat_FRINTTS
2025-04-01 08:52:50 -07:00
Alyssa Rosenzweig 166a7c7e53 FEXCore: plumb Feat_FRINTTS
we want these instructions to accelerate conversions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:01:32 -04:00
Ryan Houdek f9b369c550 Telemetry: Removes unnecessary indirection
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.

Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).

This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
2025-03-29 15:10:37 -07:00
Alyssa Rosenzweig 42ea711850 CoreState: squish and rearrange pf_raw/af_raw
to allow next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:04:08 -04:00
Tony Wasserka 7efd827e78 Merge pull request #4378 from bylaws/volmd
Implement PE volatile metadata support
2025-03-25 10:23:53 +01:00
Billy Laws 642903a7bf FEXCore: Support tracking TSO range information 2025-03-24 22:01:49 +00:00
Billy Laws f51fd6c78d Move IntervalList to FEXCore 2025-03-24 22:01:49 +00:00
Ryan Houdek 9bf47b3f23 Convert config options once
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.

Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
2025-03-22 17:32:10 -07:00
Ryan Houdek b0b41d00ee Various: More static analysis warnings cleanup
NFC
2025-03-12 17:27:41 -07:00
Ryan Houdek c9eee9bf7f Config: Stop using config values with list when unnecessary
With the previous fixes in place, we can now stop burning a fextl::list
in every single config option. This list is only required for strarray
options so reserve it for those entirely.

We also don't need to save the config option enum for each, so these
actually go from ~32 bytes per object down to their base type for most
everything.
2025-03-05 11:17:09 -08:00
Ryan Houdek de6931b1f5 OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
```diff
     "sha1rnds4 xmm0, xmm1, 10b": {
-      "ExpectedInstructionCount": 55,
+      "ExpectedInstructionCount": 10,
```

So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
2025-02-22 04:57:58 -08:00
Tony Wasserka 391f9aa97d Profiler: Add Tracy backend
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only

To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.

Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
2025-02-12 19:35:15 +01:00
Ryan Houdek a32b892787 FEXCore/Profiler: Implement support for JIT float fallbacks
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.

Performance-wise this change falls within noise of my x87 microbench.

This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
2025-02-11 14:41:16 -08:00
Ryan Houdek c8c27f26f7 Review 2025-02-11 13:42:45 -08:00
Ryan Houdek b4c47a3d24 FEXCore: Implements baseline per-thread profile stats
Not wired up, just the definitions so it lives in the
InternalThreadState.

We want this accessible from both FEXCore and the frontends so it needs
to live there.

Two types of events supported. Scoped cyclecounts and instant
increments.

This gives us JIT time and Signal handling time, plus events for number
of SIGBUS and number of SMC events.

All useful statistics for seeing stutter live.
2025-02-11 12:56:10 -08:00
Tony Wasserka 44bc3fb90b fextl: Add std::move_only_function replacement 2025-02-06 22:30:45 +01:00
Ryan Houdek 11ce97655b Allocator: Still need to return memory regions to frontend 2025-01-28 16:07:51 -08:00
Ryan Houdek c75778abeb FEX: Allocate a VMA allocator when running on a 48-bit VA
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.

This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.

Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.

This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.

On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.

Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.

This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.

Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
2025-01-28 16:07:51 -08:00
Tony Wasserka 229e7c5b61 LogManager: Remove assuming assert macros
Placing optimization hints everywhere interferes with debugging of
RelWithDebInfo builds, since the debugger won't be able to reliably
inspect variables or control flow. These hints are better placed on an
individual basis after identifying bottlenecks in a profiler.
2025-01-21 12:01:33 +01:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Ryan Houdek c16bf09310 Profiler: Setup for usage on Windows
This will get gpuviz working under Wine.
2025-01-09 14:25:48 -08:00
Ryan Houdek a47ed105e7 x87StackOptimizationPass: Minor opt to f80 fchs and fabs
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
2025-01-03 13:47:10 -08:00
Billy Laws efd6e95059 OpcodeDispatcher: Match x86 overflow behaviour for F2I conversions
ARM behaviour here is to saturate on overflow or NaN inputs, whereas
X86 returns a sentinel value of 2^(bitsize-1), explicitly emulate this.
2024-12-30 00:42:55 +00:00
Billy Laws ec003281be External: Update bundled libfmt 2024-12-18 15:25:45 +00:00
Billy Laws d5d7eec8b0 FEXCore: Expose an API to check if the current block represents a single
guest instruction

Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
2024-12-12 21:28:37 +00:00
Billy Laws 5337b9537d FEXCore: Expose an API to query intersection with the current block
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
2024-12-12 21:28:37 +00:00
Ryan Houdek e88c92de57 Merge pull request #4161 from bylaws/tf
FEXCore: Emulate EFLAGS.TF
2024-12-12 11:51:53 -08:00
Billy Laws 981c3009ee FEXCore: Emulate EFLAGS.TF
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
2024-12-10 15:20:47 +00:00
Ryan Houdek 2533ed4a63 Context: Constify GPRs passed to ReconstructCompactedEFLAGS
This only reads the GPRs passed in, doesn't modify it.
2024-12-09 15:08:59 -08:00
Ryan Houdek efb276f489 FEXCore: Removes ExitHandler and RunUntilExit
Now that all the threading behaviour has been correctly separated/moved
to the frontend, these functions serve no purpose.

- Instead of using RunUntilExit, all threads can use `ExecuteThread`
  directly, since there's nothing special about the primary thread now.
  - This also removes the public function definition of `ExecutionThread` since that was only used for threading logic.
- Instead of using an exit handler, just do the same cleanup after
  `ExecuteThread` has returned.
  - Just make gdbserver is cleaned up early if it exists since it may
    want to send some things to the connected gdb instance before
    threads are exited.
2024-12-01 10:45:38 -08:00