Commit Graph
342 Commits
Author SHA1 Message Date
Ryan Houdek 369ca5cb72 Merge pull request #4719 from Sonicadvance1/runtime_mode_switch_take2
Runtime mode switch take 2
2025-07-30 11:47:54 -07:00
Ryan Houdek aa871c797b FEXCore: Accurately store segment descriptors
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.

In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.

Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
2025-07-29 12:02:37 -07:00
Ryan Houdek 153d20ca59 Rename SHMStats 2025-07-29 12:02:19 -07:00
Ryan Houdek 392fa62dae Profiler: Decouple profile stats from the profiler option
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
2025-07-29 12:02:19 -07:00
Ryan Houdek 126c4efa0e FEXCore: Remove unused refcount_shared_mutex 2025-07-28 15:43:44 -07:00
Ryan Houdek b0c61e2b69 ArchHelpers: Remove pair usage in unaligned handler
It's being treated like optional, where a value means it has been
handled, and no value means it hasn't been handled. Stop using pair in
this case.
2025-07-25 13:01:18 -07:00
Billy Laws d67b1645c9 FEXCore: Hold a frontend allocation for the call-ret stack
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
2025-07-24 14:53:09 +01:00
Billy Laws 963a8c2f08 FEXCore: Save and restore the call/ret SP from CPUState
For simplicity in cases like signal handling, always load it in
fill and store in spill, even the SP is stored in a callee save
register.
2025-07-24 14:53:09 +01:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Paulo Matos 5267cde60e Whole-tree reformat with clang-format-19 2025-07-17 08:10:00 +02:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Tony Wasserka 648c726bb1 LogManager: Use assert log level for ERROR_AND_DIE_FMT 2025-07-15 16:55:18 +02:00
Billy Laws 6eaaab8dc2 IntervalList: Also return the full matching interval on query 2025-07-10 16:51:34 +01:00
Billy Laws 14eed89bb6 SyscallHandler: Add method to query executable memory ranges 2025-07-10 16:00:24 +01:00
Billy Laws 407c5a0f78 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-08 23:38:43 +01:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Tony Wasserka 57627d4fcf LogManager: Drop unused STDOUT/STDERR log levels 2025-06-16 13:54:03 +02:00
Tony Wasserka 3e85e60a30 LogManager: Use colors for logging when possible 2025-06-13 15:14:47 +02:00
Tony Wasserka 992d86bbc1 LogManager: Shorten debug level strings to a single letter 2025-06-13 14:31:25 +02:00
Tony Wasserka 4078840ef1 Core: Reduce JIT time by sharing CodeBuffers between threads
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
  be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
  in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
  the modifying thread. Old versions of the CodeBuffer persist as read-only
  data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
  decrease the reference count and eventually trigger deallocation of the
  old version
2025-06-01 22:44:49 +02:00
Tony Wasserka 1837aaabe4 fextl: Add shared_ptr and make_shared 2025-06-01 22:42:55 +02:00
Tony Wasserka 2a1d29d2df Config: Clean up use of templates 2025-05-29 18:38:35 +02:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
Tony Wasserka 2e24ee7a5f LogManager: Print source location when failing assertions 2025-05-24 09:35:11 +02:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Ryan Houdek 6b85fa5611 SignalScopeGuards: Review 2025-05-14 10:56:44 -07:00
Ryan Houdek 395d870814 SignalScopeGuards: Add checks for locks being held by the calling thread
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
2025-05-14 10:56:44 -07:00
Tony Wasserka d4fdb28e72 PoolBufferWithTimedRetirement: Add support for updating the buffer size
This should be done at low frequency since it may unclaim the buffer.
2025-05-14 13:37:46 +02:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
Tony Wasserka f41501444d ThreadPoolAllocator: Add a dedicated interface to try reowning a buffer without fallback 2025-05-14 13:29:44 +02:00
Tony Wasserka 86b26b80ce FixedSizePooledAllocation: Clean up documentation 2025-05-14 13:29:44 +02:00
Tony Wasserka c19119bcd6 ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
Default-constructed iterators can't be copied.
2025-05-09 15:02:49 +02:00
Ryan Houdek 92a82c3134 FEXCore: Removes unused argument on CreateThread
ParentTID is purely a Linux construct and has been moved entirely to the
frontend at this point. Remove this argument which is now unused.
2025-05-05 11:17:02 -07:00
Ryan Houdek 7ed17f68c5 FEXCore: Remove unused InvalidateGuestCodeRange with callback 2025-05-02 01:15:50 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Ryan Houdek bfee39ae70 HostFeatures: Passthrough if the host supports ECV 2025-04-15 15:38:00 -07:00
Billy Laws 039aaae041 FEXCore: Add a pre-compilation frontend callback to SyscallHandler 2025-04-11 12:07:16 +01:00
Billy Laws 4fca8fb8e6 AllocatorHooks: Correctly restore the region protection in VirtualDontNeed 2025-04-11 12:07:16 +01:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek 8aecdc536c Merge pull request #4471 from alyssarosenzweig/opt/cvtss2si
Optimize float->integer conversions with Feat_FRINTTS
2025-04-01 08:52:50 -07:00
Alyssa Rosenzweig 166a7c7e53 FEXCore: plumb Feat_FRINTTS
we want these instructions to accelerate conversions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:01:32 -04:00
Ryan Houdek f9b369c550 Telemetry: Removes unnecessary indirection
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.

Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).

This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
2025-03-29 15:10:37 -07:00
Alyssa Rosenzweig 42ea711850 CoreState: squish and rearrange pf_raw/af_raw
to allow next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-27 11:04:08 -04:00
Tony Wasserka 7efd827e78 Merge pull request #4378 from bylaws/volmd
Implement PE volatile metadata support
2025-03-25 10:23:53 +01:00
Billy Laws 642903a7bf FEXCore: Support tracking TSO range information 2025-03-24 22:01:49 +00:00
Billy Laws f51fd6c78d Move IntervalList to FEXCore 2025-03-24 22:01:49 +00:00
Ryan Houdek 9bf47b3f23 Convert config options once
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.

Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
2025-03-22 17:32:10 -07:00
Ryan Houdek b0b41d00ee Various: More static analysis warnings cleanup
NFC
2025-03-12 17:27:41 -07:00