Commit Graph
67 Commits
Author SHA1 Message Date
Ryan Houdek 37eaff93cb FEXCore: Migrate internal DiskCache header to internal
Leave the smaller public header that is required to get the current
DiskCache version variable still.

NFC
2026-09-15 19:43:27 -07:00
Justin Becker 89a13cd5fd AVX-VNNI 2026-09-10 15:56:03 -07:00
Ryan Houdek 66cad978c3 FEXCore: Removes syscall optimization
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.

Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.

This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.

Bumps the DiskCache version again because it causes codegen to change.
2026-08-31 19:23:18 -07:00
Ryan Houdek f5935b4006 InstCountCI: Enforce DiskCache version matching
Two added failure modes here. If the disk cache version has changed then
the json files must be updated to the new version to ensure correct
tracking.

Additional failure mode is that codegen actually changed but the disk
cache version hasn't. This is the expected common failure mode and we
need to do additional work before updating json results. This just means
incrementing the disk cache version before updating the instcountci
results. Have a fairly lengthy error message to showcase how much of an
impact this might have.
2026-08-30 14:26:43 -07:00
Ryan Houdek b83dd97762 FEXCore/Context: Removes InitialRIP/RSP from CreateThread
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.

Just a smidge of cleanup, NFC.
2026-08-24 18:32:35 -07:00
Lioncache e862f8f86c unittests: Add specific paths for MOPS 2026-03-19 11:55:02 -04:00
Billy Laws 86211e18d7 ALookupExecutableFileSection: Take thread argument as an optional pointer 2025-12-23 23:44:58 +00:00
Billy Laws 8c00ac78b1 Linux: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek 61719115e5 Merge pull request #4817 from neobrain/refactor_code_cache_new_interfaces
CodeCache: Introduce new interfaces
2025-09-11 12:59:33 -07:00
Tony Wasserka 0749477eb9 CodeCache: Introduce revamped interfaces 2025-09-11 17:03:50 +02:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Tony Wasserka 1dc560d45a InstCountCI: Explicitly disable TSO by default 2025-09-10 15:36:23 +02:00
Ryan Houdek 80cacc462e X86Tables: Move x87 tables to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 1b0b4ba416 InstCountCI: Add frontend handling of GDT 2025-08-18 14:03:51 -07:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Billy Laws 2e5c6283f5 CodeSizeValidation: Add dummy QueryGuestExecutableRange impl 2025-07-10 16:00:24 +01:00
Ryan Houdek 4a74bea7ab InstcountCI: Don't use a global static initializer for CodeSizeValidation
Relies on fmt facet initialization order which isn't guaranteed to have
correct initialization order.

CID 482003
2025-06-19 16:51:26 -07:00
Tony Wasserka 23b69271eb Use consistent log message formatting for all modules 2025-06-16 13:54:03 +02:00
Ryan Houdek f6b4c76d76 InstcountCI: Adds tests for instructions discovered by #4597
Apparently I completely missed that cpuid, xgetbv, syscall,
l{u,}{div,rem} were failing to hit their optimized cases for inlining
and avoiding 128-bit software divide.

The divisions are a clear performance regression for 64-bit applications
since that is the only real way to do a 64-bit division on x86, I added
those specifically because it sped up games.

CPUID depends heavily on the game, since some games use that as a
serialization instruction fairly heavily.

XGETBV is trivial since it matches behaviour of CPUID (and is basically
an extension of it).

Syscall inlining can save a decent amount of time, again heavily depends
on game.

Adds multi-inst tests for all of these situations so that once it gets
fixed (Apparently broken once RCLSE got stripped out), we can see that
they keep working. Obviously tests couldn't have existed in instcountCI
before since we didn't support multi-instruction tests.
2025-06-02 12:10:39 -07:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Ryan Houdek d41374c55a InstcountCI: Ensure RIP of blocks is consistent
This changes the instcountCI code to consistently load test data in to
RIP 0x1'0000 so we don't have any spurious changes due to json changes.

This has been a minor annoyance where if a test was added, it had the
potential to shift the rest of the data in the tests. This now ensures
it is consistent.
2025-04-05 15:55:08 -07:00
Alyssa Rosenzweig 166a7c7e53 FEXCore: plumb Feat_FRINTTS
we want these instructions to accelerate conversions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:01:32 -04:00
Ryan Houdek 9bf47b3f23 Convert config options once
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.

Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
2025-03-22 17:32:10 -07:00
Ryan Houdek 7cd4d53fa9 Fixes some instances of auto usage with unintentional copy
Just switch the uses over to `const auto&`
2025-02-12 23:51:57 -08:00
Ryan Houdek 800d447f3d InstCountCI: Add support for TSO and LRCPC1/2 2024-12-11 14:55:19 -08:00
Ryan Houdek ae9db336e7 FEXCore/Context: Removes unused features
No functional change here.
- CoreRunningMode enum and variable wasn't used anymore.
   - Code was moved to the frontend
- CustomCPUFactory wasn't used anymore
   - All special signal handling and various features were moved to
     TestHarnessRunner
   - We also don't want to support actual custom CPU cores.
   - TestHarnessRunner just runs as a host runner if compiled on an
     x86-64 device if vixl sim isn't enabled now.
   - Removes the Core config option entirely.
- Moves VDSOPointers struct to the frontend
   - Every use of this lives in the Linux frontend instead now
2024-09-09 18:38:33 -07:00
Ryan Houdek 2f8c5b4820 FEXCore: Pass HostFeatures in to CreateNewContext directly
The class constructor for ContextImpl::CPUID requires HostFeatures to be
available at construction time. Pass the host features struct directly
through during construction time instead, which cleans up the interface
slightly and fixes that issue.
2024-08-08 21:02:41 -07:00
Ryan Houdek e84848b16b FEX: Moves HostFeatures querying to the frontend
This moves the CPU feature querying to the frontend. The primary purpose
here is for the wow64 frontend to not require linux-isms for querying
these features. This is required since non-Linux environments don't have
the "CPUID" feature for reading EL1 MSRs in EL0.

Wiring up the remaining wow64 registry querying is left for a future
exercise.

This also technically removes an xbyak requirement from FEXCore for when
building the x86 Test harness runner, but that doesn't really matter for
regular use cases.
2024-08-07 05:26:02 -07:00
Ryan Houdek 5e56bdc0fd InstcountCI: Add support for SVE bitperm 2024-07-10 21:48:37 -07:00
Ryan Houdek ce4b252e5c InstCountCI: Stop disabling AVX if SVE256 is disabled. 2024-06-26 15:06:03 -07:00
Ryan Houdek ac1a096bae InstCountCI: Hardcode the offset to load tests into
Depending on where the assembly was getting loaded in to memory it was
causing slight code generation differences.

Map the entire file to the same fixed offset as our ASM tests to ensure
consistency and removing flakes in CI.
2024-05-18 17:00:28 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Alyssa Rosenzweig b50292493a InstCountCI: enable preserve_all ABI
This is what we'll actually ship (I hope), so that's the config we want to
track long-term. It's also a lot more managable resulting asm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-27 12:03:58 -04:00
Ryan Houdek 2480bab409 Fixes one mutex hang
When code invalidation is happening we currently have the issue that a
thread can acquire the code invalidation mutex in the middle of
invalidation. This is due to us acquiring and releasing the mutex
between each thread's code invalidation.

We need to hold the mutex for the entire duration for all thread's code
invalidation.
This fixes a rare hang on proton startup and resolves a consistent hang
on Proton application shutdown.

This now puts us on par with FEX-2312.1 with hanging.

This does not fix a relatively rare hang on fork (which also existed with FEX-2312.1).

This also does not fix the issue that the intersection of our mutexes
between frontend and backend are very convoluted. In part of the work
that is going to fix the rare fork mutex hang will change more of this.
2024-02-08 18:18:00 -08:00
Alyssa Rosenzweig 235f32ce8c Merge pull request #3401 from Sonicadvance1/runtime_preserve_all
HostFeatures: Supports runtime disabling of preserve_all
2024-02-05 15:34:46 -04:00
Ryan Houdek c437129ed8 Revert "Revert "FEXLoader: Moves thread management to the frontend""
This reverts commit 5358af7794.
2024-02-03 00:57:36 -08:00
Ryan Houdek 0eed73beeb HostFeatures: Supports runtime disabling of preserve_all
This is used for instcountci to ensure instruction counts don't change
when a compiler supports this feature or not. Always runtime disable
when running in instcountci.

CMake option from #3394 can still be useful so leaving that in place.
2024-02-02 08:59:04 -08:00
Ryan Houdek 36250b10f6 InstCountCI: Sanitize out adrp and adr
First time usage of adrp and adr, need to sanitize it.
2024-01-29 19:32:04 -08:00
Ryan Houdek 5358af7794 Revert "FEXLoader: Moves thread management to the frontend"
This reverts commit 58f2693954.
2023-12-27 04:33:50 -08:00
Ryan Houdek 58f2693954 FEXLoader: Moves thread management to the frontend
Lots going on here.

This moves OS thread object lifetime management and internal thread
state lifetime management to the frontend. This causes a bunch of thread
handling to move from the FEXCore Context to the frontend.

Looking at `FEXCore/include/FEXCore/Core/Context.h` really shows how
much of the API has moved to the frontend that FEXCore no longer needs
to manage. Primarily this makes FEXCore itself no longer need to care
about most of the management of the emulation state.

A large amount of the behaviour moved wholesale from Core.cpp to
LinuxEmulation's ThreadManager.cpp. Which this manages the lifetimes of
both the OS threads and the FEXCore thread state objects.

One feature lost was the instruction capability, but this was already
buggy and is going to be rewritten/fixed when gdbserver work continues.

Now that all of this management is moved to the frontend, the gdbserver
can start improving since it can start managing all thread state
directly.
2023-12-19 17:43:04 -08:00
Ryan Houdek aa2e8704bc FEXCore: Changes ParentThread ownership from the CTX to the frontend, take 2
Similar to #3284 but works around some of the bugs that one introduced.

This is the minimal amount of changes to move the ownership from FEXCore
to the frontend. Since the frontends don't yet have a full thread state
tracking, there is an opaque pointer that needs to be managed.

In the followup commits this will be changed to have the syscall handler
to be the thread object manager.
2023-12-18 14:54:07 -08:00
Ryan Houdek f090700184 FEXCore: Removes InitializeContext API
This isn't necessary anymore, just initialize everything on context
creation immediately. All use cases just called this immediately
afterwards.
2023-11-29 09:33:32 -08:00
Ryan Houdek ba1632974e InstCountCI: Actually allow disabling crypto
Forcing it was required to work around initial simulator quirks
2023-11-13 18:38:02 -08:00
Alyssa Rosenzweig bd4464bd5e InstructionCountCI: Remove Optimal flags
Instruction count CI has transformed the way we work on FEX… I love the system
and want to make it better. there’s one part of instruction count CI that isn’t
so lovable: the problematic “optimal” flag on instructions.

There are several issues with this flag, both philosophical and practical.

– it is tedious to update the optimal flag when making an implementation
optimal. The effect of that is discouraging people from making instructions,
optimal, or encouraging people to fail to update the flag, and dilute the value
of it. Either way, since we care far more about optimal implementations, then we
do about updating the flag, clearly we should prioritize the implementation and
not the flag. This issue was not obvious at the outset, when instruction count,
CI was introduced, and still quite small. The problem magnified when we started
duplicating instructions in bulk for different combinations of CPU features
(flagm, AFP, etc.) that intern multiplies the manual work required to update the
flags by the corresponding constant factor. if it comes down to a choice between
removing this extra coverage and removing the flag, I think we all agree that
removing the flag is the lesser evil.

– The definition of “optimal” is fundamentally problematic. I have often
improved the instruction count of an instruction that was already “optimal”.
This is all kinds of silly, and calls into question whether there’s any value
whatsoever in the existing classifications of the flag. Furthermore, it is often
unknowable, whether an implementation really is optimal. Is it possible to
implement BZHI (with flag calculations) in fewer than eight instructions? We
don’t know, and it’s silly to pretend that we do.

– as a consequence of the problematic definitions , there are so many errors in
both directions that I don’t think there’s much value in preserving the existing
classification at the expense of +progress. Being able to say “32% of
instructions are translated optimally” is neat, but it really doesn’t tell us
anything whatsoever when you dig a little deeper.

So, as the flag is misleading at best and perhaps harmful at worst, let’s remove
it and make the instruction count CI, more useful overall. let’s let the
expected count and the assembly speak for themselves, and cut away the chaff. if
we want a meaningless number to report to management, we can instead calculate
the average blowup factor ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:14:05 -04:00
Ryan Houdek 4edd72fc33 InstCountCI: Support disabling flagm extensions
This is necessary so #3162 can give consistent results
2023-10-23 14:02:24 -07:00
Ryan Houdek cd83d3eb24 InstCountCI: Support multiple instructions in the tests
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.

eg:
```asm
; Can be optimized to a single stp
push eax
push ebx

; Can remove half of the copy since we know the direction
cld
rep movsb

; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```

This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.

There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.

Example in the json
```json
"push ax, bx": {
  "ExpectedInstructionCount": 4,
  "Optimal": "No",
  "Comment": "0x50",
  "x86Insts": [
    "push ax",
    "push bx"
  ],
  "ExpectedArm64ASM": [
    "uxth w20, w4",
    "strh w20, [x8, #-2]!",
    "uxth w20, w7",
    "strh w20, [x8, #-2]!"
  ]
}
```
2023-10-09 21:49:53 -07:00
Ryan Houdek 22590dde77 FEXCore: Implements support for RPRES
This allows us to use reciprocal instructions which matches precision of
what x86 expects rather than converting everything to float divides.

Currently no hardware supports this, and even the upcoming X4/A720/A520
won't support it, but it was trivial to implement so wire it up.
2023-10-07 23:13:47 -07:00
Ryan Houdek 559cf6491a InstCountCI: Support overriding AFP features
Also disable AFP under the vixl simulator by default since it doesn't support it.
2023-10-07 11:48:42 -07:00
Ryan Houdek 5b7ba06d5c FEXCore: Support crypto extensions in HostFeatures override
Enables in InstCountCI so Pi users can run InstCountCI can run the tests
without breaking on crypto operations.

When crypto is enabled or disabled just wholesale change AES, CRC32, and
PMULL 128-bit in one step. We don't really care about partial support
here.
2023-10-05 17:41:08 -07:00