Commit Graph
68 Commits
Author SHA1 Message Date
Jacek Caban 2158309c51 InterpreterFallbacks: Remove unused template
Fixes -Wunused-template warning.
2026-08-10 15:05:26 +02:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek c6031a7806 Softfloat: Remove weird tail padding from X80SoftFloat
These are expected to match x87 registers in side. This was always a bit
weird.
2026-01-28 19:40:29 -08:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Paulo Matos e7a47a647c Fix quiet and signalling nan propagation
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.

Fix for __builtin_issignaling() test of SPEC2017 classify test.
2025-09-23 13:48:45 +02:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Lioncache 5ebdc01783 F80Fallbacks: Remove unnecessary memset
X80SoftFloat instances already initialize zeroed out in the default constructor.
2025-09-02 12:25:03 -04:00
Ryan Houdek c9a33c638a Dispatcher: Most minor of optimizations for f64 x87
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
2025-08-17 19:50:21 -07:00
Ryan Houdek 153d20ca59 Rename SHMStats 2025-07-29 12:02:19 -07:00
Paulo Matos 726656d0bf Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:10 +02:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek fc052efb91 FEXCore: Fixes x87 reduced precision
With the change from #4538 I had accidentally broken x87 reduced
precision.

This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.

Fixes Steam when x87 reduced precision is enabled.
2025-04-30 17:37:10 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Ryan Houdek 32764ddf81 JIT: Moves VPCMPESTRX handler to use vectors
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.

This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
2025-04-24 10:09:26 -07:00
Ryan Houdek 031afbfb18 Various: Adds missing SPDX file headers
NFC
2025-03-07 11:34:31 -08:00
Ryan Houdek 0f5aff73ab Softfloat: Define INLINE
This is embarassing. We were throwing away performance by failing to
use the softfloat library's inline helpers.

Turns out we needed to define `INLINE` to something in order for them to
work.

Feels bad.
2025-03-07 01:56:13 -08:00
Ryan Houdek a32b892787 FEXCore/Profiler: Implement support for JIT float fallbacks
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.

Performance-wise this change falls within noise of my x87 microbench.

This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
2025-02-11 14:41:16 -08:00
Ryan Houdek 3160e0a430 FEXCore: Keep PCMPISTRI arguments in vectors longer
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
2025-02-11 12:57:11 -08:00
LC 2943cff73f Merge pull request #4345 from Sonicadvance1/optimize_pcmpistri
FEXCore: Optimize VPCMPISTRX implicit length calculation
2025-02-11 12:50:55 -05:00
Ryan Houdek d46722a95c FEXCore: Optimize VPCMPISTRX implicit length calculation
With ASIMD this can be decently faster. With my microbenchmark this
makes pcmpistri ~6% faster.

With #4324 this can be made even faster since the incoming data can stay
in vector registers; Removing some overhead of umov.
2025-02-10 20:17:29 -08:00
Ryan Houdek 6abf5b90b7 FEXCore/JIT: Pass Softfloat arguments as vector registers
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.

Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.

A more tangible win from this change alone is less codegen on the JIT
side.
2025-02-10 12:53:25 -08:00
Ryan Houdek e3d7161ac5 FEXCore: Override x87 precision control when necessary
According to the documentation for x87 FCW precision control, this only
affects fadd*, fsub*, fmul*, fdiv*, and fsqrt. FEX was incorrectly
reducing precision for all x87 operations.

Precision is ignored for the following x87 ALU operations:
- fabs
- fscale
- fprem{1,}
- fcos
- fsin
- ftan
- fyl2x
- fyl2xp1
- fpatan
- fsincos
- Plus any operations just doing data movement and conversions

Next commit adds unittests to ensure this is correct for each
instruction.
2024-12-04 23:47:19 -08:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
Ryan Houdek 460a21625e JIT: Remove implicit OpSize conversions
NFC
2024-10-28 19:48:40 -07:00
Paulo Matos f35a212f06 Fix FScale'ing zero that should return zero
Added tests will fail without the current patch.
This will properly check if we return the correct sign of zero.
2024-09-26 17:57:39 +02:00
Ryan Houdek c3c0bbb060 F80Fallbacks: Fix Uninitialized variable warning that can't occur 2024-09-08 18:03:20 -07:00
Billy Laws 696503680a F80: Drop dependency on state stored in TLS
Windows cannot support the implicit TLS as was used prior, so introduce
a state structure and pass it in to functions where necessary.
2024-07-31 18:51:42 +01:00
Alyssa Rosenzweig e7d5a01c5f IR: remove F80Cmp flags
nothing is optimizing around this, it's just adding pointless complexity. if we
want to actually optimize F80Cmp, the right way would be to lift the
implementation into the OpcodeDispatcher or JIT. it wouldn't be terribly
difficult. This kludge doesn't get us closer there.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Ryan Houdek 8955f83ef6 Softfloat: Fixes Integer indefinite return for 16-bit signed values
Regardless of positive or negative value, if the converted integer
doesn't fit in to the converted int16_t then it returns INT16_MIN.
2024-07-04 17:43:28 -07:00
Alyssa Rosenzweig 0c042d1e85 VectorFallbacks: optimize PCMP*STRI flags
Return an NZCV.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:19:43 -04:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Paulo Matos 6524716404 Move const to the left in preparation for reformatting
clang-format-16 had some issues with const placement, so we are manually changing these.
2024-04-12 16:06:57 +02:00
Ryan Houdek ed3af580c5 FEXCore: Move nearly all IR definitions to internal
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.

Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.

No functional change with this, just moving headers around.
2024-03-29 17:20:18 -07:00
Ryan Houdek 0eed73beeb HostFeatures: Supports runtime disabling of preserve_all
This is used for instcountci to ensure instruction counts don't change
when a compiler supports this feature or not. Always runtime disable
when running in instcountci.

CMake option from #3394 can still be useful so leaving that in place.
2024-02-02 08:59:04 -08:00
Ryan Houdek 31564354b1 FEXCore: Removes vestigial Interpreter code 2023-09-21 15:49:49 -07:00
Ryan Houdek fea72ce19c Merge pull request #3120 from Sonicadvance1/more_optimal_x87
FEXCore: Support preserve_all ABI for interpreter fallbacks
2023-09-21 15:35:37 -07:00
Alyssa Rosenzweig c52741c813 FEXCore: Gut interpreter
It is scarcely used today, and like the x86 jit, it is a significant
maintainence burden complicating work on FEXCore and arm64 optimization. Remove
it, bringing us down to 2 backends.

1 down, 1 to go.

Some interpreter scaffolding remains for x87 fallbacks. That is not a problem
here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 12:48:12 -04:00
Ryan Houdek d588d41ab9 InterpreterFallbacks: Converts X87 and String ops to preserve_all
This improves performance!
2023-09-20 18:51:18 -07:00
Ryan Houdek 8aa8d597f6 Arm64: Supports jumping out of the JIT with preserve_all ABI
This improves perferformance when jumping out of the Arm64 JIT by
reducing the number of registers we need to save.
2023-09-20 18:51:18 -07:00
Ryan Houdek 1032224d62 FEXCore/Interface/Core/Interpreter: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Alyssa Rosenzweig fc02f38435 IR: Only invert CF for NZCV if needed
If we are going to throw away the updated value of CF anyway there is no point
wasting an instruction to invert CF. Add an IR toggle for that so the arm64 JIT
can make better choices.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:08:22 -04:00
Ryan Houdek d5782567e8 Merge pull request #3077 from Sonicadvance1/x86_shifted
FEXCore: Implements support for shifted bitwise ops
2023-09-15 08:09:35 -07:00
Ryan Houdek 9866e238d5 Merge pull request #3080 from Sonicadvance1/defer_softfloat
FEXCore: Defer setting x87 softflow rounding mode until use
2023-09-15 08:08:04 -07:00
Ryan Houdek 90ddee5f8d IR: Implements support for arm64 bfxil
This is useful for extracting a width from a register and inserting in
to the lower bits of a destination.
2023-09-14 15:52:55 -07:00
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Ryan Houdek 76bd81af15 FEXCore: Defer setting x87 softflow rounding mode until use
Currently FEX will always jump out of the JIT any time FCW was getting
written to, ensuring that the softfloat state is setup to rounding at
the time of FCW getting written.
This has the unintended side-effect that even in "x87 reduced precision"
mode we were jumping out of the JIT.
This hit a real world use case of an installer reloading FCW after every
x87 operation and generating a block with 2297 instructions.

Instead when jumping out of the JIT for handling x87 operations, load
FCW and pass it as the first argument of the handler. Setting the
softfloat state at that point.

This helps the installer's hottest block by cutting it down to 1477
instructions. 64.3% of the original size. The code block is still
burning 90% of the CPU time of the installer but the performance is
significantly better while it is doing its decompression.

In order to optimize this installer's block of code more then we will
likely need to optimize out x87 stack usage.
2023-09-12 05:21:06 -07:00
Ryan Houdek 863331b117 FEXCore: Implements support for shifted bitwise ops
This wasn't implemented initially for the interpreter and x86 JIT.

This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.

The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
2023-09-11 13:17:35 -07:00