Commit Graph
48 Commits
Author SHA1 Message Date
LC 9b8ae25491 Common/BitSet: Amend byte size retrieval
This needs to divide by 8 to get a proper byte size for all type sizes.
The only usage of this is currently a uint64_t, so it worked by
coincidence, since sizeof(uint64_t) == 8.
2026-07-13 08:17:27 -04:00
LC 46f3bec37e Common/BitSet: Ensure internal pointer is always initialized
Provides deterministic state.
2026-07-10 10:31:45 -04:00
LC f5d2e0db29 Common/BitSet: Mark getters as const
These don't modify internal state.
2026-07-10 10:31:05 -04:00
LC 88ee56f471 Common/BitSet: Amend Clear() behavior
Ensures the bits are actually being unset.
2026-07-10 10:26:07 -04:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Tony Wasserka ea45f9c694 FEXCore/VectorRegType: Use vector_size on GCC
GCC does not support neon_vector_type and silently ignores that attribute,
but vector_size(16) seems to have the same effect.
2026-03-16 19:15:05 +01:00
Ryan Houdek c6031a7806 Softfloat: Remove weird tail padding from X80SoftFloat
These are expected to match x87 registers in side. This was always a bit
weird.
2026-01-28 19:40:29 -08:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Ryan Houdek 3a9b801400 FEXCore/VectorRegType: Trivial header fix 2025-11-05 09:47:02 +01:00
Ryan Houdek b748eab4ed FEXCore: Removes some hardcoded 4096 constants
Use our defined variable instead.
2025-10-21 09:18:06 -07:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Lioncache 0b238d942e Common/StringConv: std::stoull -> std::strtoull for enum handler
Missed this when simplifying the conversion facilities, but
we should be using std::strtoull here, as per the programming
concerns.
2025-10-05 14:49:07 -04:00
Lioncache 6a19c77184 General: Migrate to std::bit_cast
We already have a few cases where we already use bit_cast in the
emitter, so we may as well move all of our temporary helper instances
over as well.
2025-09-30 22:20:10 -04:00
Lioncache 7564b6c69a StringConv: Merge integral-handling facilities together
Same behavior, but we condense all of the integral handling
into one place.
2025-09-30 21:43:57 -04:00
Lioncache 773ddbf5c4 StringConv: Fix std::enable_if usage
This neglected to use the ::type qualifier, so this candidate was
always being considered in overload resolution.
2025-09-30 19:53:45 -04:00
Paulo Matos e7a47a647c Fix quiet and signalling nan propagation
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.

Fix for __builtin_issignaling() test of SPEC2017 classify test.
2025-09-23 13:48:45 +02:00
Lioncache 831b21cd39 JitSymbols: Remove unused includes
Reduces header dependencies (and clarifies existing ones).
2025-09-05 14:27:49 -04:00
Paulo Matos 3d92c1fb32 Fix IEEE 754 unordered comparison detection in x87 floating-point operations
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.

Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
2025-09-01 18:26:48 +02:00
Tony Wasserka db94c04179 FEXCore: Use inline functions instead of maybe_unused static ones 2025-08-25 10:36:01 +02:00
Paulo Matos 726656d0bf Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:10 +02:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek 2d7f37386e Softfloat: Remove warnings
These precision warnings are no longer true!
2025-04-07 15:22:24 -07:00
Ryan Houdek 0fbb6aa02f cephes: Rewrite to use softfloat-3e 128-bit
And also use it at the same time, since the function signatures changed.

Instead of relying on the host libc math libraries for `long double` ALU
operations, rewrite the entire thing to use softfloat-3e fixed width
float128_t types.

This is a very invasive change in cephes but is a necessary requirement
for getting the precision we require in environments that map `long
  double` to be the same as `double`, like Win32 and MacOS.

This fixes the precision issue in transcendental operations when running
under WINE.
2025-04-07 15:13:04 -07:00
Ryan Houdek d8cd807520 Softfloat-3e: Moves to Externals 2025-04-05 17:26:15 -07:00
Ryan Houdek ffca27cbde FEXCore/Softfloat: Wire up cephes math library for transcendental operations
This is solving a different problem than what #4411 is specifically
trying to solve.

For our transcendental operations, we can't currently guarantee that
these functions will actually operate at the 128-bit softfloat
precision. While this is true with glibc, this is /not/ true for musl
and likely more libraries.

Instead of relying on our libc implementation to implement these,
instead include the cephes math library directly which is what most
people use for this. Including musl even, but not for all operations.

With this we are no longer beholden to the standard libraries for
providing a correct implementation.
2025-04-04 16:03:09 -07:00
Ryan Houdek 9bf47b3f23 Convert config options once
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.

Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
2025-03-22 17:32:10 -07:00
Ryan Houdek 3160e0a430 FEXCore: Keep PCMPISTRI arguments in vectors longer
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
2025-02-11 12:57:11 -08:00
Ryan Houdek 6abf5b90b7 FEXCore/JIT: Pass Softfloat arguments as vector registers
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.

Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.

A more tangible win from this change alone is less codegen on the JIT
side.
2025-02-10 12:53:25 -08:00
Ryan Houdek 00aa4ddea0 FEXCore/Softfloat: Support loading and storing SoftFloat to vector registers 2025-02-10 12:53:25 -08:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Paulo Matos 10ec6b63b6 Fix FXTRACT for 0.0 and -0.0
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.

Also removed the IR node F80XTRACTStack which is not needed anymore.
2024-10-17 09:05:10 +02:00
Alyssa Rosenzweig aa3c963df1 SoftFloat: fix FPREM
fixes incorrect signs on FPREM. in turn should fix end-to-end failures logging
into Steam.

Thank you to Sergio Lopez for tracking down the JavaScript fail, and Ryan for
finding the bug.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-10-03 13:56:14 -07:00
Paulo Matos f35a212f06 Fix FScale'ing zero that should return zero
Added tests will fail without the current patch.
This will properly check if we return the correct sign of zero.
2024-09-26 17:57:39 +02:00
Ryan Houdek f09d511ac8 FEXCore: Move JITSymbolBuffer to internal header
This definition doesn't need to be exposed in the public API.
2024-09-05 13:28:45 -07:00
Billy Laws 696503680a F80: Drop dependency on state stored in TLS
Windows cannot support the implicit TLS as was used prior, so introduce
a state structure and pass it in to functions where necessary.
2024-07-31 18:51:42 +01:00
Ryan Houdek 8955f83ef6 Softfloat: Fixes Integer indefinite return for 16-bit signed values
Regardless of positive or negative value, if the converted integer
doesn't fit in to the converted int16_t then it returns INT16_MIN.
2024-07-04 17:43:28 -07:00
Alyssa Rosenzweig 9e1e602e09 BitSet: fix memset/memclear logic
Missing a factored of 4, causing a buffer overflow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:32:54 -04:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek bce694ebb5 FEXCore: Moves BitUtils to FHU
No functional change
2023-12-25 06:38:51 -08:00
Ryan Houdek 03f63f99a8 FEXCore: Moves StringUtils to FEXCore headers
Once gdbserver gets moved to the frontend this will need to be in the
includes.
2023-11-02 20:09:12 -07:00
Ryan Houdek fea72ce19c Merge pull request #3120 from Sonicadvance1/more_optimal_x87
FEXCore: Support preserve_all ABI for interpreter fallbacks
2023-09-21 15:35:37 -07:00
Ryan Houdek d588d41ab9 InterpreterFallbacks: Converts X87 and String ops to preserve_all
This improves performance!
2023-09-20 18:51:18 -07:00
Ryan Houdek e85b90c614 FEXCore/Common: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 745729cdc2 SoftFloat-3e: Adds preserve_all attribute to all functions used
This will let FEX's JIT be more optimal
2023-09-18 17:42:48 -07:00
Ryan Houdek 0c5c146fcf FEXCore/JitSymbols: Buffer writes to reduce overhead
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.

The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.

To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
  clang benchmark and seeing how frequently it was still writing.
   - 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
  symbols.
   - In most cases 100ms is enough that you won't notice the blip in
     perf.

One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00