This needs to divide by 8 to get a proper byte size for all type sizes.
The only usage of this is currently a uint64_t, so it worked by
coincidence, since sizeof(uint64_t) == 8.
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.
Fix for __builtin_issignaling() test of SPEC2017 classify test.
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.
Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.
```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second
64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801
After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659
80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719
After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749
Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```
Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.
Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
And also use it at the same time, since the function signatures changed.
Instead of relying on the host libc math libraries for `long double` ALU
operations, rewrite the entire thing to use softfloat-3e fixed width
float128_t types.
This is a very invasive change in cephes but is a necessary requirement
for getting the precision we require in environments that map `long
double` to be the same as `double`, like Win32 and MacOS.
This fixes the precision issue in transcendental operations when running
under WINE.
This is solving a different problem than what #4411 is specifically
trying to solve.
For our transcendental operations, we can't currently guarantee that
these functions will actually operate at the 128-bit softfloat
precision. While this is true with glibc, this is /not/ true for musl
and likely more libraries.
Instead of relying on our libc implementation to implement these,
instead include the cephes math library directly which is what most
people use for this. Including musl even, but not for all operations.
With this we are no longer beholden to the standard libraries for
providing a correct implementation.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.
Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.
A more tangible win from this change alone is less codegen on the JIT
side.
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.
Also removed the IR node F80XTRACTStack which is not needed anymore.
fixes incorrect signs on FPREM. in turn should fix end-to-end failures logging
into Steam.
Thank you to Sergio Lopez for tracking down the JavaScript fail, and Ryan for
finding the bug.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.
The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.
To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
clang benchmark and seeing how frequently it was still writing.
- 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
symbols.
- In most cases 100ms is enough that you won't notice the blip in
perf.
One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>