And also use it at the same time, since the function signatures changed.
Instead of relying on the host libc math libraries for `long double` ALU
operations, rewrite the entire thing to use softfloat-3e fixed width
float128_t types.
This is a very invasive change in cephes but is a necessary requirement
for getting the precision we require in environments that map `long
double` to be the same as `double`, like Win32 and MacOS.
This fixes the precision issue in transcendental operations when running
under WINE.
This is solving a different problem than what #4411 is specifically
trying to solve.
For our transcendental operations, we can't currently guarantee that
these functions will actually operate at the 128-bit softfloat
precision. While this is true with glibc, this is /not/ true for musl
and likely more libraries.
Instead of relying on our libc implementation to implement these,
instead include the cephes math library directly which is what most
people use for this. Including musl even, but not for all operations.
With this we are no longer beholden to the standard libraries for
providing a correct implementation.
Instead of keeping the vlaue as a string array in the MetaLayer, convert
the value to its final type once.
Improves performance in some hotpaths that were doing config based
string conversion in a relatively high frequency.
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.
Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.
A more tangible win from this change alone is less codegen on the JIT
side.
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.
Also removed the IR node F80XTRACTStack which is not needed anymore.
fixes incorrect signs on FPREM. in turn should fix end-to-end failures logging
into Steam.
Thank you to Sergio Lopez for tracking down the JavaScript fail, and Ryan for
finding the bug.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.
The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.
To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
clang benchmark and seeing how frequently it was still writing.
- 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
symbols.
- In most cases 100ms is enough that you won't notice the blip in
perf.
One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>