mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-06 21:00:17 +02:00
Turns out Bayonetta hammers SINCOS, our splitting the operation is actually harming the performance of games that heavily use FSINCOS. We instead can actually combine the operation which improves performance. Not enough to get the game running full speed consistently on my Radxa, but good numbers in my microbenchmark. ``` Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second 64-bit: Before: FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319 FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239 FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801 After: FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573 FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965 FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659 80-bit: Before: FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629 FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808 FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719 After: FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922 FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084 FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749 Improvement 64-bit: 1.75x Improvement 80-bit: 1.05x ``` Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second. Disabled in the simulator because we can't easily support pairs of vector registers being returned.
FEXCore - Fast x86 Core emulation library
This is the core emulation library that is used for the FEX emulator project. This project aims to provide a fast and functional x86-64 emulation library that can meet and surpass other x86-64 emulation libraries.
Goals
- Be as fast as possible, beating and exceeding current options for x86-64 emulation
- 25% - 50% lower performance than native code would be desired target
- Use an IR to efficiently translate x86-64 to our host architecture
- Support a tiered recompiler to allow for fast runtime performance
- Support offline compilation and offline tooling for inspection and performance analysis
- Support threaded emulation. Including emulating x86-64's strong memory model on weak memory model architectures
- Support a significant portion of the x86-64 instruction space.
- Including MMX, SSE, SSE2, SSE3, SSSE3, and SSE4*
- Support fallback routines for uncommonly used x86-64 instructions
- Including x87 and 3DNow!
- Only support userspace emulation.
- All x86-64 instructions run as if they are under CPL-3(userland) security layer
- Minimal Linux Syscall emulation for testing purposes
- Portable library implementation in order to support easy integration in to applications
Target Host Architecture
The target host architecture for this library is AArch64. Specifically the ARMv8.1 version or newer. The CPU IR is designed with AArch64 in mind but should allow for other architectures as well. x86-64 host support is available for ease of development, but is not a priority.
Not desired
- Kernel space emulation
- CPL0-2 emulation
- Real Mode, Protected Mode, Virtual-8086 Mode, System Management Mode
- IRQs
- SVM
- "Cycle Accurate" emulation