mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-09 23:00:20 +02:00
Noticed this while benchmarking that the FIST* operations were converting to a GPR, and then storing to memory using an atomic TSO operation. This should be instead listening to the vector TSO configuration option. This gives a 3.8x - 6.05x improvement in my microbench. Additionally when possible, make sure to use vector conversion instructions when possible. It's lower cost to avoid the FPR->GPR transfer, but we can only use it for 64-bit FIST operations. Microbench couldn't show a difference for that on my platform, but that's because it's float pipeline bounded regardless. Should help X-class Cortex and newer Cortex-A.