mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-09 08:00:22 +02:00
For the upper-half of the registers it is more efficient to zero the context with `dc zva` on Ampere1A hardware, while Cortex implements this as equivalent uops in their store pipeline and aren't affected one way or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva` performance with 64B/c, but theoretically slightly fewer instructions. C1-Nano on the other hand, clearly loses to `dc zva`, where mops can only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or not if MOPS is a clear win once hardware is actually shipping.