The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.
8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.
Both scalar and vector wide.
Fixes#5326
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.
Additionally, FEXOfflineCompiler obviously should never discard compiled code.
Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.
Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.
Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.
I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
No need for individual calls here. While doing this I also realized
xxhash has this at one point, I guess I'll PR that there later
Signed-off-by: crueter <crueter@eden-emu.dev>