Totally forgot to update this when I opened that PR. oops!
FYI I would recommend merge-committing if I PR to other repos since it
preserves my GPG signature
Signed-off-by: crueter <crueter@eden-emu.dev>
We were passing through 32-bit signed values to the 64-bit handler,
which isn't valid and we need to make sure to sign extend.
I don't think this fixes anything because negative values aren't valid,
but a negative could have been interpreted as a large positive without
this.
So WTF can see their usage. Previously didn't add this due to some weird
madvise bug that delete anon names, but that seems to be fixed in a
newer kernel.
Fixes#5230
Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.
Note: I have no idea how fmt will handle that format string with an
array of uint8_t.
Signed-off-by: crueter <crueter@eden-emu.dev>
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.
8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.
Both scalar and vector wide.
Fixes#5326
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.