Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
Need #3348 merged first.
As I was casually thinking, this code made me realize that it was quite
branch heavy and could likely be optimized to logic.
The previous code generated some fairly nasty branch heavy code. This
can be optimized to be branchless and take roughly five instructions
per flag. Using a bitfield for each feature would turn each calculation
in to 3-4 instructions but that seems overkill.
Very minor thing.
No need to wait for initialization on for this anymore.
Ever since Init was refactored to do basically no work, this hasn't been
necessary.
CPUID does need to still be initialized after HostFeatures though, so
need to ensure correct member ordering there.
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.
Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
Hades and the vcruntime hits this very hard in memmove.
`86.56% [JIT] tid 458574 [.] JIT_0x18000c375_0x7fffc94790c8`
```asm
0x00007fffc94790f8: ldaprb w3, [x2]
0x00007fffc94790fc: stlrb w3, [x1]
0x00007fffc9479100: add x1, x1, #0x1
0x00007fffc9479104: add x2, x2, #0x1
0x00007fffc9479108: sub x0, x0, #0x1
0x00007fffc947910c: cbnz x0, 0x7fffc94790f8
```
This performance is terrible because Cortex's LRCPC performance is bottom-tier.
Work around the performance issue by forcing things to do larger moves with vector moves instead.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>