Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
When the shift size is exactly 16bytes, then it turns in to a move.
If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.
Fixes#3791.
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes#4319.
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.
We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
Only saves a handful of instructions, but still an improvement.
```
"sha1msg2 xmm0, xmm1": {
- "ExpectedInstructionCount": 11,
+ "ExpectedInstructionCount": 7,
```