When the shift size is exactly 16bytes, then it turns in to a move.
If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.
Fixes#3791.
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes#4319.
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.
We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
Only saves a handful of instructions, but still an improvement.
```
"sha1msg2 xmm0, xmm1": {
- "ExpectedInstructionCount": 11,
+ "ExpectedInstructionCount": 7,
```
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.
Discussed in #4252.
Most of this table ignores REX.W, but two encodings change behaviour
based on REX.W. These two encodings are PEXTRD/PEXTRQ and PINSRD/PINSRQ.
For every other instruction encoding, they will ignore REX.W, but FEX
was requiring that they didn't have REX.W encoding. I had special cased
this in the past by adding PALIGNR, but that didn't handle any of the
other instructions.
We can't just handle REX.W in the OpcodeDispatcher and remove the two
special cased instructions because these vector operations also interact
with instruction prefix 0x66 which changes the operating size to 16bit
with regular instructions.
So instead just generate all listings of instructions with REX.W being
zero and one and install handlers in all cases.