3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes#4319.
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.
We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
Only saves a handful of instructions, but still an improvement.
```
"sha1msg2 xmm0, xmm1": {
- "ExpectedInstructionCount": 11,
+ "ExpectedInstructionCount": 7,
```
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.
Discussed in #4252.
Most of this table ignores REX.W, but two encodings change behaviour
based on REX.W. These two encodings are PEXTRD/PEXTRQ and PINSRD/PINSRQ.
For every other instruction encoding, they will ignore REX.W, but FEX
was requiring that they didn't have REX.W encoding. I had special cased
this in the past by adding PALIGNR, but that didn't handle any of the
other instructions.
We can't just handle REX.W in the OpcodeDispatcher and remove the two
special cased instructions because these vector operations also interact
with instruction prefix 0x66 which changes the operating size to 16bit
with regular instructions.
So instead just generate all listings of instructions with REX.W being
zero and one and install handlers in all cases.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
this is a pain to track and, it turns out, buys us virtually nothing on flagm
systems. rip it out.
this fixes a bug with failing to set in all the right places.
on non-flagm systems there's a slight instcountci impact, but that is mostly
mitigated by the earlier patches in the series. so overall a wash there but
worth it for making the codebase easier
to reason about.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Ensures reads don't go past the end of the page boundary.
SVE masked loads can make this more effective but `VLoadVectorMasked`
isn't setup to be efficient for this case yet.
NFC
Finally converts the IR operations themselves to store the OpSize for
the IR operation size and element sizes.
This also finally, FINALLY, converts that remaining `_Constant` helper
to stop using a size field that is specified in bits rather than bytes
like all the other IR op handlers. That thing was so confusing and now
it's gone.