Two optimizations here:
1) The final VInsElement was generating three instructions
- This itself could have been change to vzip, which would have
removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
instruction off the calculation.
- addp odd elements and even elemnts first
- Then transpose those elements
- Then use one final addp to generate the result in the correct
order.
The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.
With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).
This is a significant improvement even on platforms that only support
128-bit SVE.
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.
Back to back bfi is actually less optimal due to dependency tracking.
With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.