Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>