Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.
We can move the masking in the AVX variant too, just to be consistent.
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.
Makes using these ops a little bit less noisy.
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.
Now we fully support VAES on SVE256.
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.
e.g.
0b00'00'00'01
0b00'01'01'01
0b11'11'00'11
are all examples of cases where we can broadcast and then insert.
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.
Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.
For example:
vpshufd ymm0, ymm1, 0b00000011
drops from 50 instructions to 9