We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!
Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.
We can move the masking in the AVX variant too, just to be consistent.
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.
Makes using these ops a little bit less noisy.
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.
Now we fully support VAES on SVE256.
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.
e.g.
0b00'00'00'01
0b00'01'01'01
0b11'11'00'11
are all examples of cases where we can broadcast and then insert.
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.