We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.
Now we fully support VAES on SVE256.
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.
e.g.
0b00'00'00'01
0b00'01'01'01
0b11'11'00'11
are all examples of cases where we can broadcast and then insert.
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.
Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.
For example:
vpshufd ymm0, ymm1, 0b00000011
drops from 50 instructions to 9
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.
In some cases, this can be quite drastic, like with:
vpblendw ymm0, ymm1, ymm2, 0b00000001
being cut down from 98 instructions to 14.
These are old paths still around from when StoreResult used to
automatically perform truncating moves.
These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
(See #3799)
I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)
This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.