stop prefixing the arguments when we generate allocate ops (in particular), this
is more convenient and simpler. in exchange we need to prefix Op to avoid a
collision on fcmpscalarinsert which has an argument named Op, but that's a local
change at least.
came up when experimenting with new IR, but I think this is probably a win by
itself.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When AFP is supported then we can actually support DAZ. This might also
fix the audio corruption in Animal Well but I can't test it until Steam
is running on Oryon. Requires a bit of plumbing for MXCSR which we were
hacking around before but now we actually want to store the value.
Fixes#3856
nothing is optimizing around this, it's just adding pointless complexity. if we
want to actually optimize F80Cmp, the right way would be to lift the
implementation into the OpcodeDispatcher or JIT. it wouldn't be terribly
difficult. This kludge doesn't get us closer there.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When FEX hits the optimal case that the destination isn't one of the
incoming sources (other than the incomingDest source) then we can
optimize out two moves per 128-bit lane.
Cuts 256-bit non-SVE gather loads from 50 instructions down to 46.
We can avoid a LOT of pointless work with some dedicated IR ops for specifically
overriding the round mode.
Small behaviour change here: we no longer reset FTZ. I think this is a bug fix?
But if it's not it's not hard to fix.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This does a gather load three ways, SVE256, SVE128, and ASIMD.
This operation is a bit special since it it can't quite handle all
gather loadstores in the 256-bit case and requires the frontend to
decompose the operation in the case that the striding hits a mode that
SVE doesn't support!
The 128-bit case is a lot simpler since both support all the cases where
stride doesn't match. I find this to be a nice compromise while there
aren't any SVE256 products on the market.
In the 128-bit case there is an SVE path which is utilized if the passed
in stride supports what SVE understands, otherwise it falls back to an
ASIMD implementation which manually emulates everything that is
necessary.
This instruction is very explicitly doing basically exactly what AVX
gather instructions want, because it's complex enough that we don't want
to try and make this a generic solution.
to be consistent with the scalar _Andn opcode, which is specifically named _Andn
and not _Bic.
noticed while reviewing AVX patches
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.
Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This avoids a bunch of sharp edges for RA at a small cost when obscure
segment registers are used.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>