Fixes#3690
When doing scalar insertions, upper bits come from different arguments
depending on the operation. These are listed in the ARM spec under the
NEP bit documentation.
to be consistent with the scalar _Andn opcode, which is specifically named _Andn
and not _Bic.
noticed while reviewing AVX patches
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.
While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.
Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.
This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
I recommend viewing the new source file as the diff is quite messy.
---
The old RA commits every "how not to write an RA" sin in the book.
Chaitin spill-one loop? Check.
Potential spilling caused by alignment issues since there's no live range
splitting? Check.
Panic spilling? Check.
Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.
...
It is wholly unsuitable to any application, and specifically unsuitable for FEX.
---
The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.
Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.
The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.
The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.
That's it :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.
Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This avoids a bunch of sharp edges for RA at a small cost when obscure
segment registers are used.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
IRListView is now purely a view type. Instead, ownership is managed on-demand
by a separate interface (IRStorageBase). Materialization of IRListViews to
owning types is moved to this interface as well.
This also avoids unneeded copies of the data.
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We were previously genrating nonsense code if the destination != source:
faddp v2.4s, v4.4s, v4.4s
faddp s2, v4.2s
The result of the first faddp is ignored, so the second merely calculates the
sum of the first 2 sources (not all 4 as needed).
The correct fix is to feed the first add into the second, regardless of the
final destination:
faddp v2.4s, v4.4s, v4.4s
faddp s2, v2.2s
Hit in an ASM test with new RA.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The frontend will provide the return logic via ExitFunctionEC, which
will be jumped to whenever there is an indirect branch/return to an addr
such that RtlIsEcCode(addr) returns true.
Generates flags for a variable shift as a dedicated IR op. This lets us optimize
around it (without generating control flow, relying on deferred flag infra,
etc). And it neatly solves our RA problem for shifts.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>