to be consistent with the scalar _Andn opcode, which is specifically named _Andn
and not _Bic.
noticed while reviewing AVX patches
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Needed something inbetween the `InlineJITBlockHeader` and `avx_high` in
order to match alignment requirements of 16-byte for avx_high. Chose the
`DeferredSignalRefCount` because we hit it quite frequently and it is
basically the only 64-bit variable that we end up touching
significantly.
In the future the CPUState object is going to need to change its view of
the object depending on if the device supports SVE256 or not, but we
don't need to frontload the work right now. It'll become significantly
easier to support that path once the RCLSE pass gets deleted.
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.
While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.
Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.
This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
This currently doesn't do much but soon this will be very important to
ensure the data prefetcher of Cortex keeps the cachelines following this
variable in L1.
No instcountci changes yet, since nothing currently spills in instcountci. This
mitigates spilling later seen with #3703, and should help for certain
pathological blocks even without those changes (maybe we should try to get some
of those blocks in instcountci?).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
slightly worse for compile time, slightly better output, honestly I'll take the
win because this is easier to reason about.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
I don't get the point, it should be handled by a combination of existing
passes/techniques just fine. no instcountci changes.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
find-and-replace across the tree, excluding IR.h itself.
also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
deduplicate all the things.
functional change:
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This was causing us to generate invalid code in Darwinia, resulting in a
crash. With assertions enabled this would be picked up in the emitter.
Only implement AddShift optimizations for now because I don't want to do
the remaining optimizations in a bug fix PR.
Fixes Darwinia.
Accidentally we were swapping which sources were the base and which was
the one getting shifted. This wasn't super common so it usually didn't
matter.
Fixes one crash in Darwinia.
use a vec. block indices will be dense in the new IR. This is memory intensive
but seems faster in practice.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.
Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If both the destination and the flags are dead for an AddWithFlags, we need to
eliminate it in one pass. If we only replace without elimiating, we would need a
second DCE pass to eliminate. We want DCE to finish in one pass, so fix this.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This has been deadcode since 2020. Drop it so we can focus on what *does* work
and what does matter.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generally, there are three reasons to track progress:
* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.
None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
I recommend viewing the new source file as the diff is quite messy.
---
The old RA commits every "how not to write an RA" sin in the book.
Chaitin spill-one loop? Check.
Potential spilling caused by alignment issues since there's no live range
splitting? Check.
Panic spilling? Check.
Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.
...
It is wholly unsuitable to any application, and specifically unsuitable for FEX.
---
The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.
Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.
The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.
The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.
That's it :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Fails with the new RA when copies are inserted, since RAValidation loses
visibility. Nontrivial to fix, but this particular assert hopefully isn't buying
us a ton. Weaken it for now, we can revisit later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.
Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
It doesn't write to its node. Fixes spurious
%7: Arg[0] expects reg0 to contain %4, but it actually contains %16
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
now that we've eliminated cross block liveness, we can do our validation locally
too for a massive simplification.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
All we actually need to validate is that each source has been previously defined
within the block. That checks everything we care about now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>