Gets rid of potential extraneous copies. We also add handling for cases
where two passes with the same name are unintentionally added.
Previously we'd blindly overwrite the mapping.
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
now obsolete!
Results for the whole series are excellent:
Difference at 95.0% confidence
-3.97603% +/- 0.254656%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
There's no reasonable way to keep this around without adding significant
complexity to RA. This series prefers to drop complexity from RA, lessening the
need for validation in the first place.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
it's only really load bearing for pf/af, which is handled as a global flag opt
now. this mitigates some of the compile time hit from globalizing flag opts.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.
While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.
Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.
This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.
Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generally, there are three reasons to track progress:
* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.
None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
I recommend viewing the new source file as the diff is quite messy.
---
The old RA commits every "how not to write an RA" sin in the book.
Chaitin spill-one loop? Check.
Potential spilling caused by alignment issues since there's no live range
splitting? Check.
Panic spilling? Check.
Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.
...
It is wholly unsuitable to any application, and specifically unsuitable for FEX.
---
The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.
Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.
The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.
The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.
That's it :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
All we actually need to validate is that each source has been previously defined
within the block. That checks everything we care about now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
RA should not depend on whether we support AVX, that's a huge layering
violation! and fortunately, it does not.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Folds reg+const memory address into addressing mode,
if the constant is within 16Kb.
Update instcountci files.
Add test 32Bit_ASM/FEX_bugs/SubAddrBug.asm
RCLSE ignores NZCV and doesn't optimize stores which doesn't help us with PF/AF
either. So, we add a new pass for dead flag elimination (cannibalizing the old
and broken dead flag elimination pass). This is a simple local optimizer that
walks each block backwards, converging in linear time & constant space in a
single iteration.
Right now, it doesn't do a ton (other than a nice reduction in silliness in
the hot Sonic block), but it provides the framework to fuse comparisons.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.
This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.
Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.
Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.
My concrete IR proposals are that:
* SSA values must be killed in the same block that they are defined.
* Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
* LoadGPR/StoreGPR are eliminated in favour of SSA within a block.
This has a lot of nice properties for our setting:
* Except for some internal REP instruction emulation (etc), we already have
registers for everything that escapes block boundaries, so this form is very
easy to go into -- straightforward local value numbering, not a full into
SSA pass.
* Spilling is entirely local (if it happens at all), since everything is in
registers at block boundaries. This is excellent, because Belady's algorithm
lets us spill nearly optimally in linear-time for individual blocks. (And
the global version of Belady's algorithm is massively more complicated...)
A nice fit for a JIT.
Relatedly, it turns out allowing spilling is probably a decent decision,
since the same spiller code can be used to rematerialize constants in a
straightforward way. This is an issue with the current RA.
* Register assignment is entirely local. For the same reason, we can assign
registers "optimally" in linear time & memory (e.g. with linear scan). And
the impl is massively simpler than a full blown SSA-based tree scan RA. For
example, we don't have to worry about parallel copies or coalescing phis or
anything. Massively nicer algorithm to deal with.
* SSA value names can be block local which makes the validation implicit :~)
It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.
Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>