No instcountci changes yet, since nothing currently spills in instcountci. This
mitigates spilling later seen with #3703, and should help for certain
pathological blocks even without those changes (maybe we should try to get some
of those blocks in instcountci?).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
find-and-replace across the tree, excluding IR.h itself.
also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generally, there are three reasons to track progress:
* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.
None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
I recommend viewing the new source file as the diff is quite messy.
---
The old RA commits every "how not to write an RA" sin in the book.
Chaitin spill-one loop? Check.
Potential spilling caused by alignment issues since there's no live range
splitting? Check.
Panic spilling? Check.
Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.
...
It is wholly unsuitable to any application, and specifically unsuitable for FEX.
---
The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.
Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.
The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.
The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.
That's it :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
RA should not depend on whether we support AVX, that's a huge layering
violation! and fortunately, it does not.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Just like #3508, clang-18 complains about VLA usage.
This vector is relatively small, only around 18 elements but is
semi-dynamic depending on arch and if FEXCore is targeting Linux or
Win32.
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.
Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.
No functional change with this, just moving headers around.
FEXCore includes was including an FHU header which would result in
compilation failure for external projects trying to link to libFEXCore.
Moves it over to fix this, it was the only FHU usage in FEXCore/include
NFC
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).
To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).
The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is the cause of a bunch of redundant moves that shows up in
InstCountCI. Fixing this aliasing and pre-colouring issue causes a ton
of 256-bit operations to become optimal.
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.
My concrete IR proposals are that:
* SSA values must be killed in the same block that they are defined.
* Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
* LoadGPR/StoreGPR are eliminated in favour of SSA within a block.
This has a lot of nice properties for our setting:
* Except for some internal REP instruction emulation (etc), we already have
registers for everything that escapes block boundaries, so this form is very
easy to go into -- straightforward local value numbering, not a full into
SSA pass.
* Spilling is entirely local (if it happens at all), since everything is in
registers at block boundaries. This is excellent, because Belady's algorithm
lets us spill nearly optimally in linear-time for individual blocks. (And
the global version of Belady's algorithm is massively more complicated...)
A nice fit for a JIT.
Relatedly, it turns out allowing spilling is probably a decent decision,
since the same spiller code can be used to rematerialize constants in a
straightforward way. This is an issue with the current RA.
* Register assignment is entirely local. For the same reason, we can assign
registers "optimally" in linear time & memory (e.g. with linear scan). And
the impl is massively simpler than a full blown SSA-based tree scan RA. For
example, we don't have to worry about parallel copies or coalescing phis or
anything. Massively nicer algorithm to deal with.
* SSA value names can be block local which makes the validation implicit :~)
It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.
Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Noticed that we hadn't ever enabled this, which was a concern when our
GPR operations weren't as strict about leaving garbage in the upper bits
when operating as a 32-bit operation.
Now that our ALU operations are more strict about enforcing upper bit
zeroing we can enable this.
This causes Half-Life: Source FPS to get to > 200FPS finally. Causes
significant performance improvements for 32-bit games because we're no
longer redundantly moving registers before and after every operation.
Causing a bunch of 3-4 instruction sequences to convert to 1.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>