Commit Graph
284 Commits
Author SHA1 Message Date
Ryan Houdek 67e1ac0442 Merge pull request #3725 from alyssarosenzweig/ir/vbic
IR: rename _VBic -> _VAndn
2024-06-18 16:34:26 -07:00
Alyssa Rosenzweig 01da5972fc IR: rename _VBic -> _VAndn
to be consistent with the scalar _Andn opcode, which is specifically named _Andn
and not _Bic.

noticed while reviewing AVX patches

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 14:00:01 -04:00
Ryan Houdek bf812aae8f CoreState: Adds avx_high structure for tracking decoupled AVX halves.
Needed something inbetween the `InlineJITBlockHeader` and `avx_high` in
order to match alignment requirements of 16-byte for avx_high. Chose the
`DeferredSignalRefCount` because we hit it quite frequently and it is
basically the only 64-bit variable that we end up touching
significantly.

In the future the CPUState object is going to need to change its view of
the object depending on if the device supports SVE256 or not, but we
don't need to frontload the work right now. It'll become significantly
easier to support that path once the RCLSE pass gets deleted.
2024-06-18 12:00:45 -04:00
Ryan Houdek 1ce27a5e6b FEXCore: Disentangle the SVE256 feature from AVX
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.

While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.

Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.

This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
2024-06-17 17:20:32 -07:00
Ryan Houdek 933d622860 Merge pull request #3710 from Sonicadvance1/avx128_1
CoreState: Move `InlineJITBlockHeader` to the start of the struct
2024-06-17 17:17:56 -07:00
Ryan Houdek a9bacc1b6b CoreState: Move InlineJITBlockHeader to the start of the struct
This currently doesn't do much but soon this will be very important to
ensure the data prefetcher of Cortex keeps the cachelines following this
variable in L1.
2024-06-17 02:59:56 -07:00
Alyssa Rosenzweig 9443b18076 RegisterAllocationPass: optimize spill loop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-16 08:15:15 -04:00
Alyssa Rosenzweig 6a314bc9cd RegisterAllocationPass: prioritize remat over spilling
No instcountci changes yet, since nothing currently spills in instcountci. This
mitigates spilling later seen with #3703, and should help for certain
pathological blocks even without those changes (maybe we should try to get some
of those blocks in instcountci?).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:21:57 -04:00
Alyssa Rosenzweig a8bf3859ea ConstProp: rm pointless constant folding
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig aa7dcffcea ConstProp: drop const pool heuristic
slightly worse for compile time, slightly better output, honestly I'll take the
win because this is easier to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig be1a5cea8e ConstProp: drop addressgen const pool stuff
I don't get the point, it should be handled by a combination of existing
passes/techniques just fine. no instcountci changes.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 402ea84aa0 RedundantFlagCalculationElimination: cleanup DCE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 19a7b06b91 ConstProp: swallow up LongDivideElimination
as usual.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 96bd643e5b ConstProp: always inline constants
x86/interpreter leftover, I think.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 6b9293979c ConstProp: swallow up InlineCallOptimization
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 7d5cee4384 InlineCallOptimization: rm x86 leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 32f5a28433 IR: use Ref instead of OrderedNode
find-and-replace across the tree, excluding IR.h itself.

also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig ce30179ed1 IR: add Ref typedef
To put new IR lipstick on the old IR pig.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig 112c49a348 ConstProp: fix inlining shifted imm to mem instructions
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 80878ae611 ConstProp: rework mem immediate inlining
deduplicate all the things.

functional change:
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 85a69be5b6 ConstProp: drop address fusion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:38:03 -04:00
Ryan Houdek ee96d60983 Merge pull request #3673 from alyssarosenzweig/ra/tied
Track tied sources in the IR
2024-05-30 10:55:15 -07:00
Alyssa Rosenzweig 55391ccbc0 RegisterAllocationPass: try to coalesce tied sources
we'll do better in the future but this is already a win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 7790d7a0b7 IR: track tied sources
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Ryan Houdek 80687c8d2d ConstProp: Limits which addressing modes can be used for vector loadstores
This was causing us to generate invalid code in Darwinia, resulting in a
crash. With assertions enabled this would be picked up in the emitter.

Only implement AddShift optimizations for now because I don't want to do
the remaining optimizations in a bug fix PR.

Fixes Darwinia.
2024-05-29 04:42:11 -07:00
Ryan Houdek 920fe60492 ConstProp: Fix bug with transposed elements from AddShift op
Accidentally we were swapping which sources were the base and which was
the one getting shifted. This wasn't super common so it usually didn't
matter.

Fixes one crash in Darwinia.
2024-05-29 04:32:51 -07:00
Ryan Houdek 8c6ce2cb3b Passes/ConstProp: Have MemExtendedAddressing return a struct rather than a tuple
Makes it less confusing about which variable is the base versus the
offset.

NFC
2024-05-29 04:32:14 -07:00
Ryan Houdek 61f30d004c IR: Document AddShift behaviour
Just to clarify that Src2 is the shifted operation.
2024-05-29 04:29:09 -07:00
Alyssa Rosenzweig 734258e23b Merge pull request #3661 from Sonicadvance1/remove_warnings2
Removes warnings
2024-05-25 11:52:37 -04:00
Ryan Houdek 74916b3757 RAPass: Remove warnings 2024-05-24 18:41:30 -07:00
Alyssa Rosenzweig bc1669b163 DeadStoreElimination: eliminate map
use a vec. block indices will be dense in the new IR. This is memory intensive
but seems faster in practice.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 83e417b2c6 DeadStoreElimination: combine GPR/FPR handling
slight speed up per profile.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cb00d9171f IR: merge general DCE with flag DCE
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.

Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cf77f2ae5d RedundantFlagCalculationElimination: fix convergence issue
If both the destination and the flags are dead for an AddWithFlags, we need to
eliminate it in one pass. If we only replace without elimiating, we would need a
second DCE pass to eliminate. We want DCE to finish in one pass, so fix this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 273d086a7b ConstProp: merge const pooling passes
walk the IR less.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 94d9cf54bc ConstProp: don't push/pop cursor
pointless

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3089e0e6de ConstProp: merge masking opts with const folding
Single pass over the IR now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3c088fb414 ConstProp: remove masking elimination opts
This has been deadcode since 2020. Drop it so we can focus on what *does* work
and what does matter.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 676c9a9be6 IR: do not return progress from passes
Generally, there are three reasons to track progress:

* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.

None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 24cb02f4ff FEXCore: remove IRCompaction
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 725d0e187a RegisterAllocationPass: rewrite RA
I recommend viewing the new source file as the diff is quite messy.

---

The old RA commits every "how not to write an RA" sin in the book.

Chaitin spill-one loop? Check.

Potential spilling caused by alignment issues since there's no live range
splitting? Check.

Panic spilling? Check.

Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.

...

It is wholly unsuitable to any application, and specifically unsuitable for FEX.

---

The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.

Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.

The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.

The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.

That's it :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig d8603cb9bd RAValidation: weaken fill validation
Fails with the new RA when copies are inserted, since RAValidation loses
visibility. Nontrivial to fix, but this particular assert hopefully isn't buying
us a ton. Weaken it for now, we can revisit later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 04c2cb5feb IR: add GPR copy/swap instructions
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.

Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Ryan Houdek c90036aeea Merge pull request #3649 from alyssarosenzweig/ra/validate-less-hard
Simplify/fix our validation passes
2024-05-21 17:12:49 -07:00
Alyssa Rosenzweig 6e0f5eccb3 RAValidation: fix spillregister validation
It doesn't write to its node. Fixes spurious

  %7: Arg[0] expects reg0 to contain %4, but it actually contains %16

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 579fb42458 RAValidation: allow filling the same slot multiple times
This can be a reasonable thing to do!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig f4b487352c RAValidation: defeature control flow analysis
now that we've eliminated cross block liveness, we can do our validation locally
too for a massive simplification.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 4448f84f29 IRValidation: merge in ValueDominanceValidation
All we actually need to validate is that each source has been previously defined
within the block. That checks everything we care about now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Ryan Houdek ca70e387ec Merge pull request #3648 from alyssarosenzweig/ra/pair-extract
Slightly improve pair coalescing + memcpy fix from RA branch
2024-05-21 16:19:35 -07:00
Ryan Houdek 9a483107e3 Merge pull request #3647 from alyssarosenzweig/ir/pass-simpler
ConstProp, RCLSE: simplifications
2024-05-21 16:11:36 -07:00