Turn this into stp to save a move in a /ton/ of AVX-128 code. It's pretty
annoying to optimize this at the dispatcher level because of the SRA cache, so
doing it in the ConstProp is a nice compromise solution.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
if we want to replace a node with one of its sources, we need to zero extend if
the source is 64-bit and the destination is 32-bit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
every pattern costs us JIT time, but not every pattern is doing anything.
especially with the flag rework of 2023-2024, a lot of patterns just don't make
sense anymore. constant folding in particular isn't too useful now.
only instcountci change is AAD which i'm contractually forbidden from caring
about.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
hash tables have high overhead! but individual blocks are small, so the
quadratic search is _much_ faster in practice. knocks 8% off node.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
more efficient to pool the output of constant folding and such. but we need to
avoid that becoming accidentally quadratic, so rework that too
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Because we already pooled the _Constants, there's no benefit to also pooling inline constants. the robin map to do so just adds extra overhead for no benefit - drop it.
without multiblock, shaves around 2% off node. with multiblock, a bit less than
1% but still a win.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
- IsImmLogical already existed in our CodeEmitter. We just forgot to
allow nullptr arguments and to use it.
- Adds an equivalent IsImmAddSub helper and uses it
This gets us closer to removing vixl's global initializers from FEXCore.
slightly worse for compile time, slightly better output, honestly I'll take the
win because this is easier to reason about.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
I don't get the point, it should be handled by a combination of existing
passes/techniques just fine. no instcountci changes.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
find-and-replace across the tree, excluding IR.h itself.
also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
deduplicate all the things.
functional change:
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This was causing us to generate invalid code in Darwinia, resulting in a
crash. With assertions enabled this would be picked up in the emitter.
Only implement AddShift optimizations for now because I don't want to do
the remaining optimizations in a bug fix PR.
Fixes Darwinia.
Accidentally we were swapping which sources were the base and which was
the one getting shifted. This wasn't super common so it usually didn't
matter.
Fixes one crash in Darwinia.
This has been deadcode since 2020. Drop it so we can focus on what *does* work
and what does matter.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generally, there are three reasons to track progress:
* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.
None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Fixed offset x86 code doesn't quite solve the issue, so adjust this
heuristic just to get instcounci to stop flaking.
This code is going to heavily change soon anyway so +50 doesn't change
much.