Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.
This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.
Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
I don't think these are worth it, and also currently they don't trigger ever.
n=100:
Difference at 95.0% confidence
-0.00245961 +/- 0.00139573
-0.524468% +/- 0.297615%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
combined results on node from this and the previous commit:
N Min Max Median Avg Stddev
x 50 0.44433582 0.46988457 0.45344824 0.45320041 0.0043288623
+ 50 0.4385365 0.46615359 0.45026179 0.44996216 0.0045742233
Difference at 95.0% confidence
-0.00323825 +/- 0.00176704
-0.71453% +/- 0.389903%
(Student's t, pooled s = 0.00445323)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
another global CFG-based optimization -- if we know that the raw PF is already
1-bit we can skip parity evaluation, saving work with floating point compares.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
find-and-replace across the tree, excluding IR.h itself.
also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.
Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If both the destination and the flags are dead for an AddWithFlags, we need to
eliminate it in one pass. If we only replace without elimiating, we would need a
second DCE pass to eliminate. We want DCE to finish in one pass, so fix this.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generally, there are three reasons to track progress:
* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.
None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Generates flags for a variable shift as a dedicated IR op. This lets us optimize
around it (without generating control flow, relying on deferred flag infra,
etc). And it neatly solves our RA problem for shifts.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.
Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.
No functional change with this, just moving headers around.