as a simple post-RA peephole. much much easier to do post-RA than pre-RA.
This isn't a post-RA /pass/ in the traditional sense... it's done while
assigning registers to coalesce the passes over the IR, since we pay per-pass
and we can merge the walks over the IR.
Closes: #4480
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Remove blows up because of use tracking, but we can do a much simpler version
for post-RA and elide lots of checks from trying to make Remove more general.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now that we can just set registers directly, we can simplify RA a lot. All the
Map/Unmap nonsense - it all goes away. We just assign registers as we go and
everything clicks into place naturally.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:
* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)
We want to get rid of RAData, so we'll move these to the header.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
this will let us encode registers directly inside OrderedNodeWrapper, rather
than pointers to OrderedNode *. that will let us speed up RA & onwards.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
There's no reasonable way to keep this around without adding significant
complexity to RA. This series prefers to drop complexity from RA, lessening the
need for validation in the first place.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
With the change from #4538 I had accidentally broken x87 reduced
precision.
This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.
Fixes Steam when x87 reduced precision is enabled.
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.
Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed
As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
- 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
- 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
- 7.3% -> 5.2% code buffer space used for RIP reconstruction
Fixes#4535
A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher
This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.
This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
TMP4 was used before we passed in a tmp register. Now use that temp
register.
Also return the amount of stack used on the push function. This will be
used in a bit.