This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.
Difference at 95.0% confidence
-0.00138911 +/- 0.00104724
-0.418608% +/- 0.315587%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.
Difference at 95.0% confidence
-0.00474273 +/- 0.00119189
-1.40908% +/- 0.354114%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.
Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.
```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second
64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801
After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659
80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719
After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749
Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```
Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.
Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)
however this special case we can (and should) optimize for now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
this will eliminate an annoying special case in post-RA opts.
No difference proven at 95.0% confidence
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:
N Min Max Median Avg Stddev
x 50 0.44474704 0.46750433 0.45455258 0.45446569 0.0044727894
+ 50 0.43149892 0.45173984 0.44252267 0.44295575 0.0045621814
Difference at 95.0% confidence
-0.0115099 +/- 0.00179263
-2.53263% +/- 0.394447%
(Student's t, pooled s = 0.00451771)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
No longer needed.
The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.
N Min Max Median Avg Stddev
x 100 0.4472467 0.46646308 0.45708424 0.45713057 0.0040838243
+ 100 0.44707586 0.46581227 0.45479448 0.45509309 0.0037548573
Difference at 95.0% confidence
-0.00203748 +/- 0.00108734
-0.445711% +/- 0.237862%
(Student's t, pooled s = 0.00392279)
...in addition to a net deletion of 144 lines of code.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
combined results on node from this and the previous commit:
N Min Max Median Avg Stddev
x 50 0.44433582 0.46988457 0.45344824 0.45320041 0.0043288623
+ 50 0.4385365 0.46615359 0.45026179 0.44996216 0.0045742233
Difference at 95.0% confidence
-0.00323825 +/- 0.00176704
-0.71453% +/- 0.389903%
(Student's t, pooled s = 0.00445323)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
as a simple post-RA peephole. much much easier to do post-RA than pre-RA.
This isn't a post-RA /pass/ in the traditional sense... it's done while
assigning registers to coalesce the passes over the IR, since we pay per-pass
and we can merge the walks over the IR.
Closes: #4480
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Remove blows up because of use tracking, but we can do a much simpler version
for post-RA and elide lots of checks from trying to make Remove more general.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now that we can just set registers directly, we can simplify RA a lot. All the
Map/Unmap nonsense - it all goes away. We just assign registers as we go and
everything clicks into place naturally.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>