Commit Graph
2776 Commits
Author SHA1 Message Date
Alyssa Rosenzweig 91828efa7a JIT: fix divisor masking
oversight. should fix Steam.

Fixes: de4becc26 ("OpcodeDispatcher: mask certain divisors")
Closes: #4652
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-11 11:34:45 -07:00
Billy Laws cce605d5e0 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-11 11:34:07 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Alyssa Rosenzweig 61150a18cc RegisterAllocationPass: fix bookkeeping with merging
this fixes Hades.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:53 -04:00
Ryan Houdek 95ca20cfee Merge pull request #4636 from alyssarosenzweig/bug/ra-invariant
RegisterAllocationPass: assert an invariant in post-RA prop
2025-07-02 18:17:08 -07:00
Alyssa Rosenzweig 5a536d47fd RegisterAllocationPass: assert an invariant in post-RA prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:53:47 -04:00
Alyssa Rosenzweig 360d8c629e RegisterAllocationPass: optimize cpuid
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)

however this special case we can (and should) optimize for now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig abb41d39e4 RegisterAllocationPass: optimize xgetbv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig d4eb4ef594 IR: plumb CPUID into RA pass
for cpuid folding.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00
Ryan Houdek c6d8e60ef8 OpcodeDispatcher: Make sure to initialize ArithRef
CID 482002
2025-06-19 16:46:15 -07:00
Ryan Houdek 79a8ed53b6 CPUBackend: Make sure to zero initialize variable
CID 482022
2025-06-19 16:41:24 -07:00
Ryan Houdek 7f71b6f1b2 Passes: Use move instead of copy semantics
To initialize in-place.

CID 482035
2025-06-19 16:22:39 -07:00
Ryan Houdek 16d5ca447f FEXCore: DebugData is never null now
This is a required data structure to exist.

CID 482036
2025-06-19 16:21:19 -07:00
Ryan Houdek 3d0c20a263 Merge pull request #4615 from neobrain/feature_logging_qol
Improve log message formatting
2025-06-19 12:50:44 -07:00
Ryan Houdek 1c38b8b046 Merge pull request #4614 from alyssarosenzweig/opt/easy-mov-elim
Merge moves that are immediately consumed
2025-06-19 12:26:07 -07:00
Alyssa Rosenzweig 5243f50ed1 RegisterAllocationPass: merge 32-bit mov + 64-bit and
mov wA, wB
  and xA, xA, ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig cde805147f RegisterAllocationPass: merge 32-bit moves
mov wA, wB
  op wA, wA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 7e39eb3df2 RegisterAllocationPass: merge full size moves
mov xA, xB
  op xA, xA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig af366d4480 RegisterAllocationPass: skip inlineconstant in RA
similar reasoning as guestopcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 33ef98aae7 RegisterAllocationPass: refactor push/pop merge
to make way for move merging.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig c16db2db4a RegisterAllocationPass: simplify an expression
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Tony Wasserka 61d77e3f9b LibraryForwarding: Infer folders for 32-bit wrappers automatically
There's no need to bother the user to select these paths manually.
Instead, just use the same folder names with _32 appended.

Fixes #4588.
2025-06-17 11:57:35 +02:00
Tony Wasserka 57627d4fcf LogManager: Drop unused STDOUT/STDERR log levels 2025-06-16 13:54:03 +02:00
Tony Wasserka 3e85e60a30 LogManager: Use colors for logging when possible 2025-06-13 15:14:47 +02:00
Tony Wasserka 992d86bbc1 LogManager: Shorten debug level strings to a single letter 2025-06-13 14:31:25 +02:00
Alyssa Rosenzweig e57130e364 JIT: drop UDiv extensions
we already extend in the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig eedcb35270 IR: merge ldiv/lrem handlers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
Ryan Houdek 6549b66cf6 Merge pull request #4606 from alyssarosenzweig/opt/cdq
Optimize CDQ
2025-06-03 12:24:15 -07:00
Alyssa Rosenzweig 6864d48dcf OpcodeDispatcher: optimize cdq
prereq to optimizing sign-ext+ldiv in a reasonable way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:06 -04:00
Tony Wasserka 45a37edd4a Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values 2025-06-03 17:45:29 +02:00
Ryan Houdek fedad275e7 Merge pull request #4597 from alyssarosenzweig/opt/drop-pile-of-constprop
Constant fold on the fly
2025-06-02 12:34:53 -07:00
Alyssa Rosenzweig d966ae145e ConstProp: merge inline + pooling
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44474704    0.46750433    0.45455258    0.45446569  0.0044727894
+  50    0.43149892    0.45173984    0.44252267    0.44295575  0.0045621814
Difference at 95.0% confidence
	-0.0115099 +/- 0.00179263
	-2.53263% +/- 0.394447%
	(Student's t, pooled s = 0.00451771)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig 4cb37e6a1b ConstProp: drop constant folding and algebraic opts
No longer needed.

The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.

    N           Min           Max        Median           Avg        Stddev
x 100     0.4472467    0.46646308    0.45708424    0.45713057  0.0040838243
+ 100    0.44707586    0.46581227    0.45479448    0.45509309  0.0037548573
Difference at 95.0% confidence
	-0.00203748 +/- 0.00108734
	-0.445711% +/- 0.237862%
	(Student's t, pooled s = 0.00392279)

...in addition to a net deletion of 144 lines of code.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:24:58 -04:00
Alyssa Rosenzweig d9da81e99b ConstProp: drop dead cross-instr opts
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 7d734740be Addressing: avoid a Bfe
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 06299cca4b OpcodeDispatcher: optimize out Bfi for storereg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig af5aaab38b OpcodeDispatcher: optimize a few ALU-with-constant ops
instead of relying on ConstProp for this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 77bb01d384 OpcodeDispatcher: don't generate pointless Xor for AF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig a9eb1bd4e6 OpcodeDispatcher: don't generate pointless Bfe for moves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 83ef2da95f OpcodeDispatcher: avoid zero shift in SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00