Commit Graph
2489 Commits
Author SHA1 Message Date
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00
Ryan Houdek c6d8e60ef8 OpcodeDispatcher: Make sure to initialize ArithRef
CID 482002
2025-06-19 16:46:15 -07:00
Ryan Houdek 79a8ed53b6 CPUBackend: Make sure to zero initialize variable
CID 482022
2025-06-19 16:41:24 -07:00
Ryan Houdek 7f71b6f1b2 Passes: Use move instead of copy semantics
To initialize in-place.

CID 482035
2025-06-19 16:22:39 -07:00
Ryan Houdek 16d5ca447f FEXCore: DebugData is never null now
This is a required data structure to exist.

CID 482036
2025-06-19 16:21:19 -07:00
Ryan Houdek 1c38b8b046 Merge pull request #4614 from alyssarosenzweig/opt/easy-mov-elim
Merge moves that are immediately consumed
2025-06-19 12:26:07 -07:00
Alyssa Rosenzweig 5243f50ed1 RegisterAllocationPass: merge 32-bit mov + 64-bit and
mov wA, wB
  and xA, xA, ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig cde805147f RegisterAllocationPass: merge 32-bit moves
mov wA, wB
  op wA, wA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 7e39eb3df2 RegisterAllocationPass: merge full size moves
mov xA, xB
  op xA, xA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig af366d4480 RegisterAllocationPass: skip inlineconstant in RA
similar reasoning as guestopcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 33ef98aae7 RegisterAllocationPass: refactor push/pop merge
to make way for move merging.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig c16db2db4a RegisterAllocationPass: simplify an expression
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Tony Wasserka 61d77e3f9b LibraryForwarding: Infer folders for 32-bit wrappers automatically
There's no need to bother the user to select these paths manually.
Instead, just use the same folder names with _32 appended.

Fixes #4588.
2025-06-17 11:57:35 +02:00
Alyssa Rosenzweig e57130e364 JIT: drop UDiv extensions
we already extend in the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig eedcb35270 IR: merge ldiv/lrem handlers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
Ryan Houdek 6549b66cf6 Merge pull request #4606 from alyssarosenzweig/opt/cdq
Optimize CDQ
2025-06-03 12:24:15 -07:00
Alyssa Rosenzweig 6864d48dcf OpcodeDispatcher: optimize cdq
prereq to optimizing sign-ext+ldiv in a reasonable way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:06 -04:00
Tony Wasserka 45a37edd4a Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values 2025-06-03 17:45:29 +02:00
Ryan Houdek fedad275e7 Merge pull request #4597 from alyssarosenzweig/opt/drop-pile-of-constprop
Constant fold on the fly
2025-06-02 12:34:53 -07:00
Alyssa Rosenzweig d966ae145e ConstProp: merge inline + pooling
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44474704    0.46750433    0.45455258    0.45446569  0.0044727894
+  50    0.43149892    0.45173984    0.44252267    0.44295575  0.0045621814
Difference at 95.0% confidence
	-0.0115099 +/- 0.00179263
	-2.53263% +/- 0.394447%
	(Student's t, pooled s = 0.00451771)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig 4cb37e6a1b ConstProp: drop constant folding and algebraic opts
No longer needed.

The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.

    N           Min           Max        Median           Avg        Stddev
x 100     0.4472467    0.46646308    0.45708424    0.45713057  0.0040838243
+ 100    0.44707586    0.46581227    0.45479448    0.45509309  0.0037548573
Difference at 95.0% confidence
	-0.00203748 +/- 0.00108734
	-0.445711% +/- 0.237862%
	(Student's t, pooled s = 0.00392279)

...in addition to a net deletion of 144 lines of code.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:24:58 -04:00
Alyssa Rosenzweig d9da81e99b ConstProp: drop dead cross-instr opts
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 7d734740be Addressing: avoid a Bfe
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 06299cca4b OpcodeDispatcher: optimize out Bfi for storereg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig af5aaab38b OpcodeDispatcher: optimize a few ALU-with-constant ops
instead of relying on ConstProp for this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 77bb01d384 OpcodeDispatcher: don't generate pointless Xor for AF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig a9eb1bd4e6 OpcodeDispatcher: don't generate pointless Bfe for moves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 83ef2da95f OpcodeDispatcher: avoid zero shift in SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ba13ccadb8 OpcodeDispatcher: use ArithRef for 8/16-bit imul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 4a2dee873d OpcodeDispatcher: use ArithRef for small rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 589906e6e6 OpcodeDispatcher: generalize AF on constants optimization
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig f0fcf6d9e6 OpcodeDispatcher: optimize SVE vmovmskpd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 1591ced5a7 OpcodeDispatcher: use ArithRef for ADC/SBB flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 0d235d63d0 OpcodeDispatcher: use ArithRef for bit tests
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 820d5c1447 OpcodeDispatcher: use ArithRef for rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 8093e8f5f1 OpcodeDispatcher: optimize LoadDir
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 802857e411 OpcodeDispatcher: add ArithRef helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig bf3a1839a2 OpcodeDispatcher: transform DF ourselves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ad7844d7da OpcodeDispatcher: do not generate useless VMov
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Ryan Houdek c865eb98ea Merge pull request #4598 from alyssarosenzweig/idc
Stop using hashmap in DCE
2025-06-02 10:26:08 -07:00
Alyssa Rosenzweig 105d0a36ad RedundantFlagCalculationElimination: dont use hashmap
combined results on node from this and the previous commit:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44433582    0.46988457    0.45344824    0.45320041  0.0043288623
+  50     0.4385365    0.46615359    0.45026179    0.44996216  0.0045742233
Difference at 95.0% confidence
	-0.00323825 +/- 0.00176704
	-0.71453% +/- 0.389903%
	(Student's t, pooled s = 0.00445323)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00