Commit Graph
43 Commits
Author SHA1 Message Date
Alyssa Rosenzweig f336261d6b IR: introduce & use more add helpers
so we can stash all the opts.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Ryan Houdek 14c1ee10b6 x87StackOptimizationPass: Removes pair usage
NFC
2025-07-28 16:14:49 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek cebcf50c65 x87OptimizationPass: Fixes {Inc,Dec}StackPop
On the slow path these were pushing and popping in the wrong direction.
Switch them around to ensure the unittests work.
2025-04-02 14:25:02 -07:00
Ryan Houdek feab0bce4b x87StackOptimizationPass: Initialize a couple of arrays
Just to silence some warnings that think these aren't zero initialized
before using.
2025-03-29 14:02:46 -07:00
Paulo Matos 4565f2b689 Initialize class fields to null/zero
Silence a couple of static analyzer warnings.
2025-03-27 15:35:40 +01:00
Paulo Matos b18f148575 Abstract AddressMode into its own header
Also refactor usage of utility functions into x87 Stack Optimization Pass.
2025-03-20 11:41:26 +01:00
Paulo Matos e18a661b50 Refactor OP_STORESTACKMEM case in x87 Stack Opt Pass 2025-03-18 18:34:06 +01:00
Ryan Houdek 031afbfb18 Various: Adds missing SPDX file headers
NFC
2025-03-07 11:34:31 -08:00
Paulo Matos 0bccb1ece5 Ensure predicate cache is reset when control flow leaves block
Whenever the control float leaves the block, it might clobber the
predicate register so we reset the cache whenever that happens.

Fixes #4264
2025-01-27 20:12:43 +01:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Paulo Matos 2d53867668 x87 fst/fld optimization for different addrmodes
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.

Discussed in #4252.
2025-01-14 16:20:33 +01:00
Paulo Matos cbda688e29 Revert "Cache predicate register generation from pattern"
This reverts commit 72a4063651.

Caused #4264
2025-01-10 12:52:11 +01:00
Ryan Houdek a47ed105e7 x87StackOptimizationPass: Minor opt to f80 fchs and fabs
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
2025-01-03 13:47:10 -08:00
Paulo Matos 72a4063651 Cache predicate register generation from pattern 2024-12-06 10:15:38 +01:00
Paulo Matos 1d3ce30e50 Generate SVE for 80bit load/stores when possible
Fixes #4166.
2024-12-06 10:15:29 +01:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
Ryan Houdek 034b62292b Passes/x86StackOptimization: Fixes implicit conversion of OpSize 2024-10-28 19:26:02 -07:00
Ryan Houdek 2dd0a82059 IR: Change F80CVTTo to use IR::OpSize 2024-10-28 02:23:47 -07:00
Ryan Houdek eccfb53bd5 IR: Change StoreStackMemory to use IR::OpSize 2024-10-28 02:01:47 -07:00
Ryan Houdek 44f9df062e IR: Change Vector_FToI to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 30317ac979 IR: Change Float_FToF to use IR::OpSize 2024-10-28 01:22:45 -07:00
Ryan Houdek 65439956bf IR: Change VCastFromGPR to use IR::OpSize 2024-10-28 01:18:21 -07:00
Ryan Houdek 4e21177988 IR: Change VBSL to use IR::OpSize 2024-10-28 01:13:53 -07:00
Ryan Houdek 365d8b9508 IR: Change VInsGPR to use IR::OpSize 2024-10-28 01:07:16 -07:00
Ryan Houdek 29115b3185 IR: Change VFAdd to use IR::OpSize 2024-10-28 00:38:55 -07:00
Ryan Houdek 0dfd5dd96f IR: Change VXor to use IR::OpSize 2024-10-28 00:15:33 -07:00
Ryan Houdek fc04b9113e IR: Change VAnd to use IR::OpSize 2024-10-28 00:11:44 -07:00
Ryan Houdek 0a34a43976 IR: Change VFSqrt to use IR::OpSize 2024-10-27 22:42:37 -07:00
Ryan Houdek 9f18de0196 IR: Change VFNeg to use IR::OpSize 2024-10-27 22:33:25 -07:00
Ryan Houdek 0bffdc4e27 IR: Change VFAbs to use IR::OpSize 2024-10-27 22:32:37 -07:00
Ryan Houdek 886db4ffca IR: Change LoadNamedVectorConstant to use IR::OpSize 2024-10-27 22:05:42 -07:00
Ryan Houdek 1a115a8ce6 IR: Change FCmp to use IR::OpSize 2024-10-27 18:06:34 -07:00
Ryan Houdek 764aacaa8f IR: Change VExtractToGPR to use IR::OpSize 2024-10-27 17:58:53 -07:00
Ryan Houdek 014917301a IR: Change StoreMem to use IR::OpSize 2024-10-27 17:14:18 -07:00
Ryan Houdek 5fd127b53a IR: Change StoreContextIndexed to use IR::OpSize 2024-10-27 15:51:30 -07:00
Ryan Houdek ece89ddeab IR: Change LoadContextIndexed to use IR::OpSize 2024-10-27 15:50:03 -07:00
Ryan Houdek 2f9b0de742 IR: Change StoreContext to use IR::OpSize 2024-10-27 15:46:32 -07:00
Ryan Houdek 40fd4bbb66 IR: Change LoadContext to use IR::OpSize 2024-10-27 15:42:07 -07:00
Paulo Matos 10ec6b63b6 Fix FXTRACT for 0.0 and -0.0
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.

Also removed the IR node F80XTRACTStack which is not needed anymore.
2024-10-17 09:05:10 +02:00
Paulo Matos eb76dbdf4d x87 small cleanups; NFC 2024-10-02 18:27:51 +02:00
Billy Laws f26bb6bf53 x87StackOptimizationPass: Default initialise StackMemberInfo members
Not doing so is UB.
2024-07-31 17:30:29 +00:00
Paulo Matos a1378f94ce X87 Code Refactoring and Optimization Pass 2024-07-22 08:44:45 +02:00