Commit Graph
256 Commits
Author SHA1 Message Date
Ryan Houdek da0e1b515a Revert "OpcodeDispatcher: Initial support for runtime long-mode switch"
This reverts commit 9e5d7aa5fe.
2024-02-01 18:14:24 -08:00
Mai ae7dc250db Merge pull request #3386 from alyssarosenzweig/opt/shift
Optimize shifts a bit
2024-01-31 14:11:58 -05:00
Alyssa Rosenzweig c9461d9997 OpcodeDispatcher: optimize BEXTR flag setting
use native test.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig d5eb99fac8 OpcodeDispatcher: optimize popcount flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig f3175848b1 OpcodeDispatcher: use lzcnt flag gen for tzcnt
as far as flags go, they're identical: set ZF for zero output, set CF for output
= DestSize, undef the rest. merge the impls, so we get the optimized lzcnt impl
for tzcnt.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3dd597a591 OpcodeDispatcher: optimize lzcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig e8e05252f0 OpcodeDispatcher: optimize BLSI
and explain why the suss thing we did before was actually right all along.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 93cef53ec0 OpcodeDispatcher: optimize blsr flags
reorder to avoid nzcv clobber

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3a19133267 OpcodeDispatcher: fix inverted BLSR carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 9b309b2102 OpcodeDispatcher: optimize blsmsk flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig fe88b904c9 OpcodeDispatcher: fix missing SF set with blsmsk 2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 2e63c6d547 OpcodeDispatcher: fix inverted CF with blsmsk
CF set if SRC = 0

per https://www.felixcloutier.com/x86/blsmsk

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 0bc9e1a409 OpcodeDispatcher: clobber OF with shift immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig ae48228943 OpcodeDispatcher: optimize vtestps/vtestpd
I don't really care about AVX but do the same thing we did for vptest.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:24:11 -04:00
Alyssa Rosenzweig e8e35e48c7 OpcodeDispatcher: optimize ptest with tst
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Alyssa Rosenzweig 8b8f27a88f OpcodeDispatcher: optimize ptest with umaxv
to check if the vector is zero, umaxv its elements and check if the reduced
scalar is zero.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Ryan Houdek 4b3792196f Merge pull request #3303 from Sonicadvance1/initial_runtime_longmode_switch
OpcodeDispatcher: Initial support for runtime long-mode switch
2024-01-04 18:17:54 -08:00
Ryan Houdek d8f20751fe FEXCore: Moves IREmitter from the public API to backend
No functional change
2023-12-25 07:00:29 -08:00
Ryan Houdek 9e5d7aa5fe OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2023-12-21 01:54:19 -08:00
Alyssa Rosenzweig e923e83efb OpcodeDispatcher: fix nzcvdirty
lets us use flagm in cmpxchg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-29 14:35:25 -04:00
Alyssa Rosenzweig 23c2a53683 OpcodeDispatcher: move fcmp flag fixup to dispatcher
simpler *and* much faster

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 82b7689ca4 OpcodeDispatcher: remove fcmp deferral
no longer load bearing, delete the abstraction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 149f3e6f6d OpcodeDispatcher: rm flagsOp unused since select rework
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 2dcae23776 Arm64Emitter: Dedicate registers for PF/AF
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).

To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).

The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Ryan Houdek c1d5fae018 Merge pull request #3273 from alyssarosenzweig/opt/shifts
Optimize shifts/rotates
2023-11-14 14:13:56 -08:00
Alyssa Rosenzweig 85b1aa4c2d OpcodeDispatcher: optimize mul flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-14 08:40:41 -04:00
Ryan Houdek c69082b1a4 OpcodeDispatcher: Optimize three sha instructions
- sha1nexte
   - Takes advantage of sha1h if supported
   - Does the operation in a vector otherwise
- sha1msg2
   - Instead of dumping everything to GPRs, we can do this with vectors
   - Mostly matches ARM's sha1su1 instruction, but it is /just/
     different enough to be annoying.
- sha256msg1
   - Directly matches sha256u0
   - Leaves the previous implementation alone
2023-11-13 18:38:02 -08:00
Alyssa Rosenzweig 4669c4541c OpcodeDispatcher: don't zero for flagm ror
missed earlier in the PR, would be annoying to rebase in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig 769a8c41c4 OpcodeDispatcher: Branch over shift=0 flags
Not supposed to touch flags at all, so don't! instead of making a terrible mess
of csels. a lot less instructions, and probably faster because the branch should
be predicted correctly in practice in hot loops.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig bec9dba2b1 OpcodeDispatcher: use shifted xor + rmif for rotates
eliminates lots of Bfe on flagm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig 2073f6d287 OpcodeDispatcher: don't zero nzcv for flagm shifts
Faster for flagm. would be slower for !flagm because bfi slowness...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig 0f25a960ee OpcodeDispatcher: remove bfe for small shl imm
We allow the garbage in flags calculation, it's ignored.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig 282ed3e309 OpcodeDispatcher: optimize bsf/bsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig cd031a7d38 OpcodeDispatcher: avoid some ubfx for flagm
Do the masking as part of the rmif, for free.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig 224a1f19a3 OpcodeDispatcher: improve bzhi flag gen
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig e25849b2cb OpcodeDispatcher: fix BZHI flag calculation
needs SF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig ef544fecf2 OpcodeDispatcher: Use predicated neg for x87 fild
Saves an instruction on non-CSSC platforms by deleting a redundant cmp.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig 584c4cc05e OpcodeDispatcher: Mask with rmif sometimes
For CF/OF calculation, this saves an instruction on flagm platforms.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 10:05:51 -04:00
Alyssa Rosenzweig 279afd88bb OpcodeDispatcher: Generalize rmif trick
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 10:05:51 -04:00
Alyssa Rosenzweig 1281145982 OpcodeDispatcher: Optimize popf
mostly for easier debugging tbh

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 5336129b58 IR: Optimize sub/sbb
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 72fc2b522d IR: Optimize add/adc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig b3055523b4 IR: Switch to dedicated NZCV load/store
Semantics differ markedly from the non-NZCV flags, splitting this out makes it a
lot easier to do things correctly imho. Gets the dest/src size correct
(important for spilling), as well as makes our existing opt passes skip this
which is needed for correctness at the moment anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 73958b9163 OpcodeDispatcher: Use DeriveOp
Replace every instance of the Op overwrite pattern, and ban that anti-pattern
from the codebase in the future. This will prevent piles of NZCV related
regressions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-07 12:05:00 -04:00
Ryan Houdek 5103f2d92b Merge pull request #3247 from alyssarosenzweig/refactor/nzcv-prereq
Preparatory patches for nzcv
2023-11-01 14:07:14 -07:00
Alyssa Rosenzweig 5522c6db9c OpcodeDispatcher: Use jump wrappers
Mostly automated replacement + renaming for build fixing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Mai 9612b2fe4b Merge pull request #3240 from Sonicadvance1/optimize_palignr_zero
OpcodeDispatcher: Optimize palignr with zero immediate
2023-11-01 05:55:52 +01:00
Ryan Houdek e5df636efd OpcodeDispatcher: Optimize pblendw
Requires #3238 to be merged first since this uses the tbx IR operation.

Worst case is now a three instruction sequence of ldr+ldr+tbx.
Some operations are special-cased, which definitely doesn't cover all
possible cases we could use without tbx, but as a worst case improvement
this is a significant improvement.
2023-10-31 21:18:02 -07:00
Ryan Houdek de10cbad98 OpcodeDispatcher: Optimize blendps
A bunch of blendps swizzles weren't optimal. This optimizes all swizzles
to be optimal.

Two instructions can be more optimal without a tbx but the rest required
tbx to be optimal since they don't match ARM's swizzle mechanics.
2023-10-31 20:06:45 -07:00
Mai 77d92872bc Merge pull request #3212 from Sonicadvance1/dpp_opt
OpcodeDispatcher: Optimize 128-bit DPPS and DPPD
2023-11-01 04:01:05 +01:00