Commit Graph
281 Commits
Author SHA1 Message Date
Alyssa Rosenzweig cc82dba1ca OpcodeDispatcher: use mvn for AF with constants
This reduces pointless constant usage. For now, it's no net change to
instcountci, but it should make it easier to get wins later. Hopefully.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:18:28 -04:00
Alyssa Rosenzweig 12cc980603 OpcodeDispatcher: shuffle adc flag order
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig f3d55dd721 OpcodeDispatcher: shuffle SBC flag order
avoids clobbering nzcv

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig 91cef6b76f OpcodeDispatcher: use native ADC even for 8/16-bit
we mask off the upper bits, and they agree in the lower bits.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig a750870abf OpcodeDispatcher: use fused sbcs calculations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:25:03 -04:00
Alyssa Rosenzweig 6994fc3a01 IR,OpcodeDispatcher,JIT: fuse adcs flags
The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!

Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 10:49:32 -04:00
Alyssa Rosenzweig 80e632db8a OpcodeDispatcher: garbage collect
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 883cca2e8f OpcodeDispatcher: use AddWithFlags
give it the same treatment we just gave sub.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig cc1c1dd047 OpcodeDispatcher: return result from SUB flag calculate
for fusion

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig dd9d3264dd OpcodeDispatcher: smarten SUB flag generation
we don't need the result, we can use subs and come out ahead in practice. also a
step towards better fusion

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 8762bc1fa3 OpcodeDispatcher: simplify CalculateAF signature
- Res is unused
- SrcSize doesn't matter since we ignore the high bits, might as well always use
  32-bit, it doesn't matter

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-21 12:48:15 -04:00
Alyssa Rosenzweig 0503c89ff6 OpcodeDispatcher: use NZCV update helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-19 14:12:54 -04:00
Alyssa Rosenzweig d7ff1b78fb IR: handle 8/16-bit AddNZCV/SubNZCV
we can do it more effectively than the current s/w lowering.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-12 12:36:09 -04:00
Alyssa Rosenzweig 8d3f0b6f02 OpcodeDispatcher: reassociate and sink W in sha1
We only need each part of W extracted in the corresponding round, so sink the
extract into the round to reduce pressure.

Further, W and E are added and then never used again. So, by reassociating we
can do the add upfront, killing W and E at the start and further reducing
pressure.

Eliminates spilling in sha1rnds4.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 60f7b9bcc4 OpcodeDispatcher: optimze sha1's 2/3 expr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig a487557173 OpcodeDispatcher: extract BitwiseAtLeastTwo
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 394b4888bb OpcodeDispatcher: reassociate and remat C0, G0
costs 2 moves and eliminates the rest of our spilling

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 142cbdd852 OpcodeDispatcher: expand, reassociate, and interleave sha256 calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 2f9102f78d OpcodeDispatcher: expand & interleave sha256 calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig c9824d04cb OpcodeDispatcher: sink sha256 extracts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 9c2a569539 OpcodeDispatcher: reexpress Major in sha256
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 515aa4ce3e OpcodeDispatcher: fuse eor+ror in sha256
This reduces instructions a ton.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig f616beb992 OpcodeDispatcher: CSE sha
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 0dcf1e12b8 OpcodeDispatcher: copyprop sha logic
prepare for clever

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 2cbf544ef5 OpcodeDispatcher: expand sha logic
no functional change, just preparing for cleverness.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Ryan Houdek da0e1b515a Revert "OpcodeDispatcher: Initial support for runtime long-mode switch"
This reverts commit 9e5d7aa5fe.
2024-02-01 18:14:24 -08:00
Mai ae7dc250db Merge pull request #3386 from alyssarosenzweig/opt/shift
Optimize shifts a bit
2024-01-31 14:11:58 -05:00
Alyssa Rosenzweig c9461d9997 OpcodeDispatcher: optimize BEXTR flag setting
use native test.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig d5eb99fac8 OpcodeDispatcher: optimize popcount flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig f3175848b1 OpcodeDispatcher: use lzcnt flag gen for tzcnt
as far as flags go, they're identical: set ZF for zero output, set CF for output
= DestSize, undef the rest. merge the impls, so we get the optimized lzcnt impl
for tzcnt.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3dd597a591 OpcodeDispatcher: optimize lzcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig e8e05252f0 OpcodeDispatcher: optimize BLSI
and explain why the suss thing we did before was actually right all along.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 93cef53ec0 OpcodeDispatcher: optimize blsr flags
reorder to avoid nzcv clobber

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3a19133267 OpcodeDispatcher: fix inverted BLSR carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 9b309b2102 OpcodeDispatcher: optimize blsmsk flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig fe88b904c9 OpcodeDispatcher: fix missing SF set with blsmsk 2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 2e63c6d547 OpcodeDispatcher: fix inverted CF with blsmsk
CF set if SRC = 0

per https://www.felixcloutier.com/x86/blsmsk

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 0bc9e1a409 OpcodeDispatcher: clobber OF with shift immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig ae48228943 OpcodeDispatcher: optimize vtestps/vtestpd
I don't really care about AVX but do the same thing we did for vptest.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:24:11 -04:00
Alyssa Rosenzweig e8e35e48c7 OpcodeDispatcher: optimize ptest with tst
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Alyssa Rosenzweig 8b8f27a88f OpcodeDispatcher: optimize ptest with umaxv
to check if the vector is zero, umaxv its elements and check if the reduced
scalar is zero.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Ryan Houdek 4b3792196f Merge pull request #3303 from Sonicadvance1/initial_runtime_longmode_switch
OpcodeDispatcher: Initial support for runtime long-mode switch
2024-01-04 18:17:54 -08:00
Ryan Houdek d8f20751fe FEXCore: Moves IREmitter from the public API to backend
No functional change
2023-12-25 07:00:29 -08:00
Ryan Houdek 9e5d7aa5fe OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2023-12-21 01:54:19 -08:00
Alyssa Rosenzweig e923e83efb OpcodeDispatcher: fix nzcvdirty
lets us use flagm in cmpxchg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-29 14:35:25 -04:00
Alyssa Rosenzweig 23c2a53683 OpcodeDispatcher: move fcmp flag fixup to dispatcher
simpler *and* much faster

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 82b7689ca4 OpcodeDispatcher: remove fcmp deferral
no longer load bearing, delete the abstraction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 149f3e6f6d OpcodeDispatcher: rm flagsOp unused since select rework
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig 2dcae23776 Arm64Emitter: Dedicate registers for PF/AF
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).

To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).

The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Ryan Houdek c1d5fae018 Merge pull request #3273 from alyssarosenzweig/opt/shifts
Optimize shifts/rotates
2023-11-14 14:13:56 -08:00