Alyssa Rosenzweig
bd0b5eceb8
Merge pull request #3545 from alyssarosenzweig/opt/pf-scalar
...
Use scalar integer code to calculate PF
2024-04-02 11:28:53 -04:00
Alyssa Rosenzweig
b632f7215c
Merge pull request #3544 from alyssarosenzweig/ra/zero-multiple
...
OpcodeDispatcher: drop ZeroMultipleFlags
2024-04-02 11:27:45 -04:00
Ryan Houdek
29c6281e11
Merge pull request #3539 from alyssarosenzweig/ra/rol-ror2
...
rewrite ROL/ROR
2024-04-02 00:17:08 -07:00
Alyssa Rosenzweig
eb4bb5875e
OpcodeDispatcher: absorb invert into PF calculation
...
with xorn
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-01 14:12:33 -04:00
Alyssa Rosenzweig
3b052e826f
OpcodeDispatcher: calculate PF with integer ops
...
based on clang's __builtin_parity
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-01 14:12:32 -04:00
Alyssa Rosenzweig
f8b68d8b5a
OpcodeDispatcher: drop ZeroMultipleFlags
...
lot of complexity for only a single interesting case. we can massively simplify.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-01 13:48:11 -04:00
Ryan Houdek
e2a095372e
Merge pull request #3534 from Sonicadvance1/move_ir_defines
...
FEXCore: Move nearly all IR definitions to internal
2024-04-01 10:00:20 -07:00
Alyssa Rosenzweig
15b86e4c5a
OpcodeDispatcher: rewrite ROL/ROR
...
single unified implementation for ROL & ROR (instead of 4 cases). no more
deferred flags because it's easy to shoot ourselves in the foot with deferred
flags w.r.t the new RA design, and rotates are rare enough with very efficient
flag calculations such that the extra JIT overhead should be minimal to DCE the
resulting calculations later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-31 14:47:19 -04:00
Ryan Houdek
ed3af580c5
FEXCore: Move nearly all IR definitions to internal
...
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.
Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.
No functional change with this, just moving headers around.
2024-03-29 17:20:18 -07:00
Alyssa Rosenzweig
c513b9685d
OpcodeDispatcher: eliminate crossblock liveness in xsave/xrstor
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-29 09:57:16 -04:00
Alyssa Rosenzweig
693d86dd67
OpcodeDispatcher: add SetAFAndFixup helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-25 12:59:19 -04:00
Ryan Houdek
fd391b1b18
JIT: Optimize pmovmaskb with a named vector constant
...
I was looking at some other JIT overheads and this cropped up as some
overhead. Instead of materializing a constant using mov+movk+movk+movk,
load it from the named vector constant array.
In a micro-benchmark this improved performance by 34%.
In bytemark this improved on subbench by 0.82%
2024-03-17 18:40:46 -07:00
Alyssa Rosenzweig
c99cbe6d0a
JIT: switch DF representation
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig
cc82dba1ca
OpcodeDispatcher: use mvn for AF with constants
...
This reduces pointless constant usage. For now, it's no net change to
instcountci, but it should make it easier to get wins later. Hopefully.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-28 10:18:28 -04:00
Alyssa Rosenzweig
12cc980603
OpcodeDispatcher: shuffle adc flag order
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
f3d55dd721
OpcodeDispatcher: shuffle SBC flag order
...
avoids clobbering nzcv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
91cef6b76f
OpcodeDispatcher: use native ADC even for 8/16-bit
...
we mask off the upper bits, and they agree in the lower bits.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig
a750870abf
OpcodeDispatcher: use fused sbcs calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-26 15:25:03 -04:00
Alyssa Rosenzweig
6994fc3a01
IR,OpcodeDispatcher,JIT: fuse adcs flags
...
The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!
Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-25 10:49:32 -04:00
Alyssa Rosenzweig
80e632db8a
OpcodeDispatcher: garbage collect
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
883cca2e8f
OpcodeDispatcher: use AddWithFlags
...
give it the same treatment we just gave sub.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
cc1c1dd047
OpcodeDispatcher: return result from SUB flag calculate
...
for fusion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
dd9d3264dd
OpcodeDispatcher: smarten SUB flag generation
...
we don't need the result, we can use subs and come out ahead in practice. also a
step towards better fusion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig
8762bc1fa3
OpcodeDispatcher: simplify CalculateAF signature
...
- Res is unused
- SrcSize doesn't matter since we ignore the high bits, might as well always use
32-bit, it doesn't matter
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-21 12:48:15 -04:00
Alyssa Rosenzweig
0503c89ff6
OpcodeDispatcher: use NZCV update helpers
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-19 14:12:54 -04:00
Alyssa Rosenzweig
d7ff1b78fb
IR: handle 8/16-bit AddNZCV/SubNZCV
...
we can do it more effectively than the current s/w lowering.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-12 12:36:09 -04:00
Alyssa Rosenzweig
8d3f0b6f02
OpcodeDispatcher: reassociate and sink W in sha1
...
We only need each part of W extracted in the corresponding round, so sink the
extract into the round to reduce pressure.
Further, W and E are added and then never used again. So, by reassociating we
can do the add upfront, killing W and E at the start and further reducing
pressure.
Eliminates spilling in sha1rnds4.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
60f7b9bcc4
OpcodeDispatcher: optimze sha1's 2/3 expr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
a487557173
OpcodeDispatcher: extract BitwiseAtLeastTwo
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
394b4888bb
OpcodeDispatcher: reassociate and remat C0, G0
...
costs 2 moves and eliminates the rest of our spilling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
142cbdd852
OpcodeDispatcher: expand, reassociate, and interleave sha256 calc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
2f9102f78d
OpcodeDispatcher: expand & interleave sha256 calc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
c9824d04cb
OpcodeDispatcher: sink sha256 extracts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
9c2a569539
OpcodeDispatcher: reexpress Major in sha256
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
515aa4ce3e
OpcodeDispatcher: fuse eor+ror in sha256
...
This reduces instructions a ton.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
f616beb992
OpcodeDispatcher: CSE sha
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
0dcf1e12b8
OpcodeDispatcher: copyprop sha logic
...
prepare for clever
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig
2cbf544ef5
OpcodeDispatcher: expand sha logic
...
no functional change, just preparing for cleverness.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-02-02 13:03:07 -04:00
Ryan Houdek
da0e1b515a
Revert "OpcodeDispatcher: Initial support for runtime long-mode switch"
...
This reverts commit 9e5d7aa5fe .
2024-02-01 18:14:24 -08:00
Mai
ae7dc250db
Merge pull request #3386 from alyssarosenzweig/opt/shift
...
Optimize shifts a bit
2024-01-31 14:11:58 -05:00
Alyssa Rosenzweig
c9461d9997
OpcodeDispatcher: optimize BEXTR flag setting
...
use native test.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
d5eb99fac8
OpcodeDispatcher: optimize popcount flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
f3175848b1
OpcodeDispatcher: use lzcnt flag gen for tzcnt
...
as far as flags go, they're identical: set ZF for zero output, set CF for output
= DestSize, undef the rest. merge the impls, so we get the optimized lzcnt impl
for tzcnt.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
3dd597a591
OpcodeDispatcher: optimize lzcnt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
e8e05252f0
OpcodeDispatcher: optimize BLSI
...
and explain why the suss thing we did before was actually right all along.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
93cef53ec0
OpcodeDispatcher: optimize blsr flags
...
reorder to avoid nzcv clobber
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
3a19133267
OpcodeDispatcher: fix inverted BLSR carry
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
9b309b2102
OpcodeDispatcher: optimize blsmsk flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
fe88b904c9
OpcodeDispatcher: fix missing SF set with blsmsk
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig
2e63c6d547
OpcodeDispatcher: fix inverted CF with blsmsk
...
CF set if SRC = 0
per https://www.felixcloutier.com/x86/blsmsk
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-01-30 22:28:06 -04:00