Alyssa Rosenzweig
d17f33a922
OpcodeDispatcher: don't emit fake 0 for condjump
...
not needed and getting in the way
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig
50a3ca0d6d
IR: introduce dedicated PF/AF instructions
...
this makes reasoning about them a little easier, e.g. for flags. about 1% win
in nodejs.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig
e9ab514962
IR: push parity evaluation down
...
so we can optimize it globally
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig
4eb0948451
IR: push down AXFLAG lowering
...
so we can get the new axflag optimizations on billy's x13s.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig
42f2851575
Merge pull request #4009 from alyssarosenzweig/opt/axflag
...
Optimize AXFLAG-less systems
2024-08-27 08:01:05 -04:00
Alyssa Rosenzweig
8d6b454455
OpcodeDispatcher: optimize AXFLAG emulation
...
this should help on x13s which has flagm but not flagm2.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-25 18:38:03 -04:00
Alyssa Rosenzweig
205ec3e14d
OpcodeDispatcher: refactor axflag
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-25 18:35:27 -04:00
Ryan Houdek
8bf4a124c8
AVX128: Fixes 256-bit float compares
...
Just like the bug in #4006 , we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.
PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek
1d00ad6030
OpcodeDispatcher: Convert VectorVariableBlend to Bind handler
2024-08-22 15:09:44 -07:00
Ryan Houdek
23a076c313
OpcodeDispatcher: Convert packed vector shifts to Bind handler
2024-08-22 15:07:09 -07:00
Ryan Houdek
57eacab654
OpcodeDispatcher: Convert VBROADCASTOp to Bind handler
2024-08-22 14:59:25 -07:00
Ryan Houdek
b4093a8888
OpcodeDispatcher: Convert packed HSub to Bind handler
2024-08-22 14:57:40 -07:00
Ryan Houdek
7c5a9b5d6a
OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler
2024-08-22 14:55:34 -07:00
Ryan Houdek
12b3c82d83
OpcodeDispatcher: Convert PExtr to Bind handler
2024-08-22 14:54:26 -07:00
Ryan Houdek
66520bce0a
OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler
2024-08-22 14:52:41 -07:00
Ryan Houdek
25cc2bdcb8
OpcodeDispatcher: Convert VPERMILImm to Bind handler
2024-08-22 14:51:02 -07:00
Ryan Houdek
f98b18800c
OpcodeDispatcher: Convert MOVMSK to Bind handler
2024-08-22 14:49:57 -07:00
Ryan Houdek
0dd687a7a1
OpcodeDispatcher: Convert SHUFOp to Bind handler
2024-08-22 14:47:30 -07:00
Ryan Houdek
1aff3acbb9
OpcodeDispatcher: Convert PSHUFW to Bind handler
2024-08-22 14:45:04 -07:00
Ryan Houdek
a6ab2ca30d
OpcodeDispatcher: Convert PUNPCKH to Bind handler
2024-08-22 14:42:24 -07:00
Ryan Houdek
ca43e2a61c
OpcodeDispatcher: Convert PUNPCKL to Bind handler
2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig
b03c613fbf
OpcodeDispatcher: optimize JP/JNP
...
fuse and+cbnz into tbz/tbnz.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
5d613e8716
OpcodeDispatcher: optimize test x, x
...
fewer uops now that we invert carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 21:17:08 -04:00
Alyssa Rosenzweig
d7a20fa28f
OpcodeDispatcher: fix tso checks
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 20:11:06 -04:00
Alyssa Rosenzweig
4c4c6e7807
OpcodeDispatcher: pair loads on cortex
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:36:50 -04:00
Alyssa Rosenzweig
a8c9c71ce3
OpcodeDispatcher: use loadcontextpair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
32d6daf558
OpcodeDispatcher: use ldp/stp for AVX load/store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
bebcb73c68
OpcodeDispatcher: pair AVX high writes
...
reduces instr count with AVX-128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
c70e44cf05
OpcodeDispatcher: extract Push helper
...
Mirrors the Pop helper we added. This cleans up a bunch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
900c62fa7b
OpcodeDispatcher: add Pop helpers
...
hide away the allocate dance
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
cab02be637
IR: remove unused pair create/extract
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
5631ff4fd5
OpcodeDispatcher: optimize variable shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
f138d7d9b8
OpcodeDispatcher: optimize JA/JNA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
3bb9d44bf5
OpcodeDispatcher: optimize SetCFDirect
...
generally same # of instructions, but potentially fewer cycles:
old:
"rmif x4, #63 , #nzCv",
"cfinv"
new:
"xor x20, x4, #1 ",
"rmif x20, #63 , #nzCv"
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
06e6a1b19e
OpcodeDispatcher: optimize logic ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
9621eca677
OpcodeDispatcher: fix setrflag masking
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
d17c427e47
OpcodeDispatcher: optimize non-flagm2 comiss
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
79c745929a
FEXCore: invert CF internally
...
Flag day change to the ABI.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Alyssa Rosenzweig
16e6163677
OpcodeDispatcher: drop unused NZCVIndexMask
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Alyssa Rosenzweig
bbbc0dc9dc
OpcodeDispatcher: fix weird formatting
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Ryan Houdek
f8ef6feff9
AVX128: Optimize blends
...
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.
One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.
Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek
3c5b59d985
AVX128: Implement support for scalar FMA with AFP
...
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.
Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos
a1378f94ce
X87 Code Refactoring and Optimization Pass
2024-07-22 08:44:45 +02:00
Alyssa Rosenzweig
d20b46e46f
IR: drop LoadFlag/StoreFlag ops
...
pointless, we can just load/store the context now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-21 15:49:09 -04:00