Commit Graph
390 Commits
Author SHA1 Message Date
Alyssa Rosenzweig d17f33a922 OpcodeDispatcher: don't emit fake 0 for condjump
not needed and getting in the way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig 50a3ca0d6d IR: introduce dedicated PF/AF instructions
this makes reasoning about them a little easier, e.g. for flags.  about 1% win
in nodejs.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig e9ab514962 IR: push parity evaluation down
so we can optimize it globally

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig 4eb0948451 IR: push down AXFLAG lowering
so we can get the new axflag optimizations on billy's x13s.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig 42f2851575 Merge pull request #4009 from alyssarosenzweig/opt/axflag
Optimize AXFLAG-less systems
2024-08-27 08:01:05 -04:00
Alyssa Rosenzweig 8d6b454455 OpcodeDispatcher: optimize AXFLAG emulation
this should help on x13s which has flagm but not flagm2.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:38:03 -04:00
Alyssa Rosenzweig 205ec3e14d OpcodeDispatcher: refactor axflag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:35:27 -04:00
Ryan Houdek 8bf4a124c8 AVX128: Fixes 256-bit float compares
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.

PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek 1d00ad6030 OpcodeDispatcher: Convert VectorVariableBlend to Bind handler 2024-08-22 15:09:44 -07:00
Ryan Houdek 23a076c313 OpcodeDispatcher: Convert packed vector shifts to Bind handler 2024-08-22 15:07:09 -07:00
Ryan Houdek 57eacab654 OpcodeDispatcher: Convert VBROADCASTOp to Bind handler 2024-08-22 14:59:25 -07:00
Ryan Houdek b4093a8888 OpcodeDispatcher: Convert packed HSub to Bind handler 2024-08-22 14:57:40 -07:00
Ryan Houdek 7c5a9b5d6a OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler 2024-08-22 14:55:34 -07:00
Ryan Houdek 12b3c82d83 OpcodeDispatcher: Convert PExtr to Bind handler 2024-08-22 14:54:26 -07:00
Ryan Houdek 66520bce0a OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler 2024-08-22 14:52:41 -07:00
Ryan Houdek 25cc2bdcb8 OpcodeDispatcher: Convert VPERMILImm to Bind handler 2024-08-22 14:51:02 -07:00
Ryan Houdek f98b18800c OpcodeDispatcher: Convert MOVMSK to Bind handler 2024-08-22 14:49:57 -07:00
Ryan Houdek 0dd687a7a1 OpcodeDispatcher: Convert SHUFOp to Bind handler 2024-08-22 14:47:30 -07:00
Ryan Houdek 1aff3acbb9 OpcodeDispatcher: Convert PSHUFW to Bind handler 2024-08-22 14:45:04 -07:00
Ryan Houdek a6ab2ca30d OpcodeDispatcher: Convert PUNPCKH to Bind handler 2024-08-22 14:42:24 -07:00
Ryan Houdek ca43e2a61c OpcodeDispatcher: Convert PUNPCKL to Bind handler 2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig b03c613fbf OpcodeDispatcher: optimize JP/JNP
fuse and+cbnz into tbz/tbnz.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig 5d613e8716 OpcodeDispatcher: optimize test x, x
fewer uops now that we invert carry

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 21:17:08 -04:00
Alyssa Rosenzweig d7a20fa28f OpcodeDispatcher: fix tso checks
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 20:11:06 -04:00
Alyssa Rosenzweig 4c4c6e7807 OpcodeDispatcher: pair loads on cortex
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:36:50 -04:00
Alyssa Rosenzweig a8c9c71ce3 OpcodeDispatcher: use loadcontextpair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 32d6daf558 OpcodeDispatcher: use ldp/stp for AVX load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig bebcb73c68 OpcodeDispatcher: pair AVX high writes
reduces instr count with AVX-128

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig c70e44cf05 OpcodeDispatcher: extract Push helper
Mirrors the Pop helper we added. This cleans up a bunch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig 900c62fa7b OpcodeDispatcher: add Pop helpers
hide away the allocate dance

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig cab02be637 IR: remove unused pair create/extract
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig 5631ff4fd5 OpcodeDispatcher: optimize variable shifts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig a4545f493e OpcodeDispatcher: add IncrementByCarry helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig f138d7d9b8 OpcodeDispatcher: optimize JA/JNA
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 3bb9d44bf5 OpcodeDispatcher: optimize SetCFDirect
generally same # of instructions, but potentially fewer cycles:

old:
        "rmif x4, #63, #nzCv",
        "cfinv"

new:
        "xor x20, x4, #1",
        "rmif x20, #63, #nzCv"

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 06e6a1b19e OpcodeDispatcher: optimize logic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 9621eca677 OpcodeDispatcher: fix setrflag masking
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 1dc22e33ae OpcodeDispatcher: inline some flag calculations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 0ef8aaebeb OpcodeDispatcher: optimize BZHI
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig d17c427e47 OpcodeDispatcher: optimize non-flagm2 comiss
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 79c745929a FEXCore: invert CF internally
Flag day change to the ABI.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 8330cc6876 OpcodeDispatcher: defer carry inverts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:05:21 -04:00
Alyssa Rosenzweig 16e6163677 OpcodeDispatcher: drop unused NZCVIndexMask
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 11:26:12 -04:00
Alyssa Rosenzweig bbbc0dc9dc OpcodeDispatcher: fix weird formatting
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 11:26:12 -04:00
Mai 4882f10536 Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws be4777110c OpcodeDispatcher: Don't apply the address-size flag to segment addresses
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Ryan Houdek f8ef6feff9 AVX128: Optimize blends
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.

One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.

Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek 3c5b59d985 AVX128: Implement support for scalar FMA with AFP
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.

Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos a1378f94ce X87 Code Refactoring and Optimization Pass 2024-07-22 08:44:45 +02:00
Alyssa Rosenzweig d20b46e46f IR: drop LoadFlag/StoreFlag ops
pointless, we can just load/store the context now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:09 -04:00