Alyssa Rosenzweig
b03c613fbf
OpcodeDispatcher: optimize JP/JNP
...
fuse and+cbnz into tbz/tbnz.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig
5d613e8716
OpcodeDispatcher: optimize test x, x
...
fewer uops now that we invert carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 21:17:08 -04:00
Alyssa Rosenzweig
d7a20fa28f
OpcodeDispatcher: fix tso checks
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 20:11:06 -04:00
Alyssa Rosenzweig
4c4c6e7807
OpcodeDispatcher: pair loads on cortex
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:36:50 -04:00
Alyssa Rosenzweig
a8c9c71ce3
OpcodeDispatcher: use loadcontextpair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
32d6daf558
OpcodeDispatcher: use ldp/stp for AVX load/store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
bebcb73c68
OpcodeDispatcher: pair AVX high writes
...
reduces instr count with AVX-128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
c70e44cf05
OpcodeDispatcher: extract Push helper
...
Mirrors the Pop helper we added. This cleans up a bunch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
900c62fa7b
OpcodeDispatcher: add Pop helpers
...
hide away the allocate dance
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
cab02be637
IR: remove unused pair create/extract
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
5631ff4fd5
OpcodeDispatcher: optimize variable shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
f138d7d9b8
OpcodeDispatcher: optimize JA/JNA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
3bb9d44bf5
OpcodeDispatcher: optimize SetCFDirect
...
generally same # of instructions, but potentially fewer cycles:
old:
"rmif x4, #63 , #nzCv",
"cfinv"
new:
"xor x20, x4, #1 ",
"rmif x20, #63 , #nzCv"
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
06e6a1b19e
OpcodeDispatcher: optimize logic ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
9621eca677
OpcodeDispatcher: fix setrflag masking
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
d17c427e47
OpcodeDispatcher: optimize non-flagm2 comiss
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
79c745929a
FEXCore: invert CF internally
...
Flag day change to the ABI.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Alyssa Rosenzweig
16e6163677
OpcodeDispatcher: drop unused NZCVIndexMask
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Alyssa Rosenzweig
bbbc0dc9dc
OpcodeDispatcher: fix weird formatting
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Ryan Houdek
f8ef6feff9
AVX128: Optimize blends
...
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.
One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.
Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek
3c5b59d985
AVX128: Implement support for scalar FMA with AFP
...
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.
Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos
a1378f94ce
X87 Code Refactoring and Optimization Pass
2024-07-22 08:44:45 +02:00
Alyssa Rosenzweig
d20b46e46f
IR: drop LoadFlag/StoreFlag ops
...
pointless, we can just load/store the context now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-21 15:49:09 -04:00
Ryan Houdek
b0bd8a62a2
AVX128: Improve VPERMILPS/PD and VPSHUFD
...
VPSHUFD and VPERMILPS are aliases of each other.
Reuses the implementation path from the PSHUFD implementation which has
a few swizzles and then a table lookup.
VPERMILPD is a very simple swizzle per 128-bit lane.
Fixes #3797
Fixes #3784
2024-07-18 04:10:58 -07:00
Alyssa Rosenzweig
1e709d1150
OpcodeDispatcher: add RecordX87 helper
...
calls will be generated.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-16 09:07:35 +02:00
Ryan Houdek
d79b7fcc49
Merge pull request #3808 from alyssarosenzweig/rclse/3
...
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek
b9a6caea8d
Merge pull request #3844 from Sonicadvance1/fix_vmovq
...
AVX128: Fixes vmovq loading too much data
2024-07-12 17:07:32 -07:00
Ryan Houdek
8021dc10a1
OpcodeDispatcher: Force noinline for the function call in the Bind helper
...
Clang was inlining a few of the functions it was calling. So force it
never to inline since this is supports to be a little shim trampoline
only.
2024-07-11 19:00:42 -07:00
Ryan Houdek
7e8d734e43
AVX256: Initial fixes just to get my unittest working
...
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Alyssa Rosenzweig
294f10fdd0
OpcodeDispatcher: reg cache mmx
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-11 13:21:14 -04:00
Tony Wasserka
b9829ed316
OpcodeDispatcher: Replace even more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
4ccec17676
OpcodeDispatcher: Replace more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
f45082043b
OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility
2024-07-11 16:19:14 +02:00
Alyssa Rosenzweig
a4f8bbff02
OpcodeDispatcher: reg cache avx high
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
cf5ab05b90
OpcodeDispatcher: reg cache AbridgedFTW
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
3a2ce240f9
OpcodeDispatcher: reg cache DF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
1f01dd53f7
OpcodeDispatcher: reg cache fprs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
72d41d70b6
OpcodeDispatcher: introduce GPR-only reg cache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
2949bc211d
OpcodeDispatcher: thunk through FlushRegisterCache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Tony Wasserka
441187470e
OpcodeDispatcher: Avoid monomorphization of some AVX functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
59fd13cc2f
OpcodeDispatcher: Avoid monomorphization of even more functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
c9e7bfdf16
OpcodeDispatcher: Avoid monomorphization of more functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
2d700c381e
OpcodeDispatcher: Avoid monomorphization of large functions
2024-07-10 17:01:30 +02:00
Ryan Houdek
72d6c8ebd6
Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
...
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00