Alyssa Rosenzweig
c70e44cf05
OpcodeDispatcher: extract Push helper
...
Mirrors the Pop helper we added. This cleans up a bunch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
900c62fa7b
OpcodeDispatcher: add Pop helpers
...
hide away the allocate dance
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
cab02be637
IR: remove unused pair create/extract
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
5631ff4fd5
OpcodeDispatcher: optimize variable shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
f138d7d9b8
OpcodeDispatcher: optimize JA/JNA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
3bb9d44bf5
OpcodeDispatcher: optimize SetCFDirect
...
generally same # of instructions, but potentially fewer cycles:
old:
"rmif x4, #63 , #nzCv",
"cfinv"
new:
"xor x20, x4, #1 ",
"rmif x20, #63 , #nzCv"
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
06e6a1b19e
OpcodeDispatcher: optimize logic ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
9621eca677
OpcodeDispatcher: fix setrflag masking
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
d17c427e47
OpcodeDispatcher: optimize non-flagm2 comiss
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
79c745929a
FEXCore: invert CF internally
...
Flag day change to the ABI.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Alyssa Rosenzweig
16e6163677
OpcodeDispatcher: drop unused NZCVIndexMask
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Alyssa Rosenzweig
bbbc0dc9dc
OpcodeDispatcher: fix weird formatting
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Ryan Houdek
f8ef6feff9
AVX128: Optimize blends
...
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.
One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.
Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek
3c5b59d985
AVX128: Implement support for scalar FMA with AFP
...
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.
Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos
a1378f94ce
X87 Code Refactoring and Optimization Pass
2024-07-22 08:44:45 +02:00
Alyssa Rosenzweig
d20b46e46f
IR: drop LoadFlag/StoreFlag ops
...
pointless, we can just load/store the context now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-21 15:49:09 -04:00
Ryan Houdek
b0bd8a62a2
AVX128: Improve VPERMILPS/PD and VPSHUFD
...
VPSHUFD and VPERMILPS are aliases of each other.
Reuses the implementation path from the PSHUFD implementation which has
a few swizzles and then a table lookup.
VPERMILPD is a very simple swizzle per 128-bit lane.
Fixes #3797
Fixes #3784
2024-07-18 04:10:58 -07:00
Alyssa Rosenzweig
1e709d1150
OpcodeDispatcher: add RecordX87 helper
...
calls will be generated.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-16 09:07:35 +02:00
Ryan Houdek
d79b7fcc49
Merge pull request #3808 from alyssarosenzweig/rclse/3
...
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek
b9a6caea8d
Merge pull request #3844 from Sonicadvance1/fix_vmovq
...
AVX128: Fixes vmovq loading too much data
2024-07-12 17:07:32 -07:00
Ryan Houdek
8021dc10a1
OpcodeDispatcher: Force noinline for the function call in the Bind helper
...
Clang was inlining a few of the functions it was calling. So force it
never to inline since this is supports to be a little shim trampoline
only.
2024-07-11 19:00:42 -07:00
Ryan Houdek
7e8d734e43
AVX256: Initial fixes just to get my unittest working
...
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Alyssa Rosenzweig
294f10fdd0
OpcodeDispatcher: reg cache mmx
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-11 13:21:14 -04:00
Tony Wasserka
b9829ed316
OpcodeDispatcher: Replace even more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
4ccec17676
OpcodeDispatcher: Replace more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
f45082043b
OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility
2024-07-11 16:19:14 +02:00
Alyssa Rosenzweig
a4f8bbff02
OpcodeDispatcher: reg cache avx high
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
cf5ab05b90
OpcodeDispatcher: reg cache AbridgedFTW
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
3a2ce240f9
OpcodeDispatcher: reg cache DF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
1f01dd53f7
OpcodeDispatcher: reg cache fprs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
72d41d70b6
OpcodeDispatcher: introduce GPR-only reg cache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
2949bc211d
OpcodeDispatcher: thunk through FlushRegisterCache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Tony Wasserka
441187470e
OpcodeDispatcher: Avoid monomorphization of some AVX functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
59fd13cc2f
OpcodeDispatcher: Avoid monomorphization of even more functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
c9e7bfdf16
OpcodeDispatcher: Avoid monomorphization of more functions
2024-07-10 17:01:30 +02:00
Tony Wasserka
2d700c381e
OpcodeDispatcher: Avoid monomorphization of large functions
2024-07-10 17:01:30 +02:00
Ryan Houdek
72d6c8ebd6
Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
...
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00
Ryan Houdek
ec7c8fd922
AVX128: Optimize QPS/QD variant of gather loads!
...
SVE has a special version of their gather instruction that gets similar
behaviour to x86's VGATHERQPS/VPGATHERQD instructions.
The quirk of these instructions that the previous SVE implementation
didn't handle and required ASIMD fallback, was that most gather
instructions require the data element size and address element size to
match. This x86 instruction uses a 64-bit address size while loading 32-bit
elements. This matches this specific variant of the SVE instruction, but
the data is zero-extended once loaded, requiring us to shuffle the data
after it is loaded.
This isn't the worst but the implementation is different enough that
stuffing it in to the other gather load will cause headaches.
Basically gets 32 instruction variants to use the SVE version!
Fixes #3827
2024-07-08 17:19:18 -07:00
Ryan Houdek
c5a0ae7b34
IR: Adds new QPS gather load variant!
2024-07-08 17:19:18 -07:00
Ryan Houdek
0d4414fdd0
AVX128: Removes templated AddrElementSize and add as argument
...
NFC
2024-07-06 18:32:35 -07:00
Alyssa Rosenzweig
adc709db2f
OpcodeDispatcher: drop remnants of deferred flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 17:22:41 -04:00
Alyssa Rosenzweig
395573720d
OpcodeDispatcher: drop pointless flag defers for shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
0e62759d24
OpcodeDispatcher: stop deferring logical
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
926b6c3117
OpcodeDispatcher: don't defer mul flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00