Alyssa Rosenzweig
6d4693cbc1
IR: fix scalar FMA tied sources
...
needs to be modelled explicitly or else we lose information when translating
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-04 07:42:44 -04:00
Mai
2478abba29
Merge pull request #4008 from Sonicadvance1/fix_avx128_vfcmp
...
AVX128: Fixes 256-bit float compares
2024-08-25 12:37:04 -04:00
Ryan Houdek
8bf4a124c8
AVX128: Fixes 256-bit float compares
...
Just like the bug in #4006 , we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.
PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek
7e6ba184f9
AVX128: Fixes incorrect size usage in AVX128_Vector_CVT_Int_To_Float
...
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".
The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.
Fixes Aperture Desk Job in FEX.
2024-08-25 05:18:53 -07:00
Ryan Houdek
1d00ad6030
OpcodeDispatcher: Convert VectorVariableBlend to Bind handler
2024-08-22 15:09:44 -07:00
Ryan Houdek
23a076c313
OpcodeDispatcher: Convert packed vector shifts to Bind handler
2024-08-22 15:07:09 -07:00
Ryan Houdek
57eacab654
OpcodeDispatcher: Convert VBROADCASTOp to Bind handler
2024-08-22 14:59:25 -07:00
Ryan Houdek
b4093a8888
OpcodeDispatcher: Convert packed HSub to Bind handler
2024-08-22 14:57:40 -07:00
Ryan Houdek
7c5a9b5d6a
OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler
2024-08-22 14:55:34 -07:00
Ryan Houdek
12b3c82d83
OpcodeDispatcher: Convert PExtr to Bind handler
2024-08-22 14:54:26 -07:00
Ryan Houdek
66520bce0a
OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler
2024-08-22 14:52:41 -07:00
Ryan Houdek
25cc2bdcb8
OpcodeDispatcher: Convert VPERMILImm to Bind handler
2024-08-22 14:51:02 -07:00
Ryan Houdek
f98b18800c
OpcodeDispatcher: Convert MOVMSK to Bind handler
2024-08-22 14:49:57 -07:00
Ryan Houdek
0dd687a7a1
OpcodeDispatcher: Convert SHUFOp to Bind handler
2024-08-22 14:47:30 -07:00
Ryan Houdek
1aff3acbb9
OpcodeDispatcher: Convert PSHUFW to Bind handler
2024-08-22 14:45:04 -07:00
Ryan Houdek
a6ab2ca30d
OpcodeDispatcher: Convert PUNPCKH to Bind handler
2024-08-22 14:42:24 -07:00
Ryan Houdek
ca43e2a61c
OpcodeDispatcher: Convert PUNPCKL to Bind handler
2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig
a8c9c71ce3
OpcodeDispatcher: use loadcontextpair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8a4bd5f22c
OpcodeDispatcher: pair avx128 save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
138a36c69a
OpcodeDispatcher: pair avx128 restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
3400ca5d42
OpcodeDispatcher: pair avx restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
5d8164da4e
OpcodeDispatcher: pair restore x87
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8b89b30a6f
OpcodeDispatcher: pair restore SSE
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
e66d7cfd6c
OpcodeDispatcher: pair mm save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
63afa29dae
OpcodeDispatcher: pair save AVX
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
926a9b40e2
OpcodeDispatcher: pair save mxcsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
9bb43264b5
OpcodeDispatcher: pair SSE store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig
32d6daf558
OpcodeDispatcher: use ldp/stp for AVX load/store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
34301319bf
OpcodeDispatcher: optimize IncrementByCarry
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
8eac3198b6
OpcodeDispatcher: use carry increment for ADC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
832edd4da3
OpcodeDispatcher: use carry increment for SBC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
630285c589
OpcodeDispatcher: optimize SetPackedRFLAG
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
7eee50d929
OpcodeDispatcher: optimize mul/umul
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
f73fb62c6e
OpcodeDispatcher: optimize V(P)TEST
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
5976b712ca
OpcodeDispatcher: optimize bl*
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
63f5e64adb
OpcodeDispatcher: optimize add/sub
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Ryan Houdek
e613876e9d
AVX128: Optimize all cases of vpermq
...
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.
Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.
Fixes #3795
2024-08-06 09:08:30 -07:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Paulo Matos
5933a59c09
Intersperse flag retrieval and FSW insertion
2024-07-31 12:04:54 +02:00
Paulo Matos
c2136272bf
Reuse Top in ReconstructFSW_Helper
...
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek
87fbcf754d
Vector: Optimize pblendw
...
Using a brute force solver to add in more optimized code paths
- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations
Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek
d92b6a9ac4
Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
...
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek
c2092bfed0
Merge pull request #3893 from pmatos/FNINITFix
...
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Alyssa Rosenzweig
5ff09f5091
OpcodeDispatcher/X87: use less creative Refs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-26 14:30:42 -04:00
Paulo Matos
d1e36f264f
Fix call to FNINITF64 and refactor
2024-07-26 14:56:06 +02:00