Commit Graph
567 Commits
Author SHA1 Message Date
Alyssa Rosenzweig 6d4693cbc1 IR: fix scalar FMA tied sources
needs to be modelled explicitly or else we lose information when translating

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-04 07:42:44 -04:00
Mai 2478abba29 Merge pull request #4008 from Sonicadvance1/fix_avx128_vfcmp
AVX128: Fixes 256-bit float compares
2024-08-25 12:37:04 -04:00
Ryan Houdek 8bf4a124c8 AVX128: Fixes 256-bit float compares
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.

PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek 7e6ba184f9 AVX128: Fixes incorrect size usage in AVX128_Vector_CVT_Int_To_Float
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".

The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.

Fixes Aperture Desk Job in FEX.
2024-08-25 05:18:53 -07:00
Ryan Houdek 1d00ad6030 OpcodeDispatcher: Convert VectorVariableBlend to Bind handler 2024-08-22 15:09:44 -07:00
Ryan Houdek 23a076c313 OpcodeDispatcher: Convert packed vector shifts to Bind handler 2024-08-22 15:07:09 -07:00
Ryan Houdek 57eacab654 OpcodeDispatcher: Convert VBROADCASTOp to Bind handler 2024-08-22 14:59:25 -07:00
Ryan Houdek b4093a8888 OpcodeDispatcher: Convert packed HSub to Bind handler 2024-08-22 14:57:40 -07:00
Ryan Houdek 7c5a9b5d6a OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler 2024-08-22 14:55:34 -07:00
Ryan Houdek 12b3c82d83 OpcodeDispatcher: Convert PExtr to Bind handler 2024-08-22 14:54:26 -07:00
Ryan Houdek 66520bce0a OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler 2024-08-22 14:52:41 -07:00
Ryan Houdek 25cc2bdcb8 OpcodeDispatcher: Convert VPERMILImm to Bind handler 2024-08-22 14:51:02 -07:00
Ryan Houdek f98b18800c OpcodeDispatcher: Convert MOVMSK to Bind handler 2024-08-22 14:49:57 -07:00
Ryan Houdek 0dd687a7a1 OpcodeDispatcher: Convert SHUFOp to Bind handler 2024-08-22 14:47:30 -07:00
Ryan Houdek 1aff3acbb9 OpcodeDispatcher: Convert PSHUFW to Bind handler 2024-08-22 14:45:04 -07:00
Ryan Houdek a6ab2ca30d OpcodeDispatcher: Convert PUNPCKH to Bind handler 2024-08-22 14:42:24 -07:00
Ryan Houdek ca43e2a61c OpcodeDispatcher: Convert PUNPCKL to Bind handler 2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig a8c9c71ce3 OpcodeDispatcher: use loadcontextpair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8a4bd5f22c OpcodeDispatcher: pair avx128 save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 138a36c69a OpcodeDispatcher: pair avx128 restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 3400ca5d42 OpcodeDispatcher: pair avx restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 5d8164da4e OpcodeDispatcher: pair restore x87
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8b89b30a6f OpcodeDispatcher: pair restore SSE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig e66d7cfd6c OpcodeDispatcher: pair mm save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 63afa29dae OpcodeDispatcher: pair save AVX
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 926a9b40e2 OpcodeDispatcher: pair save mxcsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 9bb43264b5 OpcodeDispatcher: pair SSE store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig 32d6daf558 OpcodeDispatcher: use ldp/stp for AVX load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 34301319bf OpcodeDispatcher: optimize IncrementByCarry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 8eac3198b6 OpcodeDispatcher: use carry increment for ADC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 832edd4da3 OpcodeDispatcher: use carry increment for SBC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig a4545f493e OpcodeDispatcher: add IncrementByCarry helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 630285c589 OpcodeDispatcher: optimize SetPackedRFLAG
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 7eee50d929 OpcodeDispatcher: optimize mul/umul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 1dc22e33ae OpcodeDispatcher: inline some flag calculations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 0ef8aaebeb OpcodeDispatcher: optimize BZHI
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig f73fb62c6e OpcodeDispatcher: optimize V(P)TEST
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 5976b712ca OpcodeDispatcher: optimize bl*
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 63f5e64adb OpcodeDispatcher: optimize add/sub
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 8330cc6876 OpcodeDispatcher: defer carry inverts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:05:21 -04:00
Mai 4882f10536 Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Ryan Houdek e613876e9d AVX128: Optimize all cases of vpermq
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.

Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.

Fixes #3795
2024-08-06 09:08:30 -07:00
Billy Laws be4777110c OpcodeDispatcher: Don't apply the address-size flag to segment addresses
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Paulo Matos 5933a59c09 Intersperse flag retrieval and FSW insertion 2024-07-31 12:04:54 +02:00
Paulo Matos c2136272bf Reuse Top in ReconstructFSW_Helper
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek 87fbcf754d Vector: Optimize pblendw
Using a brute force solver to add in more optimized code paths

- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations

Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek d92b6a9ac4 Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek c2092bfed0 Merge pull request #3893 from pmatos/FNINITFix
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Alyssa Rosenzweig 5ff09f5091 OpcodeDispatcher/X87: use less creative Refs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-26 14:30:42 -04:00
Paulo Matos d1e36f264f Fix call to FNINITF64 and refactor 2024-07-26 14:56:06 +02:00