Ryan Houdek
23a076c313
OpcodeDispatcher: Convert packed vector shifts to Bind handler
2024-08-22 15:07:09 -07:00
Ryan Houdek
57eacab654
OpcodeDispatcher: Convert VBROADCASTOp to Bind handler
2024-08-22 14:59:25 -07:00
Ryan Houdek
b4093a8888
OpcodeDispatcher: Convert packed HSub to Bind handler
2024-08-22 14:57:40 -07:00
Ryan Houdek
7c5a9b5d6a
OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler
2024-08-22 14:55:34 -07:00
Ryan Houdek
12b3c82d83
OpcodeDispatcher: Convert PExtr to Bind handler
2024-08-22 14:54:26 -07:00
Ryan Houdek
66520bce0a
OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler
2024-08-22 14:52:41 -07:00
Ryan Houdek
25cc2bdcb8
OpcodeDispatcher: Convert VPERMILImm to Bind handler
2024-08-22 14:51:02 -07:00
Ryan Houdek
f98b18800c
OpcodeDispatcher: Convert MOVMSK to Bind handler
2024-08-22 14:49:57 -07:00
Ryan Houdek
0dd687a7a1
OpcodeDispatcher: Convert SHUFOp to Bind handler
2024-08-22 14:47:30 -07:00
Ryan Houdek
1aff3acbb9
OpcodeDispatcher: Convert PSHUFW to Bind handler
2024-08-22 14:45:04 -07:00
Ryan Houdek
a6ab2ca30d
OpcodeDispatcher: Convert PUNPCKH to Bind handler
2024-08-22 14:42:24 -07:00
Ryan Houdek
ca43e2a61c
OpcodeDispatcher: Convert PUNPCKL to Bind handler
2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig
a8c9c71ce3
OpcodeDispatcher: use loadcontextpair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8a4bd5f22c
OpcodeDispatcher: pair avx128 save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
138a36c69a
OpcodeDispatcher: pair avx128 restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
3400ca5d42
OpcodeDispatcher: pair avx restore
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
5d8164da4e
OpcodeDispatcher: pair restore x87
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
8b89b30a6f
OpcodeDispatcher: pair restore SSE
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
e66d7cfd6c
OpcodeDispatcher: pair mm save
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
63afa29dae
OpcodeDispatcher: pair save AVX
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
926a9b40e2
OpcodeDispatcher: pair save mxcsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig
9bb43264b5
OpcodeDispatcher: pair SSE store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig
32d6daf558
OpcodeDispatcher: use ldp/stp for AVX load/store
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig
34301319bf
OpcodeDispatcher: optimize IncrementByCarry
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
8eac3198b6
OpcodeDispatcher: use carry increment for ADC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
832edd4da3
OpcodeDispatcher: use carry increment for SBC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
630285c589
OpcodeDispatcher: optimize SetPackedRFLAG
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
7eee50d929
OpcodeDispatcher: optimize mul/umul
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
f73fb62c6e
OpcodeDispatcher: optimize V(P)TEST
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
5976b712ca
OpcodeDispatcher: optimize bl*
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
63f5e64adb
OpcodeDispatcher: optimize add/sub
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Ryan Houdek
e613876e9d
AVX128: Optimize all cases of vpermq
...
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.
Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.
Fixes #3795
2024-08-06 09:08:30 -07:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Paulo Matos
5933a59c09
Intersperse flag retrieval and FSW insertion
2024-07-31 12:04:54 +02:00
Paulo Matos
c2136272bf
Reuse Top in ReconstructFSW_Helper
...
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek
87fbcf754d
Vector: Optimize pblendw
...
Using a brute force solver to add in more optimized code paths
- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations
Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek
d92b6a9ac4
Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
...
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek
c2092bfed0
Merge pull request #3893 from pmatos/FNINITFix
...
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Alyssa Rosenzweig
5ff09f5091
OpcodeDispatcher/X87: use less creative Refs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-26 14:30:42 -04:00
Paulo Matos
d1e36f264f
Fix call to FNINITF64 and refactor
2024-07-26 14:56:06 +02:00
Ryan Houdek
7816b150d0
FEXCore: Removes CPUBackendFeatures
...
We were only ever hardcoding true for TBL2 and Flags now. Get rid of it.
2024-07-24 17:19:30 -07:00
Ryan Houdek
f8ef6feff9
AVX128: Optimize blends
...
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.
One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.
Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek
3c5b59d985
AVX128: Implement support for scalar FMA with AFP
...
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.
Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos
a1378f94ce
X87 Code Refactoring and Optimization Pass
2024-07-22 08:44:45 +02:00
Ryan Houdek
f8c6baae97
Merge pull request #3883 from Sonicadvance1/implement_daz
...
Arm64: Implements support for DAZ using AFP.FIZ
2024-07-21 10:03:34 -07:00