Commit Graph
552 Commits
Author SHA1 Message Date
Ryan Houdek a6ab2ca30d OpcodeDispatcher: Convert PUNPCKH to Bind handler 2024-08-22 14:42:24 -07:00
Ryan Houdek ca43e2a61c OpcodeDispatcher: Convert PUNPCKL to Bind handler 2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig a8c9c71ce3 OpcodeDispatcher: use loadcontextpair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8a4bd5f22c OpcodeDispatcher: pair avx128 save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 138a36c69a OpcodeDispatcher: pair avx128 restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 3400ca5d42 OpcodeDispatcher: pair avx restore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 5d8164da4e OpcodeDispatcher: pair restore x87
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 8b89b30a6f OpcodeDispatcher: pair restore SSE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig e66d7cfd6c OpcodeDispatcher: pair mm save
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 63afa29dae OpcodeDispatcher: pair save AVX
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 926a9b40e2 OpcodeDispatcher: pair save mxcsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:41 -04:00
Alyssa Rosenzweig 9bb43264b5 OpcodeDispatcher: pair SSE store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:33:37 -04:00
Alyssa Rosenzweig 32d6daf558 OpcodeDispatcher: use ldp/stp for AVX load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-20 18:31:19 -04:00
Alyssa Rosenzweig 34301319bf OpcodeDispatcher: optimize IncrementByCarry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 8eac3198b6 OpcodeDispatcher: use carry increment for ADC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 832edd4da3 OpcodeDispatcher: use carry increment for SBC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig a4545f493e OpcodeDispatcher: add IncrementByCarry helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig 630285c589 OpcodeDispatcher: optimize SetPackedRFLAG
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 7eee50d929 OpcodeDispatcher: optimize mul/umul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 1dc22e33ae OpcodeDispatcher: inline some flag calculations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 0ef8aaebeb OpcodeDispatcher: optimize BZHI
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig f73fb62c6e OpcodeDispatcher: optimize V(P)TEST
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 5976b712ca OpcodeDispatcher: optimize bl*
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 63f5e64adb OpcodeDispatcher: optimize add/sub
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig 8330cc6876 OpcodeDispatcher: defer carry inverts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-10 13:05:21 -04:00
Mai 4882f10536 Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Ryan Houdek e613876e9d AVX128: Optimize all cases of vpermq
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.

Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.

Fixes #3795
2024-08-06 09:08:30 -07:00
Billy Laws be4777110c OpcodeDispatcher: Don't apply the address-size flag to segment addresses
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Paulo Matos 5933a59c09 Intersperse flag retrieval and FSW insertion 2024-07-31 12:04:54 +02:00
Paulo Matos c2136272bf Reuse Top in ReconstructFSW_Helper
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek 87fbcf754d Vector: Optimize pblendw
Using a brute force solver to add in more optimized code paths

- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations

Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek d92b6a9ac4 Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek c2092bfed0 Merge pull request #3893 from pmatos/FNINITFix
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Alyssa Rosenzweig 5ff09f5091 OpcodeDispatcher/X87: use less creative Refs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-26 14:30:42 -04:00
Paulo Matos d1e36f264f Fix call to FNINITF64 and refactor 2024-07-26 14:56:06 +02:00
Ryan Houdek 7816b150d0 FEXCore: Removes CPUBackendFeatures
We were only ever hardcoding true for TBL2 and Flags now. Get rid of it.
2024-07-24 17:19:30 -07:00
Ryan Houdek f8ef6feff9 AVX128: Optimize blends
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.

One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.

Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek 3c5b59d985 AVX128: Implement support for scalar FMA with AFP
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.

Fixes #3793
2024-07-22 12:58:19 -07:00
Paulo Matos a1378f94ce X87 Code Refactoring and Optimization Pass 2024-07-22 08:44:45 +02:00
Ryan Houdek f8c6baae97 Merge pull request #3883 from Sonicadvance1/implement_daz
Arm64: Implements support for DAZ using AFP.FIZ
2024-07-21 10:03:34 -07:00
Ryan Houdek b78da2e5ad Arm64: Implements support for DAZ using AFP.FIZ
When AFP is supported then we can actually support DAZ. This might also
fix the audio corruption in Animal Well but I can't test it until Steam
is running on Oryon. Requires a bit of plumbing for MXCSR which we were
hacking around before but now we actually want to store the value.

Fixes #3856
2024-07-20 15:34:54 -07:00
Ryan Houdek 1c35eeffeb Vector: Optimize PSHUFD with brute force search
With a brute force search of methods between 1-3 instructions we cover a
lot more cases more optimally.

There's definitely still more cases (and probably some that can reduce
from 3 instruction to 2), but covering 44 cases is a pretty good margin
already.
2024-07-18 04:10:58 -07:00
Ryan Houdek b0bd8a62a2 AVX128: Improve VPERMILPS/PD and VPSHUFD
VPSHUFD and VPERMILPS are aliases of each other.

Reuses the implementation path from the PSHUFD implementation which has
a few swizzles and then a table lookup.

VPERMILPD is a very simple swizzle per 128-bit lane.

Fixes #3797
Fixes #3784
2024-07-18 04:10:58 -07:00
Ryan Houdek da51169ba9 Merge pull request #3875 from alyssarosenzweig/ir/gethostflag
IR: garbage collect premature F80Cmp optimizations
2024-07-17 03:05:48 -07:00
Alyssa Rosenzweig e7d5a01c5f IR: remove F80Cmp flags
nothing is optimizing around this, it's just adding pointless complexity. if we
want to actually optimize F80Cmp, the right way would be to lift the
implementation into the OpcodeDispatcher or JIT. it wouldn't be terribly
difficult. This kludge doesn't get us closer there.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Alyssa Rosenzweig 0c3a8d0bc8 IR: remove GetHostFlag
it doesn't get host flags, it's just an extra Bfe used in x87. pointless and
confusing!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:44:34 -04:00
Alyssa Rosenzweig c4ba7eee87 X87: save uop in ReconstructFTW
noticed while reviewing Paulo's work

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 13:54:09 -04:00
Ryan Houdek d79b7fcc49 Merge pull request #3808 from alyssarosenzweig/rclse/3
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek 7e8d734e43 AVX256: Initial fixes just to get my unittest working
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Ryan Houdek 3c7318d7c8 AVX128: Fixes vmovq loading too much data
This was doing a 128-bit load from memory and then a 64-bit zero extend
which looked like a spurious move but it was trying to match the
behaviour of vmovq where it needed the zero extend.

Also adds a unit test to ensure that we aren't loading too much data by
loading right up against a page boundary.

Fixes #3787
2024-07-11 18:34:05 -07:00