mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-10 16:00:31 +02:00
When loading 256-bits of data with only 128-bits of address indices, we can sign extend the source indices to be 64-bit. Thus falling down the ideal path for SVE where each 128-bit lane is loading the data to addresses in a 1:1 element ratio. This means we use the SVE path more often because of this. Based on top of #3825 because the prescaling behaviour was introduced there. This implements its own prescaling when the sign extension occurs because ARM's SSHLL{,2} instruction gives us that for free. This additionally fixes a bug where we were accidentally loading the top 128-bit half of the addresses for gathers when it was unnecessary, and on the AVX256 side it was duplicating and doing some additional work when it shouldn't have. It'll be good to walk the commits when looking at this one, as there are a couple of incremental changes that are easier to follow that way. Fixes #3806