Commit Graph
14478 Commits
Author SHA1 Message Date
Ryan Houdek 5d8d052a77 Merge pull request #5640 from lioncash/move
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka 7fa4d78269 Merge pull request #5635 from Sonicadvance1/178
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek 110313e7de Merge pull request #5638 from lioncash/swap
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek 44e24c9e6b Merge pull request #5637 from mrpippy/unicodestring
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks 6c58fef220 Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism. 2026-06-30 15:37:55 -07:00
LC 3d593ce87d Merge pull request #5626 from Sonicadvance1/177
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek 36e1b5107e InstcountCI: Update 2026-06-30 15:01:46 -07:00
Ryan Houdek f89123f489 unittests/ASM: Allow up to 3-bits of precision loss 2026-06-30 15:00:16 -07:00
Paulo Matos b7280a765d asm_tests: Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek a4f89b79a3 Merge pull request #5636 from lioncash/move
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek 5bde4d875a FEXOfflineCompiler: Fixes HostFeature detection under Win32
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
  differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
  WINE

Ensure that when built for Win32 that it uses the correct feature
fetching.

One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek a79c471c31 Merge pull request #5634 from lioncash/shuffle
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
Ryan Houdek 417bd8604c Merge pull request #5632 from lioncash/vpblendd_test
unittests: Add test for stress-testing VPBLENDD selectors
2026-06-29 22:13:05 -07:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek d138c854f3 Merge pull request #5631 from lioncash/vpblendd
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek e26a792b70 Merge pull request #5630 from lioncash/comiss
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek 7e2d3b07c0 Merge pull request #5629 from lioncash/comment
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek 394a6f28db Merge pull request #5628 from lioncash/movmsk
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek 72e01274c9 Merge pull request #5627 from lioncash/vpblendw
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 59919a0b0c unittests: Expand VPBLENDW selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 07:03:16 -04:00
LC bf1857ecc7 unittests: Expand VPBLENDD selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 06:56:37 -04:00
LC b40920596a unittests: Add selector stress tests for VPSHUFD
Ensures any added optimization paths result in the same output.
2026-06-29 06:56:34 -04:00
LC 3b14c322e8 unittests: Add selector stress tests for VPSHUFLW/VPSHUFHW
Ensures any added optimization paths result in the same output.
2026-06-29 06:45:18 -04:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 6ca2d27c82 unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator 2026-06-29 08:00:58 +02:00
Simon Scherer d2f26d1969 InstcountCI: Update 2026-06-29 07:55:05 +02:00
LC fc25443827 unittests: Move full_vpblendw_imm test into the VEX folder
Keeps all of the selector tests in the same location.
2026-06-28 23:17:19 -04:00
LC 986800885a unittests: Add test for stress-testing VPBLENDD selectors
Drops a test in like the one for VPERMQ to ensure that, even if different
optimization paths are introduced, the behavior remains consistent.
2026-06-28 23:13:39 -04:00
LC 52c6ab1cec instcountci/VEX_map3: Add missing third param to VPBLENDD
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek 83a989c6dc Merge pull request #5625 from lioncash/perm
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC 32b96c259b Merge pull request #5623 from Sonicadvance1/176
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek 954581c750 Windows/UnixLib: Adds remaining helpers
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.

Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC 24da43f823 Merge pull request #5613 from Sonicadvance1/175
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek 9740f488cf Merge pull request #5624 from lioncash/psad
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC f81467fd30 instcountci/VEX_map1: Remove obsolete comments
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC a215bb9709 instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek 848c4b2d68 Merge pull request #5621 from wsxarcher/fixfexfolder
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek 4f995cbc1c Merge pull request #5620 from wsxarcher/fixuafbash
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
Ryan Houdek 3470dd1e7b Merge pull request #5619 from lioncash/dpps
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00