Ryan Houdek
5d8d052a77
Merge pull request #5640 from lioncash/move
...
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka
7fa4d78269
Merge pull request #5635 from Sonicadvance1/178
...
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek
110313e7de
Merge pull request #5638 from lioncash/swap
...
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek
44e24c9e6b
Merge pull request #5637 from mrpippy/unicodestring
...
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks
6c58fef220
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism.
2026-06-30 15:37:55 -07:00
LC
3d593ce87d
Merge pull request #5626 from Sonicadvance1/177
...
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek
36e1b5107e
InstcountCI: Update
2026-06-30 15:01:46 -07:00
Ryan Houdek
f89123f489
unittests/ASM: Allow up to 3-bits of precision loss
2026-06-30 15:00:16 -07:00
Paulo Matos
b7280a765d
asm_tests: Re-optimize FYL2X for reduced precision x87 path
2026-06-30 15:00:16 -07:00
Paulo Matos
37b010795e
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 15:00:16 -07:00
Ryan Houdek
a4f89b79a3
Merge pull request #5636 from lioncash/move
...
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek
5bde4d875a
FEXOfflineCompiler: Fixes HostFeature detection under Win32
...
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
WINE
Ensure that when built for Win32 that it uses the correct feature
fetching.
One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek
a79c471c31
Merge pull request #5634 from lioncash/shuffle
...
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
Ryan Houdek
417bd8604c
Merge pull request #5632 from lioncash/vpblendd_test
...
unittests: Add test for stress-testing VPBLENDD selectors
2026-06-29 22:13:05 -07:00
LC
3d289f4489
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
...
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.
Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek
d138c854f3
Merge pull request #5631 from lioncash/vpblendd
...
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek
e26a792b70
Merge pull request #5630 from lioncash/comiss
...
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek
7e2d3b07c0
Merge pull request #5629 from lioncash/comment
...
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek
394a6f28db
Merge pull request #5628 from lioncash/movmsk
...
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek
72e01274c9
Merge pull request #5627 from lioncash/vpblendw
...
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek
9f2e982944
Merge pull request #5617 from simon902/vcvtps2ph
...
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC
c6b0f360fe
AVX: Make use of table swapping constant to trim down relevant ops
...
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC
ef19242be3
IR: Add constant for swapping midsections of 256-bit vectors around
...
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
LC
2cb4f8b6f5
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
...
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC
59919a0b0c
unittests: Expand VPBLENDW selector test to cover 128-bit paths
...
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 07:03:16 -04:00
LC
bf1857ecc7
unittests: Expand VPBLENDD selector test to cover 128-bit paths
...
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 06:56:37 -04:00
LC
b40920596a
unittests: Add selector stress tests for VPSHUFD
...
Ensures any added optimization paths result in the same output.
2026-06-29 06:56:34 -04:00
LC
3b14c322e8
unittests: Add selector stress tests for VPSHUFLW/VPSHUFHW
...
Ensures any added optimization paths result in the same output.
2026-06-29 06:45:18 -04:00
LC
3e60aa5738
Merge pull request #5610 from Sonicadvance1/171
...
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer
6ca2d27c82
unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator
2026-06-29 08:00:58 +02:00
Simon Scherer
d2f26d1969
InstcountCI: Update
2026-06-29 07:55:05 +02:00
LC
fc25443827
unittests: Move full_vpblendw_imm test into the VEX folder
...
Keeps all of the selector tests in the same location.
2026-06-28 23:17:19 -04:00
LC
986800885a
unittests: Add test for stress-testing VPBLENDD selectors
...
Drops a test in like the one for VPERMQ to ensure that, even if different
optimization paths are introduced, the behavior remains consistent.
2026-06-28 23:13:39 -04:00
LC
52c6ab1cec
instcountci/VEX_map3: Add missing third param to VPBLENDD
...
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek
83a989c6dc
Merge pull request #5625 from lioncash/perm
...
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC
32b96c259b
Merge pull request #5623 from Sonicadvance1/176
...
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek
954581c750
Windows/UnixLib: Adds remaining helpers
...
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.
Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC
24da43f823
Merge pull request #5613 from Sonicadvance1/175
...
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC
600e2ddecf
Vector: Only signify 128-bit vector loads in UCOMISxOp
...
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.
No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek
9740f488cf
Merge pull request #5624 from lioncash/psad
...
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC
f81467fd30
instcountci/VEX_map1: Remove obsolete comments
...
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC
2ec2c39cf1
AVX: Lessen codegen for VMOVMSKPD
...
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek
c85e426cc5
Merge pull request #5622 from lioncash/perm
...
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC
41fc57f46c
AVX: Lessen codegen for 256-bit VMOVMSKPS
...
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC
a215bb9709
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
...
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek
b23fa30099
Merge pull request #5618 from wsxarcher/fixsmcfullvector
...
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek
848c4b2d68
Merge pull request #5621 from wsxarcher/fixfexfolder
...
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek
4f995cbc1c
Merge pull request #5620 from wsxarcher/fixuafbash
...
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher
b21c49352e
Core: Start a new block for the next op in full SMC check
2026-06-28 18:07:28 +02:00
Ryan Houdek
3470dd1e7b
Merge pull request #5619 from lioncash/dpps
...
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00