Commit Graph
14454 Commits
Author SHA1 Message Date
Ryan Houdek d138c854f3 Merge pull request #5631 from lioncash/vpblendd
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek e26a792b70 Merge pull request #5630 from lioncash/comiss
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek 7e2d3b07c0 Merge pull request #5629 from lioncash/comment
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek 394a6f28db Merge pull request #5628 from lioncash/movmsk
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek 72e01274c9 Merge pull request #5627 from lioncash/vpblendw
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 6ca2d27c82 unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator 2026-06-29 08:00:58 +02:00
Simon Scherer d2f26d1969 InstcountCI: Update 2026-06-29 07:55:05 +02:00
LC 52c6ab1cec instcountci/VEX_map3: Add missing third param to VPBLENDD
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek 83a989c6dc Merge pull request #5625 from lioncash/perm
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC 32b96c259b Merge pull request #5623 from Sonicadvance1/176
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek 954581c750 Windows/UnixLib: Adds remaining helpers
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.

Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC 24da43f823 Merge pull request #5613 from Sonicadvance1/175
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek 9740f488cf Merge pull request #5624 from lioncash/psad
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC f81467fd30 instcountci/VEX_map1: Remove obsolete comments
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC a215bb9709 instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek 848c4b2d68 Merge pull request #5621 from wsxarcher/fixfexfolder
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek 4f995cbc1c Merge pull request #5620 from wsxarcher/fixuafbash
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
Ryan Houdek 3470dd1e7b Merge pull request #5619 from lioncash/dpps
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00
wsxarcher 2ba35286ef Do not exec FEX if it is a folder 2026-06-28 17:28:41 +02:00
wsxarcher bbd8212ced Fix UAF of PS1 2026-06-28 17:18:05 +02:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
Simon Scherer b9052ed7f0 unittests/ASM: Test zeroing of vcvtps2ph 2026-06-28 13:14:26 +02:00
Ryan Houdek 4160a92621 Windows: Fixes duplicated hardware TSO handling
This was handled in both the Module.cpp files and also the common
TSOHandlerConfig on accident. Wouldn't have caused an issue but it was
definitely a bit weird.
2026-06-27 20:59:15 -07:00
Ryan Houdek 201bb73980 Windows/UnixLib: Adds support for Hardware TSO support
Including fallback to non-unixlib path because we need to support both.

Showcases how these are going to be implemented without throwing the
entire world at it right away. Next PR will be implementing the
remaining four necessary unixlib handlers that we will require:

- Kernel unaligned atomic control
- shm_stats thing
- madvise operation
- prctl vma naming
2026-06-27 19:59:35 -07:00
LC 78832cc0d0 Merge pull request #5612 from Sonicadvance1/174
Windows: Load unixlib if possible
2026-06-27 22:25:43 -04:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek 70fe9a4405 Merge pull request #5614 from lioncash/whoops
Vector: Fix typo in VPERMQOp
2026-06-26 21:45:33 -07:00
Ryan Houdek c09225f868 Windows: Load unixlib if possible
Currently does nothing other than load it (as the library also doesn't
do anything yet). Ensured it was working by temporarily creating a test
entrypoint and doing `Call` on to it.

Next step after this is to reimplement some of the nasty hacks FEX is
doing inside the unixlib code itself.
2026-06-26 20:50:24 -07:00
Ryan Houdek 126bcd365d winternl: Update enums
Newer WINE has a better mechanism for asking to load unix libraries.
Older WINE like what is in Proton doesn't have this yet. Add definitions
for both so we can try either one.
2026-06-26 20:50:19 -07:00
LC fe4d2bc6c5 Merge pull request #5611 from Sonicadvance1/173
Windows: Adds empty Linux side unix library
2026-06-26 23:48:51 -04:00
Ryan Houdek dbaf22372c Windows: Adds empty Linux side unix library
We are going to need a unix library. Going to take this one step at a
time without AI/ML so I fully understand all the pieces of the puzzle,
and to ensure we don't lose any functionality before we're ready.

This only ensures that we are building the Linux facing .so files for
arm64ec and wow64, but they are empty today. Next PR will be
initializing it on the PE side.
2026-06-26 19:05:08 -07:00
LC 43bd243457 Merge pull request #5609 from Sonicadvance1/170
OpcodeDispatcher: Fixes CRC32 with high 8-bit register
2026-06-26 15:30:00 -04:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
Ryan Houdek 26266c6a94 InstcountCI: Adds CRC32 with high 8-bit register 2026-06-26 11:38:08 -07:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Simon Scherer 0c1a35f297 unittests/ASM: Add unit test for crc32 with 8bit register operand 2026-06-26 11:10:06 -07:00
Tony Wasserka 9ac608ca43 Merge pull request #5606 from Sonicadvance1/169
LibraryForwarding/cuda: Convert constexpr to const
2026-06-26 11:17:49 +02:00