Compare commits

...
547 Commits
Author SHA1 Message Date
Ryan Houdek 1cc4b93e7a Docs: Update for release FEX-2607 2026-07-02 17:47:31 -07:00
LC b4e2f5118a Merge pull request #5645 from Sonicadvance1/180
Windows: Fixes SHM stats reallocation
2026-07-02 19:33:37 -04:00
Ryan Houdek 812b6398e5 Windows: Fixes SHM stats reallocation
This was accidentally setting `CurrentSize` instead of just returning
the newly allocated size to the frontend. This was causing the frontend
to then fail to detect the reallocation actually occured and no longer
get stats for new threads.

Also happened to not use `NewSize` but instead `CurrentSize * 2` which
didn't matter as it matched the growth pattern, but was technically
incorrect.

Fixes SHM stats since the introduction of the unixlib, ezpz.
2026-07-02 13:38:14 -07:00
LC 1db45e2a70 Merge pull request #5639 from Sonicadvance1/179
FEXServerClient: Workaround sun_path 108 byte limit
2026-07-02 04:03:48 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
Ryan Houdek 5f1c8efe0e Merge pull request #5643 from lioncash/rec
VectorOps: Avoid temporary if able in 256-bit VFRecp
2026-07-01 17:25:33 -07:00
Ryan Houdek b9aeccf13b Merge pull request #5642 from lioncash/feature
HostFeatures: Put SVE support querying into single function
2026-07-01 15:29:02 -07:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
Ryan Houdek 5d8d052a77 Merge pull request #5640 from lioncash/move
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka 7fa4d78269 Merge pull request #5635 from Sonicadvance1/178
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek 6bf0db7df6 FEXServerClient: Workaround sun_path 108 byte limit
We really don't want to do this, but in the case that the AF_UNIX path
is longer than the 108-byte limit that sun_path provides we don't really
have a choice. The alternative choice would be to switch /entirely/ away
from AF_UNIX and instead use pipes. We need a bandage fix for now, so
throw the socket in to a temp folder if the path is too long.
2026-06-30 17:31:23 -07:00
Ryan Houdek 110313e7de Merge pull request #5638 from lioncash/swap
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek 44e24c9e6b Merge pull request #5637 from mrpippy/unicodestring
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks 6c58fef220 Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism. 2026-06-30 15:37:55 -07:00
LC 3d593ce87d Merge pull request #5626 from Sonicadvance1/177
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek 36e1b5107e InstcountCI: Update 2026-06-30 15:01:46 -07:00
Ryan Houdek f89123f489 unittests/ASM: Allow up to 3-bits of precision loss 2026-06-30 15:00:16 -07:00
Paulo Matos b7280a765d asm_tests: Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek a4f89b79a3 Merge pull request #5636 from lioncash/move
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek 5bde4d875a FEXOfflineCompiler: Fixes HostFeature detection under Win32
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
  differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
  WINE

Ensure that when built for Win32 that it uses the correct feature
fetching.

One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek a79c471c31 Merge pull request #5634 from lioncash/shuffle
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
LC d3a85e14d9 VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
Ryan Houdek 417bd8604c Merge pull request #5632 from lioncash/vpblendd_test
unittests: Add test for stress-testing VPBLENDD selectors
2026-06-29 22:13:05 -07:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek d138c854f3 Merge pull request #5631 from lioncash/vpblendd
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek e26a792b70 Merge pull request #5630 from lioncash/comiss
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek 7e2d3b07c0 Merge pull request #5629 from lioncash/comment
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek 394a6f28db Merge pull request #5628 from lioncash/movmsk
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek 72e01274c9 Merge pull request #5627 from lioncash/vpblendw
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 59919a0b0c unittests: Expand VPBLENDW selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 07:03:16 -04:00
LC bf1857ecc7 unittests: Expand VPBLENDD selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 06:56:37 -04:00
LC b40920596a unittests: Add selector stress tests for VPSHUFD
Ensures any added optimization paths result in the same output.
2026-06-29 06:56:34 -04:00
LC 3b14c322e8 unittests: Add selector stress tests for VPSHUFLW/VPSHUFHW
Ensures any added optimization paths result in the same output.
2026-06-29 06:45:18 -04:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 6ca2d27c82 unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator 2026-06-29 08:00:58 +02:00
Simon Scherer d2f26d1969 InstcountCI: Update 2026-06-29 07:55:05 +02:00
LC fc25443827 unittests: Move full_vpblendw_imm test into the VEX folder
Keeps all of the selector tests in the same location.
2026-06-28 23:17:19 -04:00
LC 986800885a unittests: Add test for stress-testing VPBLENDD selectors
Drops a test in like the one for VPERMQ to ensure that, even if different
optimization paths are introduced, the behavior remains consistent.
2026-06-28 23:13:39 -04:00
LC 52c6ab1cec instcountci/VEX_map3: Add missing third param to VPBLENDD
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek 83a989c6dc Merge pull request #5625 from lioncash/perm
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC 32b96c259b Merge pull request #5623 from Sonicadvance1/176
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek 954581c750 Windows/UnixLib: Adds remaining helpers
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.

Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC 24da43f823 Merge pull request #5613 from Sonicadvance1/175
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek 9740f488cf Merge pull request #5624 from lioncash/psad
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC f81467fd30 instcountci/VEX_map1: Remove obsolete comments
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC a215bb9709 instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek 848c4b2d68 Merge pull request #5621 from wsxarcher/fixfexfolder
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek 4f995cbc1c Merge pull request #5620 from wsxarcher/fixuafbash
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
Ryan Houdek 3470dd1e7b Merge pull request #5619 from lioncash/dpps
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00
wsxarcher 2ba35286ef Do not exec FEX if it is a folder 2026-06-28 17:28:41 +02:00
wsxarcher bbd8212ced Fix UAF of PS1 2026-06-28 17:18:05 +02:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
Simon Scherer b9052ed7f0 unittests/ASM: Test zeroing of vcvtps2ph 2026-06-28 13:14:26 +02:00
Ryan Houdek 4160a92621 Windows: Fixes duplicated hardware TSO handling
This was handled in both the Module.cpp files and also the common
TSOHandlerConfig on accident. Wouldn't have caused an issue but it was
definitely a bit weird.
2026-06-27 20:59:15 -07:00
Ryan Houdek 201bb73980 Windows/UnixLib: Adds support for Hardware TSO support
Including fallback to non-unixlib path because we need to support both.

Showcases how these are going to be implemented without throwing the
entire world at it right away. Next PR will be implementing the
remaining four necessary unixlib handlers that we will require:

- Kernel unaligned atomic control
- shm_stats thing
- madvise operation
- prctl vma naming
2026-06-27 19:59:35 -07:00
LC 78832cc0d0 Merge pull request #5612 from Sonicadvance1/174
Windows: Load unixlib if possible
2026-06-27 22:25:43 -04:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek 70fe9a4405 Merge pull request #5614 from lioncash/whoops
Vector: Fix typo in VPERMQOp
2026-06-26 21:45:33 -07:00
Ryan Houdek c09225f868 Windows: Load unixlib if possible
Currently does nothing other than load it (as the library also doesn't
do anything yet). Ensured it was working by temporarily creating a test
entrypoint and doing `Call` on to it.

Next step after this is to reimplement some of the nasty hacks FEX is
doing inside the unixlib code itself.
2026-06-26 20:50:24 -07:00
Ryan Houdek 126bcd365d winternl: Update enums
Newer WINE has a better mechanism for asking to load unix libraries.
Older WINE like what is in Proton doesn't have this yet. Add definitions
for both so we can try either one.
2026-06-26 20:50:19 -07:00
LC fe4d2bc6c5 Merge pull request #5611 from Sonicadvance1/173
Windows: Adds empty Linux side unix library
2026-06-26 23:48:51 -04:00
Ryan Houdek dbaf22372c Windows: Adds empty Linux side unix library
We are going to need a unix library. Going to take this one step at a
time without AI/ML so I fully understand all the pieces of the puzzle,
and to ensure we don't lose any functionality before we're ready.

This only ensures that we are building the Linux facing .so files for
arm64ec and wow64, but they are empty today. Next PR will be
initializing it on the PE side.
2026-06-26 19:05:08 -07:00
LC 43bd243457 Merge pull request #5609 from Sonicadvance1/170
OpcodeDispatcher: Fixes CRC32 with high 8-bit register
2026-06-26 15:30:00 -04:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
Ryan Houdek 26266c6a94 InstcountCI: Adds CRC32 with high 8-bit register 2026-06-26 11:38:08 -07:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Simon Scherer 0c1a35f297 unittests/ASM: Add unit test for crc32 with 8bit register operand 2026-06-26 11:10:06 -07:00
Tony Wasserka 9ac608ca43 Merge pull request #5606 from Sonicadvance1/169
LibraryForwarding/cuda: Convert constexpr to const
2026-06-26 11:17:49 +02:00
Ryan Houdek 7ae55d73c1 Merge pull request #5607 from lioncash/broadcast
AVX: Handle easily broadcastable permutations in VPERMQ
2026-06-25 20:23:55 -07:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek f8491794d3 thunks/cuda: Convert constexpr to const
Apparently some compilers or libstdc++ or libc++ takes offence to
constexpr std::array that gets filled by GOT. My compiler this generates
the same code regardless but I guess this'll probably fix #5582.
2026-06-25 16:53:59 -07:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
Ryan Houdek 3d66be9e5a Merge pull request #5604 from lioncash/pcmpstr
instcountci: Add 16-bit pcmpxstrx variants
2026-06-25 10:23:59 -07:00
Ryan Houdek ec1b24d05b Merge pull request #5603 from lioncash/permq
AVX: Handle transpose cases in VPERMQ
2026-06-25 01:18:01 -07:00
Ryan Houdek d93997c1cb Merge pull request #5602 from lioncash/same
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek e19aa975c8 Merge pull request #5597 from simon902/BTOpTypo
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00
Simon Scherer 3f02dd0a36 OpcodeDispatcher: Update BTOp comment to clarify AMD vs Intel flag behavior 2026-06-25 07:36:32 +02:00
Ryan Houdek ab4e0f653a Merge pull request #5600 from lioncash/palign
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek fda023e7dc Merge pull request #5601 from lioncash/pshufb
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00
Ryan Houdek ad618be979 Merge pull request #5599 from lioncash/unused
Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
2026-06-24 09:47:38 -07:00
Ryan Houdek 5881266256 Merge pull request #5598 from lioncash/ilpd
AVX: Wire up helper to VPERMILPD
2026-06-24 09:03:22 -07:00
Simon Scherer e03187852b OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved. 2026-06-24 16:11:13 +02:00
Ryan Houdek b846cb6c2c Merge pull request #5596 from lioncash/ilps
AVX: Wire up lane helper for VPERMILPS imm variant
2026-06-24 00:56:36 -07:00
Ryan Houdek 62ecdd650f Merge pull request #5595 from lioncash/pspd
AVX: Wire up lane helper for VSHUFPD/VSHUFPS
2026-06-24 00:22:12 -07:00
Ryan Houdek 9f195ff377 Merge pull request #5594 from lioncash/pshuf
AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
2026-06-23 20:31:04 -07:00
Ryan Houdek 27a5f09185 Merge pull request #5592 from ShadowCurse/fixes
JIT: Arm64: fix the loop in CacheLineClear/Clean
2026-06-23 17:15:07 -07:00
Ryan Houdek 01b0b4e653 Merge pull request #5593 from lioncash/dup
VectorOps: Avoid dup if able in VInsElement 128-bit element path
2026-06-23 17:10:35 -07:00
Egor Lazarchuk ff5dfff5bb IR: fix typo in CacheLineClean description 2026-06-24 00:38:59 +01:00
Egor Lazarchuk e3e9777ee6 JIT: Arm64: fix the loop in CacheLineClear/Clean
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
2026-06-24 00:38:48 +01:00
Ryan Houdek 7dc2dc8749 Merge pull request #5591 from lioncash/typo
OpcodeDispatcher: Fix typo in comment
2026-06-23 16:16:29 -07:00
Ryan Houdek 4eb5694872 Merge pull request #5590 from lioncash/insert
AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
2026-06-23 15:21:45 -07:00
Ryan Houdek 681c5e8097 Merge pull request #5589 from lioncash/selector
OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
2026-06-23 13:01:41 -07:00
Ryan Houdek 5ad06b0255 Merge pull request #5588 from lioncash/pinsr
AVX: Remove unnecessary moves from PINSRX ops
2026-06-23 12:34:03 -07:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC a4cd5f7584 instcountci: Add 16-bit pcmpxstrx variants
Also adds expanded mask variants. Lets us get a better whole picture on
all the main paths of these instructions.
2026-06-23 09:25:37 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek 1619374252 Merge pull request #5587 from lioncash/pd
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
Ryan Houdek c13064e201 Merge pull request #5586 from lioncash/mov
[SVE256] Remove unnecessary move in VCVTPS2PD
2026-06-22 20:44:23 -07:00
LC 20647f2287 AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
Ryan Houdek 37e32fbcb9 Merge pull request #5585 from lioncash/cmp-move
[SVE256] Remove heavy handed moves from scalar compares
2026-06-22 17:26:37 -07:00
Ryan Houdek 27acbba52e Merge pull request #5584 from lioncash/cmp_scalar
[SVE256] More comprehensively test SSE insertions for scalar comparisons
2026-06-22 15:49:57 -07:00
Ryan Houdek 6f33d2b4c4 Merge pull request #5583 from lioncash/scalar
[SVE256] Add more SSE scalar variant unit tests
2026-06-22 12:41:52 -07:00
LC a2e4209f4f AVX: Reduce moves in 256-bit VPSHUFB
Just a minor reduction by avoiding insertion overhead.
2026-06-22 10:47:49 -04:00
LC 73d6716828 AVX: Shave some moves off 256-bit VPALIGNR
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
2026-06-22 10:09:13 -04:00
LC 1fb2419be2 Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
No behavior change, just a reduction in noise.
2026-06-22 09:47:46 -04:00
LC 3394808c06 AVX: Wire up helper to VPERMILPD
We can just leverage the shuffle handler for this, since VPSHUFD
essentially functions like VPERMILPD
2026-06-22 09:03:22 -04:00
LC c252b58a15 AVX: Wire up lane helper for VPERMILPS imm variant
Makes for some more trivial savings. Will need handling for VPERMILPD
added separately, since selector behavior is different.
2026-06-22 05:21:10 -04:00
LC 39ae8c3ea0 AVX: Wire up lane helper for VSHUFPD/VSHUFPS
Also allows collapsing quite a bit of emitted code, like with
the shuffles in #5594
2026-06-22 04:46:22 -04:00
LC eb8c2d964c Vector: Factor out 128-bit path in SHUFOpImpl
We can leverage this for the 256-bit path
2026-06-22 03:47:20 -04:00
LC bd9cf9ca11 AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.

Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.

For example:

vpshufd ymm0, ymm1, 0b00000011

drops from 50 instructions to 9
2026-06-22 00:16:39 -04:00
Ryan Houdek 280568df2f Merge pull request #5581 from lioncash/fwd
Passes: Trim unnecessary forward declarations
2026-06-21 20:05:23 -07:00
Ryan Houdek b87ff1e2dc Merge pull request #5580 from lioncash/list
IntrusiveIRList: Amend signature for PostRA()
2026-06-21 20:04:51 -07:00
Ryan Houdek 55c90cfc38 Merge pull request #5579 from lioncash/typo
JIT: Amend op typos in implementations
2026-06-21 20:04:19 -07:00
Ryan Houdek 500d2374a5 Merge pull request #5578 from lioncash/tidy
Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
2026-06-21 20:03:47 -07:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC 9ab0920f01 OpcodeDispatcher: Fix typo in comment
It's the bits in general, not just the even ones (whoops).
2026-06-21 19:16:38 -04:00
LC a6e7fba433 AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.

In some cases, this can be quite drastic, like with:

vpblendw ymm0, ymm1, ymm2, 0b00000001

being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC 8905e39439 OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC df41b85827 OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC 9fa3b9345e AVX: Remove unnecessary moves from PINSRX ops
These are old paths still around from when StoreResult used to
automatically perform truncating moves.

These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
LC 42af6c8508 [SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC b7df1bc259 [SVE256] Remove unnecessary move in VCVTPS2PD
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC b40f9db735 [SVE256] Remove heavy handed moves from scalar compares
(See #3799)

I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)

This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC 938841c658 [SVE256] More comprehensively test SSE insertions for scalar comparisons
See #3799

Drops in the facilities to ensure all of the available SSE scalar comparison
paths are tested for proper insertion behavior.
2026-06-21 00:35:47 -04:00
LC 34344e5769 Merge pull request #5576 from Sonicadvance1/167
Proton: Fixes Mafia 3
2026-06-20 23:18:33 -04:00
Ryan Houdek 538a9624ec ArchHelpers/Arm64: Fixes zero register usage
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
2026-06-20 19:56:18 -07:00
Ryan Houdek c4c69ca8de ArchHelpers/Arm64: Support CAS/CASP in non-JIT SIGBUS handler
Proton was using this
2026-06-20 19:55:39 -07:00
LC 90330ab3e4 [SVE256] Add more SSE scalar variant unit tests
See #3799

These were technically already covered when the work was done to make
SSE insertion behavior conform to hardware, so this just adds tests
that ensure that behavior holds over time.
2026-06-20 19:25:19 -04:00
Ryan Houdek c5880e7618 Merge pull request #5572 from lioncash/str
[SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
2026-06-20 15:13:21 -07:00
Ryan Houdek 9d0c05d9cc Merge pull request #5575 from lioncash/movq2dq
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek f6a68cb7fd Merge pull request #5574 from lioncash/sse4a
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
Ryan Houdek 462c785418 Merge pull request #5573 from lioncash/cvtpi
[SVE256] Handle SSE insertions for CVTPI2PD
2026-06-20 12:44:21 -07:00
LC a1f90dd8d3 Passes: Trim unnecessary forward declarations
Less visual noise and lingering types left in the header.
2026-06-20 13:59:18 -04:00
LC 2a67261eac IntrusiveIRList: Amend signature for PostRA()
PostRA is a bool, not an unsigned value. We can also adjust SpillSlots()
to use uint32_t like its returned data member.
2026-06-20 13:39:35 -04:00
LC 1d3403fdc2 JIT: Amend op typos in implementations
Mostly benign, but ensures that they're correct in the event any of
their IR definitions change.
2026-06-20 13:16:12 -04:00
LC 53301b0f56 Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
Same thing, just a little less verbose.
2026-06-20 12:32:23 -04:00
LC 8989ce1766 [SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
See #3799
2026-06-19 18:53:50 -04:00
LC faa121e9ef Vector: Move MOVQ2DQ over to Bind
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC f48759e83a [SVE256] Handle SSE insertions for MOVQ2DQ
See #3799
2026-06-19 17:42:21 -04:00
LC 2eca733603 [SVE256] Handle SSE insertions for EXTRQ/INSERTQ
See #3799
2026-06-19 17:05:59 -04:00
LC 14b65cec43 [SVE256] Handle SSE insertions for CVTPI2PD
See #3799

CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
Ryan Houdek f5477039fa Merge pull request #5571 from simon902/16bitleave
Fix incorrect RSP update for 16bit leave
2026-06-19 10:27:33 -07:00
Simon Scherer 9fa3221687 FEXCore: Fix incorrect RSP update for 16bit leave 2026-06-19 14:23:16 +02:00
Simon Scherer c8c63faf15 unittests/ASM: Adds unit test for 16bit leave 2026-06-19 14:16:15 +02:00
Ryan Houdek ee4794c99e Merge pull request #5570 from lioncash/pmadd
[SVE256] Handle SSE insertions for PMADDWD
2026-06-17 22:21:16 -07:00
Ryan Houdek 3a23bb4b73 Merge pull request #5569 from lioncash/cmp
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00
Ryan Houdek 46ffb25f84 Merge pull request #5568 from lioncash/mov3
[SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
2026-06-17 21:21:50 -07:00
Ryan Houdek 069a4025c5 Merge pull request #5567 from lioncash/mov2
[SVE256] Handle SSE insertion for aligned/unaligned loads and non-temporal loads
2026-06-17 20:18:20 -07:00
LC edd044752d [SVE256] Handle SSE insertions for PMADDWD
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek 3929d25dcc Merge pull request #5566 from lioncash/mov
[SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
2026-06-17 18:55:22 -07:00
LC 99baa4f3d9 [SVE256] Handle SSE insertions for CMPSD/CMPSS
See #3799
2026-06-17 21:05:31 -04:00
LC f64d4c571b [SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
See #3799
2026-06-17 20:38:58 -04:00
LC 9372fa169a [SVE256] Handle SSE insertions for MOVSD/MOVSS
See #3799
2026-06-17 20:09:17 -04:00
Ryan Houdek f98ac7f268 Merge pull request #5565 from lioncash/xor
[SVE256] Handle SSE insertions for XOR special case
2026-06-17 16:47:50 -07:00
LC daaa6ec129 [SVE256] Handle SSE insertions for aligned and unaligned moves
See #3799
2026-06-17 19:39:06 -04:00
LC 88afc22d5b [SVE256] Handle SSE insertions for MOVNTDQA
See #3799
2026-06-17 19:38:57 -04:00
Ryan Houdek 23099100b6 Merge pull request #5564 from lioncash/misc2
[SVE256] Handle SSE insertions for INSERTPS, PSIGN, PINSR, and shuffles
2026-06-17 16:24:26 -07:00
LC b7ea9e30df [SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
See #3799

Gets a few of the moves out of the way.
2026-06-17 17:39:13 -04:00
LC 844c3bb197 [SVE256] Handle SSE insertions for XOR special case
See #3799

Ensures that our special case maintains insertion behavior
2026-06-17 16:33:05 -04:00
LC 208c6d3eac [SVE256] Handle SSE insertions for vector unary ops
See #3799
2026-06-17 14:43:56 -04:00
LC 886a2e74ac [SVE256] Handle SSE insertions for pack ops 2026-06-17 13:36:18 -04:00
Ryan Houdek adad3c27dd Merge pull request #5563 from lioncash/shift
[SVE256] Handle SSE insertions for shifts
2026-06-17 08:02:41 -07:00
Ryan Houdek 470aeab215 Merge pull request #5562 from lioncash/misc 2026-06-17 05:54:59 -07:00
LC db1d90ec9d [SVE256] Handle SSE insertions for shuffles
See #3799
2026-06-17 08:44:45 -04:00
LC 124ce8420a [SVE256] Handle SSE insertions for PINSR(B,D,Q,W)
See #3799
2026-06-17 08:22:26 -04:00
LC 4433eaf242 [SVE256] Handle SSE insertions for INSERTPS
See #3799
2026-06-17 08:11:49 -04:00
LC 5ba070f600 [SVE256] Handle SSE insertions for PSIGN(B,D,W)
See #3799
2026-06-17 07:59:33 -04:00
LC 2d1a42aa00 [SVE256] Handle SSE insertions for shifts
See #3799
2026-06-17 06:06:08 -04:00
LC 74d9f5a3e2 [SVE256] Handle SSE insertions for MOVDDUP
See #3799
2026-06-17 05:13:08 -04:00
LC 634fbb5a73 [SVE256] Handle SSE insertions for Float->Int/Int->Float conversions
See #3799
2026-06-17 04:44:39 -04:00
LC e16948bf80 [SVE256] Handle SSE insertions for CVTPD2PS/CVTPS2PD
See #3799
2026-06-17 03:51:36 -04:00
LC 76c9833ee5 [SVE256] Handle SSE insertions for CMPPD/CMPPS
See #3799
2026-06-17 03:33:26 -04:00
Ryan Houdek 99662b70ff Merge pull request #5560 from lioncash/psad
[SVE256] Handle SSE insertions for more misc ops
2026-06-17 00:14:30 -07:00
LC fcde9eabbf [SVE256] Handle SSE insertions for PALIGNR
See #3799
2026-06-17 02:13:37 -04:00
LC 4ef15951a8 [SVE256] Handle SSE insertions for PACKSS/PACKUS ops
See #3799
2026-06-17 02:08:24 -04:00
LC 1bc51c2290 [SVE256] Handle SSE insertions for PMULUDQ
See #3799
2026-06-17 02:00:11 -04:00
LC be1025901a [SVE256] Handle SSE insertions for ADDSUBPD/ADDSUBPS
See #3799
2026-06-17 01:54:55 -04:00
Ryan Houdek 1a606de29f Merge pull request #5559 from lioncash/phmin
[SVE256] Handle SSE insertions for PHMINPOSUW, DPPD, and DPPS
2026-06-16 22:50:53 -07:00
LC 454c0b31cb [SVE256] Handle SSE insertions for MPSADBW
See #3799
2026-06-17 01:47:21 -04:00
Ryan Houdek 223e0f4e53 Merge pull request #5558 from lioncash/blend
[SVE256] Handle SSE insertions for blends
2026-06-16 22:34:32 -07:00
LC 3a84091945 [SVE256] Handle SSE insertions for DPPD/DPPS
See #3799
2026-06-17 01:32:36 -04:00
LC a0e8f1097f [SVE256] Handle SSE insertions for PHMINPOSUW
See #3799
2026-06-17 01:21:19 -04:00
Ryan Houdek 08ed4fb983 Merge pull request #5546 from neobrain/feature_woa_fexofflinecompiler
CodeCache: Support targeting WOW64/ARM64EC in FEXOfflineCompiler
2026-06-16 22:15:49 -07:00
Ryan Houdek 0b1f336e03 Merge pull request #5557 from lioncash/round 2026-06-16 22:09:20 -07:00
LC 0258fcb116 [SVE256] Handle SSE insertions for blends
See #3799
2026-06-17 01:05:46 -04:00
LC 3272aa3f08 [SVE256] Handle SSE insertions for ROUNDPD/ROUNDPS
See #3799
2026-06-17 00:42:20 -04:00
Ryan Houdek d9ea6651f8 Merge pull request #5556 from lioncash/madd
[SVE256] Handle more SSE insertions for some one-off instructions
2026-06-16 21:22:29 -07:00
LC 32b11603d8 [SVE256] Handle SSE insertions for PMOVSX/PMOVZX ops 2026-06-17 00:04:03 -04:00
LC 538fd2672d [SVE256] Handle SSE insertions for PSADBW
See #3799
2026-06-17 00:04:03 -04:00
LC 3e37724e3e [SVE256] Handle SSE insertions for PHADDSW
See #379
2026-06-17 00:04:03 -04:00
LC 1b249ba76b [SVE256] Handle SSE insertion for PHSUBD/PHSUBW/PHSUBSW 2026-06-17 00:04:03 -04:00
LC df1295fbd0 [SVE256] Handle SSE insertions for HSUBPD/HSUBPS
See #3799
2026-06-17 00:04:03 -04:00
LC dcd71fe126 [SVE256] Handle SSE insertions for PMULHW/PMULHRSW
See #3799
2026-06-17 00:04:00 -04:00
LC 610ee5db76 [SVE256] Handle SSE insertions for PMADDUBSW
See #3799
2026-06-16 23:02:25 -04:00
Ryan Houdek 9d5494d9f0 Merge pull request #5555 from lioncash/alu
[SVE256] Vector: Handle SSE insertion properly for various ALU operations
2026-06-16 19:56:38 -07:00
Ryan Houdek dd44bc8d00 Merge pull request #5554 from lioncash/vmov
OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
2026-06-16 19:47:01 -07:00
Ryan Houdek 110c7cb62b Merge pull request #5553 from lioncash/bind
OpcodeDispatcher: Make use of Bind consistently
2026-06-16 19:44:44 -07:00
Ryan Houdek 09aa5abbdf Merge pull request #5552 from lioncash/literal
OpcodeDispatcher: Move a few stray literal accesses to Literal()
2026-06-16 19:37:28 -07:00
LC 1d8b6df630 [SVE256] Vector: Handle SSE insertion properly for various ALU operations
See #3799 for the bulk of the issue explanation. Ensures that emulated
SSE operation on aarch64 don't end up obliterating the upper 128-bit
lane when SVE-256 is present (Adv. SIMD operations zero-extend)

Knocks out quite a few SSE instructions right off the jump.
2026-06-16 22:26:28 -04:00
LC 195058752e OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
Will make removing the TODO in LoadSource regarding partial loads a
little easier.
2026-06-16 17:47:01 -04:00
LC c62805e86d OpcodeDispatcher: Make use of Bind consistently
We had a few places that were using Bind, and a few other places
that were using specializations as a means to composing the instruction
tables. Instead, we can just use Bind consistently, which lets us tidy
up a bunch of the implementations (and gets rid of some unnecessary
codegen).
2026-06-16 15:07:59 -04:00
LC d15b175c33 OpcodeDispatcher: Move a few stray literal accesses to Literal()
Same core behavior, but ensures that the immediates are valid literals
when assertions are enabled.
2026-06-16 11:50:48 -04:00
Tony Wasserka 6cb73adfd5 CodeCache: Integrate FEXOfflineCompiler backend for WoA 2026-06-16 17:33:16 +02:00
Billy Laws 6d1cd67900 Windows: Implement WOW64/ARM64EC offline compiler backend
Implements an offline JIT compiler backend that compiles x86 code blocks
from a given code map into an ARM64 code cache on Windows. Separate
binaries are built for WOW64 (32-bit) and ARM64EC (64-bit) targets.

NOTE: This patch originally added a separate binary; instead the new
      functionality is added to the existing FEXOfflineCompiler and will
      be properly integrated in the next patches.
2026-06-16 17:30:32 +02:00
Ryan Houdek e12bd27106 Merge pull request #5536 from Sonicadvance1/105
LibraryForwarding: Implement support for CUDA
2026-06-12 15:14:38 -07:00
Ryan Houdek cb018257cf Merge pull request #5547 from neobrain/feature_fexofflinecompiler_process_all
FEXOfflineCompiler: Add "process-all" verb
2026-06-12 02:11:33 -07:00
LC e02953dc17 Merge pull request #5548 from Sonicadvance1/166
Frontend: Fix vsyscall page tracking.
2026-06-03 21:20:20 -04:00
Ryan Houdek ba56f8e0c5 unittests/FEXLinuxTests: Adds 64-bit vsyscall test
We never actually had a unittest to ensure these keep working, so add
one now.
2026-06-03 18:03:31 -07:00
Ryan Houdek ac12dd55c3 Frontend: Fix vsyscall page tracking.
Now that NX is tracked in the frontend, we need to ensure that adjusted
RIP pages are tracked correctly. Keep around both instruction stream
pointers, validate the the original RIP is executable, and read from the
adjusted RIP as appropriate.

Fixes #5544
2026-06-03 18:03:31 -07:00
Ryan Houdek 24647820d7 Linux: Ensure 64-bit vsyscall page is tracked
It's a purely virtual page even on x86, so we need to manually add it to
tracking.
2026-06-03 18:03:31 -07:00
Ryan Houdek 7e8aa711ef Linux: Pass gettimeofday through glibc
This ensures that it hits the VDSO path if possible.
2026-06-03 17:45:41 -07:00
Tony Wasserka 329e561eff FEXOfflineCompiler: Add "process-all" verb
This operation will take care of any pending code cache operations:
* import new code maps from `CACHE_DIR/codemap/new` and process them to `CACHE_DIR/codemap/ready`
* generate caches for updated code maps with new blocks
* ensure caches already exist for all other code maps (and generate them if needed)
2026-06-03 18:42:37 +02:00
LC d848cbbc0f Merge pull request #5543 from Sonicadvance1/165
FEXCore: Ensure LOCK prefix instructions are handled correctly
2026-06-02 23:39:02 -04:00
Ryan Houdek cd46e43c20 unittests/FEXLinuxTests: Add a LOCK prefix test
Ensures we handle lock prefixing correctly.
2026-06-02 19:41:04 -07:00
Ryan Houdek 98674c1cc8 unittests/FEXLinuxTests: Support redirecting RIP entirely 2026-06-02 19:41:04 -07:00
Ryan Houdek 8bfae631b1 Frontend: Support raising unimplemented instruction on LOCK failure
When the LOCK prefix is on an instruction that doesn't support LOCK then
it raises a SIGILL. Make sure to pass that up.

Additionally if the instruction does support lock prefix, has a lock
prefix, but the destination is not memory then that is also invalid.
2026-06-02 19:41:03 -07:00
Ryan Houdek d00c7cf3a3 Frontend: Support passing the decode failure type through decoding
Only used for invalid inst currently
2026-06-02 19:41:03 -07:00
Ryan Houdek b45665fee4 OpcodeDispatcher: Support instruction type of unimplement operation 2026-06-02 19:41:03 -07:00
Ryan Houdek 1b58664541 X86Tables: Describe instructions that support LOCK prefix 2026-06-02 19:41:02 -07:00
LC ed6a178ae3 Merge pull request #5542 from Sonicadvance1/164
OpcodeDispatcher: Fixes 64-bit LODs with address size override
2026-06-02 22:40:53 -04:00
Ryan Houdek e925ca509d unittests/ASM: Adds lods tests with address size override
If the override isn't handled then it'll fall back to either using the
wrong address and getting the wrong data or crashing depending on how it
is broken.
2026-06-02 18:21:00 -07:00
Ryan Houdek ca310cf815 OpcodeDispatcher: Fixes 64-bit LODs with address size override
Fairly trivial but just need to be careful with address size wraparound
as usual.
2026-06-02 18:21:00 -07:00
Ryan Houdek 6fa27aac42 Thunks: Implement support for CUDA
This is enough to get less complex cuda applications running, and is a
good starting spot to slowly finish off the remaining implementation.

Some information:
- 429 functions in total
- 153 only compiled for 64-bit (35.6%)
- 7 functions disabled entirely (1.6%)

The main thing /not/ working with this initial implementation is .cu
files compiled in to an ELF using the static cuda runtime. This is due
to the `cuGetExportTable` function being stubbed out and the static cuda
RT requires at least two interfaces from that function before it
continues.

This function isn't publicly documented by NVIDIA but has been publicly
reverse engineered to be fairly trivial. It's just a jump table with the
first element being the size of the table in bytes.

That will be the next step of the implementation.
2026-06-01 16:32:45 -07:00
Tony Wasserka a5c3fc4751 Merge pull request #5514 from peppergrayxyz/snd_htimestamp_t
LibraryForwarding: Add annotation for snd_htimestamp_t
2026-06-01 12:53:42 +02:00
LC 65b05fa8c1 Merge pull request #5540 from Sonicadvance1/163
arm64ec: Single instruction optimization in EC map lookup
2026-06-01 05:41:34 -04:00
Pepper Gray 3ee556d858 add template for snd_htimestamp_t
building on musl fails with:
`error: Unsupported parameter type 'snd_htimestamp_t *' (aka 'timespec *')`

due to empty padding members in `alltypes.h`:
```
STRUCT timespec {
  time_t tv_sec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER==4321);
  long tv_nsec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER!=4321);
};
```
add (missing) annotation to libasound interface

fix: #5513
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-31 17:38:25 +02:00
Ryan Houdek fd3546a999 arm64ec: Single instruction optimization in EC map lookup
We can merge the lsr+and in to a single lsr by 18 and then use ldr with
LSL of 3 to accomplish the same result. Modern Cortex doesn't even
generate an additional integer pipeline uop for this ldr+lsl instruction
anymore.
2026-05-29 19:17:13 -07:00
Ryan Houdek f5fafa5b96 Merge pull request #5492 from FrontMage/codex/int29-failfast-probe
Windows: Trace interrupt translation and prototype INT 0x29 fail-fast mapping
2026-05-29 16:59:16 -07:00
Ryan Houdek 7ff2069e60 Merge pull request #5539 from Sonicadvance1/162
arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
2026-05-29 16:49:14 -07:00
Ryan Houdek 154ff43d7f arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
We were accidentally allocating some things without this flag and it was
causing us to dump memory in to the lower 32-bits on arm64ec.

This was causing the game
[Below](https://store.steampowered.com/app/250680/BELOW/) to run out of
memory to allocate for its LUA JIT and causes it to crash.
2026-05-29 16:17:04 -07:00
Ryan Houdek b754fe4810 External/rpmalloc: update 2026-05-29 16:16:46 -07:00
Tony Wasserka a1071ec01a Merge pull request #5538 from peppergrayxyz/syscall_headers
LinuxSyscalls: add missing thread header
2026-05-29 14:35:43 +02:00
Pepper Gray 92dce9a2ea add missing header <thread>
building using musl fails due to missing defintions:

```
Source/Tools/LinuxEmulation/LinuxSyscalls/Syscalls.cpp:913:23: error: no member named 'sleep_for' in namespace 'std::this_thread'
  913 |     std::this_thread::sleep_for(std::chrono::milliseconds {10});
      |                       ^~~~~~~~~
1 error generated.
```

Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-29 13:19:24 +02:00
LC 5fd917ec2f Merge pull request #5537 from Sonicadvance1/161
Format: Fix missed clang-format
2026-05-29 00:52:34 -04:00
Ryan Houdek d7cda23b25 Format: Fix missed clang-format
Minor clang-format version differences causing differing behaviour
again.
2026-05-28 12:12:24 -07:00
Ryan Houdek cae5da5777 Merge pull request #5530 from neobrain/fix_ccache_time_macros
Build: Enable ccache sloppiness for time macros
2026-05-25 10:55:27 -07:00
Ryan Houdek 97f1f47fa5 Merge pull request #5531 from neobrain/refactor_drop_cmake_settings
Drop unused CMakeSettings.json
2026-05-25 10:53:44 -07:00
Tony Wasserka 53c269ee25 Merge pull request #5516 from peppergrayxyz/thunk_rootfs
set sysroot to X86_DEV_ROOTFS for guest toolchain
2026-05-25 18:08:45 +02:00
Pepper Gray 1420d3cc10 set sysroot to X86_DEV_ROOTFS for guest toolchain
**Faulty Behaviour:**
When not building on Ubuntu `unittests/ThunkLibs` fails to find
c++ header and fails:

```
gen_input.cpp:2:10: fatal error: 'cstddef' file not found
    2 | #include <cstddef>
      |          ^~~~~~~~~
1 error generated.
```

This also includes building with nix-shell:
```
nix-shell ../Data/nix/LibraryForwarding/shell.nix --run "cmake --build . --target thunkgen_tests"
```

**Root Cause**:
`X86_DEV_ROOTFS` is not used, but  header paths are hard
coded to Ubuntu's multilib layout (`/usr/i686-linux-gnu/include/`,
`/usr/x86_64-linux-gnu/include/`). It works on Ubuntu but the
mechanism to use to another directory is broken and the include
paths always point to the host.

**Solution**:
set `--sysroot ${X86_DEV_ROOTFS}` for guest builds and remove
hard coded include paths.

Fix: #5515
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-25 17:18:47 +02:00
Tony Wasserka ce401b5ca1 Build: Enable ccache sloppiness for time macros
Ccache won't attempt to cache files that use __DATE__/__TIME__ since their
contents would always be out of date. Setting sloppiness disables this
behavior, which works for us since we don't use __DATE__/__TIME__ for anything
that needs accurate values.
2026-05-25 12:03:33 +02:00
Tony Wasserka cd412bd0f5 Drop unused CMakeSettings.json
Visual Studio (but not VS Code) used to read this file, but this convention is
discouraged nowadays.
2026-05-25 11:27:29 +02:00
FrontMage f33f88072e Windows: Fix INT exception formatting 2026-05-24 11:42:51 +08:00
LC 1240a00fa5 Merge pull request #5522 from Sonicadvance1/160
New CPL0 instructions from #5510 but with unittests
2026-05-23 22:57:53 -04:00
LC 83f325de0d Merge pull request #5519 from Sonicadvance1/157
meta: Add CONTRIBUTING.md
2026-05-23 22:55:26 -04:00
LC 0b871bf54e Merge pull request #5518 from Sonicadvance1/156
code-format-helper: More dependabot changes
2026-05-23 22:54:36 -04:00
LC 07f7aa3c8f Merge pull request #5521 from Sonicadvance1/159
HostFeatures: Don't capture CTR/MIDR under simulator
2026-05-22 22:10:38 -04:00
Ryan Houdek 7208bc6cdd unittests: Extend unittests for CPL0 instructions 2026-05-22 15:33:42 -07:00
Ryan Houdek fef5a98602 Fix build failure. 2026-05-22 15:33:41 -07:00
Daniel Lu 5cce65cdfa OpcodeDispatcher: Decode INVD and WBINVD through privileged op handling 2026-05-22 15:25:22 -07:00
Ryan Houdek 268081e5d0 Merge pull request #5520 from Sonicadvance1/158
Cherry-pick #5508 with instcountci changes
2026-05-22 14:46:20 -07:00
Ryan Houdek f5f179117e HostFeatures: Don't capture CTR/MIDR under simulator
If the simulator was selected, we would still capture CTR and MIDR on
the host AArch64 system. Potentially modifying codegen in unexpected
ways.

Ensure we return 0/0 like under x86 with simulator to simulate
"unknown".
2026-05-22 14:04:52 -07:00
Ryan Houdek 8c85096f98 meta: Add CONTRIBUTING.md 2026-05-22 14:03:14 -07:00
Ryan Houdek c5e7675c4b InstcountCI: Update 2026-05-22 13:58:10 -07:00
Daniel Lu 03009912ac JIT: Avoid clobbering guest rdx while raising generated faults 2026-05-22 13:56:40 -07:00
Ryan Houdek e4a1138291 code-format-helper: More dependabot changes 2026-05-22 13:44:22 -07:00
FrontMage a5bc54d2d5 Windows: Preserve INT 0x2D exception parameter source 2026-05-22 12:56:44 +08:00
LC df73e84725 Merge pull request #5507 from Sonicadvance1/155
CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588
2026-05-21 22:39:14 -04:00
Ryan Houdek 5bf07c2e77 CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588 2026-05-21 18:04:29 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
LC c98cef0da1 Merge pull request #5506 from Sonicadvance1/154
HostFeatures: Only enable `dc zva` optimization on Ampere CPUs
2026-05-19 22:29:11 -04:00
Ryan Houdek 1d9c52be02 InstcountCI: Update 2026-05-19 17:38:10 -07:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00
Ryan Houdek f66368b191 Merge pull request #5505 from fixedcat/main
Arm64: Fix byte-size handling in unaligned STLXR emulation
2026-05-19 12:28:23 -07:00
LC e4daea406e Merge pull request #5503 from Sonicadvance1/154
FEXCore: Allow InterruptFaultPage to be significantly further away
2026-05-19 12:22:18 -04:00
fixedcat 6216f22cb9 Arm64: Fix byte-size handling in unaligned STLXR emulation 2026-05-19 18:53:44 +08:00
LC d4c80d9094 Merge pull request #5422 from Sonicadvance1/138
FEXCore: Add support for developer single stepping, read/write watching.
2026-05-19 01:11:04 -04:00
Ryan Houdek 27324ded87 Merge pull request #5499 from bylaws/winstuff
Windows additions for code caching
2026-05-18 16:05:59 -07:00
Ryan Houdek 7d1c625e32 FEXCore: Allow InterruptFaultPage to be significantly further away
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.

Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
2026-05-18 15:57:25 -07:00
Ryan Houdek b4fe65f2c0 Merge pull request #5496 from neobrain/fix_libfwd_findpkg
Library Forwarding: Various build system improvements
2026-05-18 14:55:53 -07:00
Ryan Houdek f5efdac2e5 Merge pull request #5423 from bylaws/depenencey
WOW64: Support disabling DEP
2026-05-18 14:39:44 -07:00
Billy Laws 23de875516 Windows: Add NtUnmapViewOfSection prototype 2026-05-17 23:06:02 +01:00
Billy Laws 82030b8286 Windows: Declare winternl relocation APIs 2026-05-17 23:03:11 +01:00
Ryan Houdek af4da43bb8 Merge pull request #5494 from neobrain/fix_determine_va
Allocator: Fix and optimize VA range detection
2026-05-17 14:51:59 -07:00
Ryan Houdek 33f3b8659c Merge pull request #5495 from neobrain/fix_portable_config
Config: Use more sensible default for portable config location
2026-05-17 14:51:04 -07:00
Ryan Houdek ed724a61a7 Merge pull request #5498 from bylaws/evmd
ImageTracker: Support using image IDs as an extended volatile metadata key
2026-05-17 14:49:16 -07:00
Billy Laws 053bd74aa3 Windows/Common: Add ScopedHandle::reset() 2026-05-17 22:35:06 +01:00
Billy Laws 3c0410e59f ImageTracker: Support using image IDs as an extended volatile metadata key 2026-05-17 19:34:41 +01:00
Tony Wasserka e621f6c753 LibraryForwarding/Build: Explicitly look up LLVM headers
This could previously set up incorrect header paths when clang and LLVM were
installed in different directories (such as when using nix).
2026-05-15 16:07:13 +02:00
Tony Wasserka b05f000f42 LibraryForwarding/Build: Try harder to properly discover header locations 2026-05-15 16:07:13 +02:00
Tony Wasserka e60bfc6d23 LibraryForwarding/Build: Allow specifying system header location externally 2026-05-15 16:07:09 +02:00
Tony Wasserka 9a3d3201f9 Config: Use more sensible default for portable config location 2026-05-15 15:59:22 +02:00
Tony Wasserka abf9724424 Allocator: Fix and optimize VA range detection 2026-05-15 15:47:38 +02:00
LC ab9a8c62ab Merge pull request #5493 from Sonicadvance1/153
FEXGetConfig: Even more correctness changes for X2E
2026-05-14 09:39:31 -04:00
FrontMage e2fe936152 Windows: Handle INT 0x29 as fast-fail 2026-05-14 12:05:08 +08:00
Ryan Houdek 1d71650379 FEXGetConfig: Even more correctness changes for X2E
Some of the information was incorrect, so make sure it shows the
hardware correctly.
2026-05-13 20:00:57 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Tony Wasserka 52ad434d24 LinuxSyscalls: Defer MappedResource deletion until after code invalidation
This ensures that any code buffer memory owned by the MappedResource is
invalidated before being deallocated.
2026-05-13 21:25:47 +02:00
Tony Wasserka d69d111bb6 CodeCache: Align code section within cache files
This allows mapping the code directly into memory for execution.
2026-05-13 21:25:47 +02:00
LC 50f4494875 Merge pull request #5490 from Sonicadvance1/152
FEXGetConfig: Showcase RMW versus loadstore atomic differences
2026-05-12 16:32:04 -04:00
Ryan Houdek 8a4982383a FEXGetConfig: Showcase RMW versus loadstore atomic differences
This differs on X2E, so it's good to showcase it.
2026-05-12 12:38:44 -07:00
Ryan Houdek 0d72890482 Merge pull request #5432 from pmatos/f64-fprem
JIT-inline FPREM/FPREM1 for reduced precision x87 path
2026-05-11 14:31:31 -07:00
Ryan Houdek 9be7d6d112 Merge pull request #5489 from neobrain/feature_better_fexbash
FEXBash: Drop implicit -c and add colored PS1
2026-05-11 12:22:00 -07:00
Tony Wasserka 3c4121ba07 FEXBash: Use a shiny rainbow for PS1 2026-05-11 20:30:21 +02:00
Tony Wasserka b6e44b04d6 FEXBash: Clean up path handling 2026-05-11 20:30:21 +02:00
Tony Wasserka 02c11afbc6 FEXBash: Don't imply "-c" to behave more closely like bash
Implicitly adding "-c" breaks argument passing for scripts. For example, the
command "FEXBash ./steam.sh -silent" will process steam.sh but the script
wouldn't see the "-silent" argument previously.
2026-05-11 19:49:15 +02:00
Paulo Matos 2b8f5b57eb instcountci: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 0519c9467c asm_tests: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00
Ryan Houdek 7506cb4ea1 Merge pull request #5483 from neobrain/fix_guest_wrapper_code_cache
CodeCache: Delay cache loading for guest library wrappers until after LoadLib
2026-05-06 18:04:55 -07:00
Tony Wasserka 60bc5944db CodeCache: Slightly optimize cache file writing
ftruncate only requires one call (and one extra seek) instead up to 64 manual
zero writes.
2026-05-06 17:08:15 +02:00
Tony Wasserka 8f0572283a Windows/CRT: Implement ftruncate and _chsize 2026-05-06 17:08:04 +02:00
Tony Wasserka b13b46eefe CodeCache: Delay cache loading for guest library wrappers until after LoadLib
These libraries need to be initialized before relocating their caches,
since the guest function hashes won't be registered before.
2026-05-05 16:29:39 +02:00
LC 4db2a98d7f Merge pull request #5481 from Sonicadvance1/148
win32: Query DCZID_EL0 so clzero works
2026-05-05 08:42:31 -04:00
LC 05ebb07753 Merge pull request #5482 from Sonicadvance1/149
OpcodeDispatcher: Optimize MMX pshufw
2026-05-05 08:41:22 -04:00
Ryan Houdek 694e68b838 Merge pull request #5468 from peppergrayxyz/proc_self_stat
read /proc/self/stat using %lu
2026-05-04 20:21:16 -07:00
Pepper Gray 78320e1433 read /proc/self/stat using %lu
building on clang/musl causes warnings:

```
FEX/Source/Tools/FEXInterpreter/ELFCodeLoader.h:782:29: warning: format specifies type 'unsigned long long *' but the argument has type 'uint64_t *' (aka 'unsigned long *') [-Wformat]
  776 |                             "%llu %llu %llu %*u %*u "   // 26 to 30
      |                              ~~~~
      |                              %lu
  777 |                             "%*u %*u %*u %*u %*u "      // 31 to 35
  778 |                             "%*u %*u %*d %*d %*u "      // 36 to 40
  779 |                             "%*u %*u %*u %*d %llu "     // 40 to 45
  780 |                             "%llu %llu %llu %llu %llu " // 46 to 50
  781 |                             "%llu",                     // 51
  782 |                             &map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
      |                             ^~~~~~~~~~~~~~~
```

according to the [man page](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html)
`/proc/self/stat` uses `%lu`:

read the values as unsigned long (%lu) and then write them to
prctl_mm_map (platform specific format).

fixes: #5467
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-05 04:24:31 +02:00
Ryan Houdek 082e7b2695 InstcountCI: Update 2026-05-04 17:52:54 -07:00
Ryan Houdek cdbffb80c7 unittests: Adds a full coverage pshufw test 2026-05-04 17:52:53 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 47e173e549 Merge pull request #5472 from peppergrayxyz/format
use portable format specifiers
2026-05-04 14:18:28 -07:00
Ryan Houdek 162bd4be97 win32: Query DCZID_EL0 so clzero works
EL0 registers are readable without going through the registry, but we
were failing to populate this register, which was causing clzero to not
be supported.
2026-05-04 12:42:18 -07:00
Ryan Houdek f0764aeafe Merge pull request #5480 from neobrain/refactor_musl_sigmask
SignalDelegator: Simplify support for musl's sigset_t
2026-05-04 10:39:42 -07:00
Ryan Houdek 1efed71696 Merge pull request #5479 from neobrain/refactor_drop_compile_service
FEXCore: Drop unused CompileService
2026-05-04 10:36:44 -07:00
Tony Wasserka 06d77c1c19 SignalDelegator: Simplify support for musl's sigset_t 2026-05-04 16:28:47 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray abae5dd93b use portable format specifiers
building with clang/musl causes these warnings:

```
FEX/Source/Tools/FEXServer/ProcessPipe.cpp:100:96: warning: format specifies type 'ssize_t' (aka 'long') but the argument has type 'rlim_t' (aka 'unsigned long long') [-Wformat]
```

- cast platform specific MaxFDs members to uintmax_t and print as PRIuMAX
- use %zu for GetNumFilesOpen (size_t)

fix: #5471
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-04 11:20:55 +02:00
Ryan Houdek 93015a0266 Merge pull request #5458 from peppergrayxyz/largefile64
make 64bit symbols visible to enhance portability (musl)
2026-05-03 23:36:59 -07:00
Ryan Houdek d238db69d3 Merge pull request #5477 from peppergrayxyz/unistd
include missing header unistd.h
2026-05-03 15:18:41 -07:00
Ryan Houdek e0ead236b6 Merge pull request #5473 from peppergrayxyz/ObjectCacheRefCounter
remove dead code (ObjectCacheRefCounter)
2026-05-03 15:16:57 -07:00
Pepper Gray c548262664 include missing header unistd.h
build on clang/musl fails with:

```
FEX/unittests/APITests/Allocator.cpp:17:5: error: use of undeclared identifier 'close'
FEX/unittests/APITests/Allocator.cpp:23:5: error: use of undeclared identifier 'lseek'; did you mean 'fseek'?
FEX/unittests/APITests/Allocator.cpp:23:11: error: cannot initialize a parameter of type 'FILE *' (aka 'struct _IO_FILE *') with an lvalue of type 'int'
FEX/unittests/APITests/Allocator.cpp:24:5: error: use of undeclared identifier 'write'; did you mean '_IO_cookie_io_functions_t::write'?
FEX/unittests/APITests/Allocator.cpp:24:5: error: invalid use of non-static data member 'write'
```

include header to provide defintions

fix: #5476
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:49:40 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Ryan Houdek 85773995e1 Merge pull request #5470 from peppergrayxyz/header_redirect
fix include redirect for <poll.h> and <signal.h>
2026-05-03 04:41:18 -07:00
Pepper Gray 00b4777290 fix include redirect for <poll.h> and <signal.h>
building on musl/clang causes redirecting incorrect #includes warnings:

```
warning: redirecting incorrect #include <sys/poll.h> to <poll.h> [-W#warnings]
warning: redirecting incorrect #include <sys/signal.h> to <signal.h> [-W#warnings]
```

include headers instead of sys/headers.

fix: #5469
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 13:22:32 +02:00
Ryan Houdek 197e6de194 Merge pull request #5462 from peppergrayxyz/uc_sigmask
determine sigset_t fieldname to enhance portability (musl)
2026-05-03 03:48:53 -07:00
Ryan Houdek 9908ea4c2f Merge pull request #5466 from peppergrayxyz/libgen_h
include <libgen.h> for basename
2026-05-03 03:46:43 -07:00
Ryan Houdek c402b15bd3 Merge pull request #5464 from peppergrayxyz/tgkill
add header and classpath for tgkill
2026-05-03 03:46:07 -07:00
Ryan Houdek a3f3118ac5 Merge pull request #5460 from peppergrayxyz/sigset_t
use <signal.h> instead of glibc header to enhance portability (musl)
2026-05-03 03:18:20 -07:00
Pepper Gray b1aab0e498 determine sigset_t fieldname to enhance portability (musl)
musl build fails due to access to internal glibc member:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
  655 |       .SigMask = _context->uc_sigmask.__val[0],
      |                  ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```

add check to determine private glibc or musl member name or throw an error.

fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:17:09 +02:00
Pepper Gray 927c0ce54a include <libgen.h> for basename
building on clang/musl build fails due to missing symbol:

```
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:321:35: error: use of undeclared identifier 'basename'
  321 |   auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
      |                                   ^~~~~~~~
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:327:43: error: use of undeclared identifier 'basename'
  327 |     fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
      |                                           ^~~~~~~~
```

include missing header.

fixes: #5465
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:10:07 +02:00
Pepper Gray 57d9dc037d add header and classpath for tgkill
building on clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
 1174 |     tgkill(::getpid(), ::getpid(), SIGKILL);
      |     ^~~~~~
1 error generated.
```

include and use `FHU::Syscalls::tgkill`.

fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:50:23 +02:00
Pepper Gray e3cfe28848 use <signal.h> instead of glibc header to enhance portability (musl)
building using musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/ThreadManager.h:33:10: fatal error: 'bits/types/sigset_t.h' file not found
   33 | #include <bits/types/sigset_t.h>
      |          ^~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
```

`<bits/types/sigset_t.h>` is a glibc internal header, use <signal.h> instead.

fixes: #5459
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:03:38 +02:00
Pepper Gray ac91f583b8 make 64bit symbols visible to enhance portability (musl)
building using clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/x32/Types.h:548:5: error: member access into incomplete type 'const struct statfs64'
  548 |     COPY(f_bsize);
      |     ^
```

add `_LARGEFILE64_SOURCE` to define large-file feature macros

fix: #5457
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 10:52:09 +02:00
Ryan Houdek 215658bf29 Merge pull request #5455 from peppergrayxyz/sys_prctl
use <sys/prctl.h> to enhance portability (clang)
2026-05-02 15:11:33 -07:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek e91bda7765 Merge pull request #5447 from bylaws/claudefix5
SoftFloat: Fix FSCALE(0, +Inf) to raise IE and return a quiet NaN
2026-04-30 14:41:24 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Ryan Houdek b1381fd3b7 Merge pull request #5444 from bylaws/claudefix2
VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
2026-04-30 14:39:56 -07:00
Ryan Houdek e5f6a7d85e Merge pull request #5450 from neobrain/fix_code_cache_portable
CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
2026-04-30 14:17:16 -07:00
Ryan Houdek feae76fc4f Merge pull request #5451 from Sonicadvance1/146
arm64ec: Fixes crash in many games with SDL+Dualsense
2026-04-30 14:16:06 -07:00
Ryan Houdek 015f3cffb9 arm64ec: Fixes crash in many games with SDL+Dualsense
We were pointing to an incorrect function pointer and exploding when a
pending suspend doorbell had occured.
2026-04-30 12:59:53 -07:00
Tony Wasserka 3e5c17ae80 CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
In portable mode, FEXOfflineCompiler may not be in PATH (and if it is, it's
most likely not a compatible version). Instead, use the executable next to
the FEXServer binary.
2026-04-30 17:14:22 +02:00
Billy Laws 9a1d06c6ab WOW64: Support disabling DEP
Required for older 32-bit games that assumes the execute bit is implicit
from read.
2026-04-29 03:17:03 +00:00
Billy Laws 3e278b42f8 InstcountCI: Update 2026-04-29 02:54:34 +00:00
Billy Laws fa953445e9 InstcountCI: Update 2026-04-29 02:53:07 +00:00
Billy Laws a412b1d3b7 InstcountCI: Update 2026-04-29 02:43:19 +00:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c260b45e1 unittests/ASM: Adds tests for FXTRACT Inf/NaN 2026-04-29 02:29:23 +00:00
Ryan Houdek 098c4c57b4 Merge pull request #5443 from bylaws/claudefix1
X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel
2026-04-28 19:25:15 -07:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Billy Laws 1bd2ff3fc3 unittests/ASM: Adds test for FSCALE(0, +Inf) raising IE 2026-04-29 02:16:57 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Billy Laws cf20647b25 unittests/ASM: Adds test for 16-bit FIST with denormal input not setting IE 2026-04-29 02:09:54 +00:00
Billy Laws 8d7071e549 VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
2026-04-29 02:00:58 +00:00
Billy Laws fb006b2c6d unittests/ASM: Adds test for MAXPS/MAXPD NaN and signed-zero tie 2026-04-29 01:57:41 +00:00
Billy Laws 9039eeb3cd X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel 2026-04-29 01:57:18 +00:00
Billy Laws dd0702d30f unittests/ASM: Mark SSE4a/CLZERO as required for tests using them 2026-04-29 01:56:35 +00:00
LC 886faf0bd4 Merge pull request #5442 from neobrain/refactor_code_cache_check
CodeCache: Move bounds check to FEXOfflineCompiler
2026-04-28 19:47:52 -04:00
Tony Wasserka 86e28c6d34 CodeCache: Move bounds check to FEXOfflineCompiler
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.

It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
2026-04-28 17:31:20 +02:00
LC 821efab8aa Merge pull request #5440 from Sonicadvance1/145
OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
2026-04-28 07:55:10 -04:00
Ryan Houdek 5295365dd0 InstcountCI: Update 2026-04-27 17:55:49 -07:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
LC dd145aaa88 Merge pull request #5439 from Sonicadvance1/144
Fix push/pop fs/gs segments and unittests
2026-04-27 19:43:44 -04:00
Ryan Houdek 34b3adc23d unittests/ASM: Adds unit test to ensure push/pop segment of o16 works
Only ensures we are pushing and popping the correct size, not any of the
selector data within it, as 64-bit systems with the FSGSBase extension
don't use them selectors anyway.

Can't test the 32-bit side currently because we would corrupt FS/GS in
CI and the host testharnessrunner can't fix that right now.
2026-04-27 15:21:47 -07:00
Simon Scherer ab14882761 FEXCore: Fix 2byte stack access for 0x66 PUSH/POP FS/GS 2026-04-27 15:21:11 -07:00
LC 7dc1f54fb6 Merge pull request #5435 from Sonicadvance1/143
unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting
2026-04-25 10:34:07 -04:00
Ryan Houdek c09fb03eda unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting 2026-04-25 00:46:40 -07:00
Ryan Houdek dbf2761fb7 InstcountCI: Update 2026-04-25 00:45:05 -07:00
Ryan Houdek fd1378f778 InstcountCI: Fix incorrect instruction 2026-04-25 00:43:50 -07:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
LC 4b02c04afc Merge pull request #5429 from Sonicadvance1/142
Steam/CompatTool: Fixes Graphics Provider path handling
2026-04-23 20:13:13 -04:00
Ryan Houdek deed99e7a3 Merge pull request #5425 from pmatos/f64-atan-fyl2x
JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path
2026-04-23 15:16:04 -07:00
Ryan Houdek 7bffc4a177 Steam/CompatTool: Fixes Graphics Provider path handling
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.

Misunderstood the assignment before.
2026-04-23 15:13:42 -07:00
Ryan Houdek 701555e400 Merge pull request #5428 from Sonicadvance1/141
Steam/CompatTool: Support `STEAM_COMPAT_GRAPHICS_PROVIDER` for rootfs path
2026-04-21 12:50:22 -07:00
Ryan Houdek 49fa86d0b5 Steam/CompatTool: Support STEAM_COMPAT_GRAPHICS_PROVIDER for rootfs path
If we have been provided a graphics provider path through an environment
variable, then use that path directly rather than searching.
2026-04-21 12:37:40 -07:00
Paulo Matos adbace8810 instcountci: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos 050138bcea asm_tests: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
LC 59755ec115 Merge pull request #5426 from Sonicadvance1/139
Snapdragon X2 Elite fixes
2026-04-20 11:53:51 -04:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Ryan Houdek d41d52b889 Merge pull request #5419 from pmatos/f64-scale-f2xm1
JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path
2026-04-17 14:38:22 -07:00
Paulo Matos 18f69fb16d instcountci: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-17 17:28:13 +02:00
Ryan Houdek 739e85032b FEXCore: Add support for developer single stepping, read/write watching. 2026-04-16 14:13:19 -07:00
Paulo Matos d165711f2e asm_tests: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:24 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 441116e1e6 Merge pull request #5421 from Sonicadvance1/136
Scripts: Move arch check first in InstallFEX
2026-04-15 16:24:04 -04:00
Ryan Houdek ce97ef0ab1 Scripts: Move arch check first in InstallFEX
Don't give people false hope that the script might work on distros that
aren't Ubuntu.

Fixes #5420
2026-04-15 13:01:37 -07:00
LC 2ea0de92f4 Merge pull request #5418 from Sonicadvance1/135
ArchHelpers: Allow atomic memory operations in non-JIT handler
2026-04-15 07:22:25 -04:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
Ryan Houdek 9681559d56 Docs: Update for release FEX-2604 2026-04-09 13:45:35 -07:00
Ryan Houdek b478e4845f Merge pull request #5417 from tiopex/main
FEXRootFSFetcher: clear Unknown when distro is set on the CLI
2026-04-09 13:42:55 -07:00
tpietrus 1fa5104076 FEXRootFSFetcher: clear Unknown when distro is set on the CLI 2026-04-09 07:53:53 +02:00
LC ce65f5376f Merge pull request #5415 from Sonicadvance1/133
FEX: Workaround Docker seccomp bug
2026-04-08 11:39:56 -04:00
LC 0695249fc8 Merge pull request #5413 from Sonicadvance1/132
Arm64EC: Invert suspend doorbell and move out of hot path
2026-04-06 19:49:49 -04:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
LC 251398a7cb Merge pull request #5416 from Sonicadvance1/134
OpcodeDispatcher: Fixes nop encoded prefetch instruction
2026-04-06 19:45:52 -04:00
Ryan Houdek 2e6a7f869c OpcodeDispatcher: Fixes nop encoded prefetch instruction
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.

Fixes `Devil May Cry 4`
2026-04-06 11:04:46 -07:00
Ryan Houdek f308162334 FEX: Workaround Docker seccomp bug
Docker's seccomp filter fails to follow AAPCS64 and SysV zero-extension
rules.  For values smaller than 64-bit they were required in their
seccomp filters to truncate the value to the specific size but do not.
Instead they do a 64-bit comparison operation against smaller arguments
(in this case 32-bit). This means 64-bit -1 and 32-bit -1 passed through
have different values for this `personality` syscall.

The real fix would be for Docker to audit their seccomp filter rules and
ensure they zero-extend every argument that is smaller than 64-bit, but
we don't control that. So there is likely to be more bugs in their
filter that we encounter, this is just an easy one to resolve.
2026-04-06 09:51:15 -07:00
Ryan Houdek db4867839c Arm64EC: Invert suspend doorbell and move out of hot path
This was causing a surprisingly high amount of branch mispredicts in
Death Stranding. Suspend doorbell is fairly rare so just invert the
check and move the target down out of the hot path. Then the doorbell
handling code will trampoline to the correct location still.
2026-04-03 20:07:37 -07:00
LC 73ffff7d22 Merge pull request #5409 from Sonicadvance1/130
OpcodeDispatcher: Special case optimize a broadcast
2026-04-03 20:23:39 -04:00
LC dc48a4f73c Merge pull request #5406 from Sonicadvance1/128
FEXRootFSFetcher: Improve hashing performance
2026-04-03 00:26:44 -04:00
Ryan Houdek efbccccdc0 InstcountCI: Update 2026-04-02 18:45:22 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Ryan Houdek 12cfe8fc37 InstcountCI: Add instruction found in Death Stranding 2 2026-04-02 18:25:19 -07:00
Tony Wasserka c6d2ce043f Merge pull request #5388 from Sonicadvance1/123
Config: Finish wiring up Regex app overrides
2026-04-02 11:00:37 +02:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek 4018d23c39 Config: Finish wiring up Regex app overrides
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.

Instead compare it to the appname that now gets passed over to it for
matching.

This allows us to do the following:
```
{
    "Config": {
        "ProfileStats": "1",
        "X87ReducedPrecision": "1",
        "TSOEnabled": "1",
        "VectorTSOEnabled": "0",
        "MemcpySetTSOEnabled": "0",
        "HalfBarrierTSOEnabled":"1",
        "MaxInst": "500",
        "Multiblock": "1"
    },
    "AppOverrides" : {
        "setup*" : {
            "Comment": [
                "292030 - The Witcher 3: Wild Hunt"
            ],
            "X87ReducedPrecision": "0"
        }
     }
}
```

Based on #121 which needs to get merged first.

Code Review

Code Review: Class deletion
2026-04-01 10:46:02 -07:00
badumbatish 81d4e8fe9d Initial implementation for regex engine
Add support for question mark and plus mark in regex, supply testing for star

Added more characters to the regex alphabets, add more test case

Added support for regex matching of configs, awaiting reviews

Rename variable to CamelCase

Addresses PR reviews

Remove unnecessary features and test cases

Rewrite to naive regex with dp

Addresses PR reviews

Build fixes

Code Review
2026-04-01 10:45:59 -07:00
Tony Wasserka 34b48c4069 Merge pull request #5383 from Sonicadvance1/119
FEXGetConfig: Test for showing fault granularity
2026-04-01 10:48:18 +02:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
LC ae3fa6a836 Merge pull request #5403 from Sonicadvance1/127
IR: Adds support for printing strings
2026-03-31 16:23:42 -04:00
Ryan Houdek ba93bdd66d FEXRootFSFetcher: Improve hashing performance
Don't use pread, instead map the file and madvise larger blocks. This
removes copying overhead as its just mapping file pages in instead.
Also splits the implementation of file reading from hashing to make
tinkering less involved, as if I want more performance out of this (say
due to live hashing) then it's easier to tinker.

Improves hashing performance from ~2.2GB/s to ~3.6GB/s on my system,
which is CPU bounded by xxhash here.
2026-03-30 18:16:46 -07:00
Ryan Houdek 1fa0b37fac FEXGetConfig: Test for showing fault granularity
Useful for seeing if behaviour has changed. Useful with the
`--tso-emulation-info` option to show hardware behaviour
2026-03-30 12:32:00 -07:00
LC 5b4a5969cc Merge pull request #5401 from Sonicadvance1/126
Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
2026-03-28 22:25:59 -04:00
Ryan Houdek 941a7934ef InstcountCI: Update 2026-03-28 18:06:54 -07:00
Ryan Houdek 474439ab4b InstcountCI: Update 2026-03-28 17:56:23 -07:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek e9a9cc5bc3 unittests/FEXLinuxTests: Adds an MXCSR signal test
Ensures that the MXCSR value stays the same with a signal inbetween that
modifies it.
2026-03-28 17:49:12 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Ryan Houdek f78e194cf7 Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
This was causing an unfortunate set of circumstances where Dark Souls
III was modifying MXCSR and we weren't saving it, cause the value to change
from 0x9fc0 to 0.

This "enabled" float exceptions by unmasking the exception masks in
MXCSR. This in turn had Dark Souls III's `expf` function to fault out,
as it checks if the MXCSR exception masks are set or not for determining
if underflow should assert or not.

Wow64/arm64ec has a similar problem where it always sets back to default
on signal. Which means game lose DAZ, but I'm not fixing that bug right
now.

Fixes #5391
2026-03-28 17:44:45 -07:00
LC b77ddcf1a7 Merge pull request #5398 from Sonicadvance1/125
Allocators: Remove legacy NOREPLACE handling
2026-03-27 08:14:05 -04:00
Ryan Houdek 3ef677537d Merge pull request #5397 from neobrain/fix_elfreads
LinuxSyscalls: Skip reading ELF files when code caching is disabled
2026-03-26 14:33:26 -07:00
Tony Wasserka 1df1265ed1 LinuxSyscalls: Skip reading ELF files when code caching is disabled
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.

Closes #5390
2026-03-26 21:40:20 +01:00
Ryan Houdek f0854a16fe Allocators: Remove legacy NOREPLACE handling
We needed this handling on old kernels that didn't understand the
NOREPLACE flag. We no longer support kernels this old, so remove some of
this vestigial code.
2026-03-26 13:36:29 -07:00
Ryan Houdek 6bd476fb03 Merge pull request #5395 from lioncash/cpuid
CPUID: Add basic stub handling for AVX10 info
2026-03-25 17:16:24 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Tony Wasserka 5149ebc70e Merge pull request #5389 from Sonicadvance1/124
SMCTracking: Remove relocation log
2026-03-24 11:12:10 +01:00
Ryan Houdek e92a6a5803 SMCTracking: Remove relocation log
Holy jeez does this thing spam.
2026-03-23 19:55:11 -07:00
Ryan Houdek 6da963a695 Merge pull request #5394 from neobrain/fix_eager_relocation_parsing
LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded
2026-03-23 14:01:20 -07:00
Tony Wasserka 2a74489858 LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded 2026-03-23 17:47:47 +01:00
Ryan Houdek 65a436ca98 Merge pull request #5392 from neobrain/fix_code_cache_dupfd
LinuxSyscalls: Re-open file descriptors for parsing ELF headers
2026-03-23 09:42:45 -07:00
LC bc533c8050 Merge pull request #5393 from neobrain/fix_code_cache_glibc_assert
CodeCache: Fix glibc debug mode assertion
2026-03-22 14:07:27 -04:00
Tony Wasserka fd6cea4698 CodeCache: Fix glibc debug mode assertion
If begin == end, the first vector::erase() call would invalidate the begin
iterator.
2026-03-22 10:11:53 +01:00
Tony Wasserka a1aa1658ec LinuxSyscalls: Re-open file descriptors for parsing ELF headers
File descriptors returned by dup() share state with the original FD, so we
need to use open() to create a fully independent object.

Fixes #5379
2026-03-22 10:09:37 +01:00
LC 5c4c468d13 Merge pull request #5387 from Sonicadvance1/122
github: Stop running unittests always on build failure
2026-03-20 00:58:09 -04:00
Ryan Houdek 69ef1658cc github: Stop running unittests always on build failure
This was taking too much time.
2026-03-19 20:08:10 -07:00
Ryan Houdek 8c72aa76a0 Merge pull request #5385 from lioncash/ilog
MemoryOps: Collapse duplicate add/sub in Memset
2026-03-19 18:59:17 -07:00
Lioncache 928a932a43 MemoryOps: Collapse duplicate add/sub in Memset
We can just use ilog2 to deduplicate this a bit.
2026-03-19 21:34:03 -04:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Ryan Houdek 68480f6e43 Merge pull request #5356 from lioncash/mops
MemoryOps: Drop MOPS handling into place for MemSet/MemCpy
2026-03-19 18:04:14 -07:00
Ryan Houdek 42291540ab Merge pull request #5343 from pmatos/f64-sin-cos-tan
JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path
2026-03-19 17:51:50 -07:00
LC 194eb69838 Merge pull request #5384 from Sonicadvance1/120
Cmake: Default to release builds with a message
2026-03-19 19:42:39 -04:00
Ryan Houdek c1d27fa453 Cmake: Default to release builds with a message
People keep forgetting to set this and have a worse experience.
Default to a Release build, which ensures optimizations are enabled and
assertions are disabled.
2026-03-19 16:25:40 -07:00
Paulo Matos 9d5f7caa79 instcountci: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Paulo Matos a1d78dceb0 asm_tests: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Lioncache 600f4bb2b6 unittests: Add overlapping tests for MemCpy 2026-03-19 11:55:02 -04:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache e862f8f86c unittests: Add specific paths for MOPS 2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
LC c18fb3cb78 Merge pull request #5382 from Sonicadvance1/118
FEXpidof: Fixes another missing std::filesystem throw
2026-03-18 18:07:08 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Ryan Houdek 740350c8ea FEXpidof: Fixes another missing std::filesystem throw
Turns out std::filesystem::exists throws as well if there was an
underlying OS API failure.
2026-03-18 12:10:22 -07:00
Tony Wasserka 9f9b20eac0 Merge pull request #5378 from Sonicadvance1/117
SMCTracking: Move read check up for ELF parsing
2026-03-18 12:04:23 +01:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Tony Wasserka 1ea24f3d6c LinuxSyscalls: Enable delayed code cache load for ELF files
Specifically this is needed if any ELF relocations cover read-only code
sections, which is indicated in the ELF headers via DT_TEXTREL/DF_TEXTREL.
2026-03-18 11:59:41 +01:00
Tony Wasserka 226bd51afe LinuxSyscalls: Implement delayed cache load for binaries that require ELF/PE relocations 2026-03-18 11:58:26 +01:00
Tony Wasserka 293568be36 FEXOfflineCompiler: Apply relocations to loaded ELF binaries 2026-03-18 11:57:38 +01:00
Tony Wasserka 152fe81d16 LinuxSyscalls: Parse and provide ELF relocation information to the JIT 2026-03-18 11:57:28 +01:00
Ryan Houdek 494dd64c50 Merge pull request #5372 from CxnYusuf/add-fisttp-tests
Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives
2026-03-17 16:55:56 -07:00
Ryan Houdek 56de0d1ab4 Merge pull request #5377 from Sonicadvance1/116
FEXServer: Try both fusermount and fusermount3
2026-03-17 16:55:42 -07:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Ryan Houdek de11c05d2a Merge pull request #5380 from neobrain/fix_cc_32bit_constants
Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
2026-03-17 13:34:20 -07:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Ryan Houdek 63e37b7cbb SMCTracking: Move read check up for ELF parsing 2026-03-17 12:43:54 -07:00
Ryan Houdek 9ee329034f FEXServer: Try both fusermount and fusermount3
Apparently some distros don't symlink these, so try both with the newer
fusermount3 going first as its the common path now.

Fixes #5375
2026-03-17 12:36:49 -07:00
Ryan Houdek f3e904207b Merge pull request #5376 from OFFTKP/flag
Add test for shifts preserving flags
2026-03-17 09:40:59 -07:00
LC e4ae6ce635 Merge pull request #5373 from Sonicadvance1/115
code-format-helper: Another dependabot upgrade
2026-03-17 10:10:46 -04:00
Paris Oplopoios f7d76255ad Add test for shift preserving flags
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-03-17 15:50:33 +02:00
Tony Wasserka ea45f9c694 FEXCore/VectorRegType: Use vector_size on GCC
GCC does not support neon_vector_type and silently ignores that attribute,
but vector_size(16) seems to have the same effect.
2026-03-16 19:15:05 +01:00
Tony Wasserka 1d449c0f58 FEXCore/Utils: Add quotes around preprocessor errors 2026-03-16 19:15:05 +01:00
Tony Wasserka 9ecc991043 LibraryForwarding: Remove unnecessary const qualifier 2026-03-16 19:15:05 +01:00
Tony Wasserka 4ef834859c SignalDelegator: Don't use the same name for two different symbols 2026-03-16 19:15:05 +01:00
Tony Wasserka fa082bc5c4 FileManagement: Fix ambiguous name reference 2026-03-16 19:15:05 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Tony Wasserka 321d4d84d7 LinuxSyscalls: Don't cast away qualifiers 2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Tony Wasserka d27c9d3f98 CodeEmitter: Fix ambigious ExtendedType declaration 2026-03-16 18:50:01 +01:00
Tony Wasserka fbef482265 CMake: Link against libatomic if compiling with GCC 2026-03-16 18:50:01 +01:00
Tony Wasserka a57926ac57 CMake: Explicitly demote -Wchanges-meaning diagnostics to warnings on GCC 2026-03-16 18:50:01 +01:00
Ryan Houdek 70a7137e62 code-format-helper: Another dependabot upgrade 2026-03-15 20:45:17 -07:00
CxnYusuf 2d3a08a362 Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives 2026-03-16 03:10:38 +01:00
LC cc02edb3f6 Merge pull request #5371 from Sonicadvance1/114
Syscalls: Fixes crash in ELF parsing code
2026-03-15 20:36:00 -04:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 24675969cc Merge pull request #5366 from Sonicadvance1/112
JIT: Use struct for Spill/Fill default arguments
2026-03-15 15:20:13 -07:00
Ryan Houdek a0cba1194e Syscalls: Fixes crash in ELF parsing code
When an application maps a file as PROT_NONE, we can't check if it is an
ELF. Was causing a crash in `Cisco Packet Tracer`.
2026-03-15 15:16:29 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek 6177ab957b Merge pull request #5363 from Sonicadvance1/109
External/code-format-helper: Update dependencies
2026-03-13 11:52:07 -07:00
Ryan Houdek 5c1300a2b9 Merge pull request #5365 from Sonicadvance1/111
Config: Fix issue with config overrides
2026-03-13 11:51:51 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek dc0162122f Config: Fix issue with config overrides
Accidentally was checking for Config override in the combination of
portable config and `FEX_APP_CONFIG_LOCATION`.

Fixes an early crash in PV.
2026-03-12 17:46:05 -07:00
Ryan Houdek ae491fb15b External/code-format-helper: Update dependencies
Removes dependabot alert.
2026-03-12 15:31:02 -07:00
Ryan Houdek 957c1fc420 Merge pull request #5357 from Sonicadvance1/106
Config: Enable Dynamic L1 and Disabled L2 caches by default
2026-03-11 14:12:38 -07:00
Ryan Houdek 86acfb35aa Config: Enable Dynamic L1 and Disabled L2 caches by default
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.

Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
2026-03-11 13:42:32 -07:00
LC e27d12ee5e Merge pull request #5359 from Sonicadvance1/107
OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
2026-03-11 08:40:04 -04:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek 9d4a71b57a Merge pull request #5355 from wsxarcher/patch-1
Handle zero length in ChangeProtectionFlags
2026-03-10 08:40:16 -07:00
Marco Bartoli 3462dc3e14 Handle zero length in ChangeProtectionFlags
Add a no-op for zero length in ChangeProtectionFlags.

This fixes AMD Vivado 2025.2 which tries to mprotect with 0 as size and merge strategies fails:

```
Unexpected ChangeProtectionFlags Merge strategy! [0x400000, 0x401000) Versus [0x0, 0x0)
```
2026-03-10 11:34:52 +01:00
LC d21351e66e Merge pull request #5354 from Sonicadvance1/104
Config: Fixes
2026-03-09 21:51:05 -04:00
Ryan Houdek 3204d20335 Config: Fix priorities of config paths
Fixes b4a87d8c0b

`FEX_APP_CONFIG_LOCATION` wasn't overriding paths properly anymore once
that commit landed. Instead legacy `~/.fex-emu/` path would get returned
if it existed first.

Ensures that it returns first, before `STEAM_COMPAT_DATA_PATH` even.
2026-03-09 18:20:56 -07:00
Ryan Houdek a519489d80 CMake: Make sure not to compile Steam tools on mingw 2026-03-09 17:25:59 -07:00
Ryan Houdek 5558c3a35a Merge pull request #5353 from lioncash/group
HostFeatures: Group feature ifdefs together more
2026-03-09 15:45:55 -07:00
Ryan Houdek 58d9755314 Merge pull request #5352 from Sonicadvance1/103
gitlab-ci: Update requirements
2026-03-09 15:45:48 -07:00
Lioncache 498ba0a384 HostFeatures: Group feature ifdefs together more
Makes it a little nicer to see everything grouped together.
2026-03-09 18:27:56 -04:00
Ryan Houdek 5a5477e895 gitlab-ci: Update requirements 2026-03-09 14:00:08 -07:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
LC afc7248912 Merge pull request #5347 from Sonicadvance1/102
CPUBackend: Enable Transparent Huge Pages on JIT buffers
2026-03-09 14:54:08 -04:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
520 changed files with 41258 additions and 20746 deletions

No files matched your search

+14 -13
View File
@@ -49,6 +49,7 @@ jobs:
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
id: build
run: cmake --build build
- name: Install
@@ -56,40 +57,40 @@ jobs:
# GCC tests
- name: GCC64 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_64
- name: GCC32 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_32
# API tests
- name: API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: api_tests
- name: FEXCore API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fexcore_apitests
# ARM emission tests
- name: ARM Emitter Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: emitter_tests
# Linux tests
- name: FEX Linux Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fex_linux_tests_all
@@ -98,13 +99,13 @@ jobs:
# Thunking
- name: Thunkgen tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: thunkgen_tests
- name: Test GL No-Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_nothunks
@@ -112,7 +113,7 @@ jobs:
DISPLAY: ':0'
- name: Test GL Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_thunks
@@ -121,28 +122,28 @@ jobs:
# ASM tests
- name: ASM Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: asm_tests
# POSIX tests
- name: POSIX Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: posix_tests
# GVisor tests
- name: GVisor Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gvisor_tests
# Struct verifier tests
- name: Struct verifier tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: struct_verifier
+14
View File
@@ -33,3 +33,17 @@ runs:
- name: Install
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_${{ inputs.target }} -t install
- name: Configure UnixLib
shell: bash
run: |
cmake -S Source/Windows/UnixLib -B build_unixlib_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE \
-G Ninja -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-unix -DCMAKE_INSTALL_PREFIX=/usr
- name: Build UnixLib
shell: bash
run: cmake --build build_unixlib_${{ inputs.target }}
- name: Install UnixLib
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_unixlib_${{ inputs.target }} -t install
+3 -1
View File
@@ -50,6 +50,8 @@ jobs:
with:
overwrite: true
name: wine_dll_artifacts
path: ${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
path: |
${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
${{ github.workspace }}/install/usr/lib/wine/aarch64-unix/lib*.so
retention-days: 60
compression-level: 9
+40 -1
View File
@@ -1,3 +1,17 @@
spec:
inputs:
PROMOTE_BRANCH:
description: "Branch to promote the build to. Empty means no promotion."
default: "bleeding-edge"
---
workflow:
rules:
- when: always
variables:
PROMOTE_BRANCH: $[[ inputs.PROMOTE_BRANCH ]]
variables:
DEBIAN_FRONTEND: noninteractive
GIT_SUBMODULE_STRATEGY: recursive
@@ -5,7 +19,8 @@ variables:
CC: clang
CXX: clang++
aarch64:
build:
stage: build
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
@@ -30,3 +45,27 @@ aarch64:
untracked: false
paths:
- install/
promote:
stage: deploy
variables:
GIT_STRATEGY: none
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
- linux
- arm64
- aarch64
rules:
- if: '$PROMOTE_BRANCH'
before_script:
- apt-get -y update
- apt-get install -y tmux curl
script:
# comment out to debug: SSH in via GCP, go down the container and attach to the session (with `tmux attach -t debug`)
# - tmux new-session -d -s debug
# - while tmux has-session -t debug 2>/dev/null; do sleep 1; done
# ref controls which fex-depot code runs the pipeline, while VERSION_PARAM controls which fex branch's artifacts that pipeline downloads.
- >
curl --fail --location --request POST --form token=${FEX_DEPOT_TRIGGER_TOKEN} --form ref=master --form "variables[PROMOTE_BRANCH]=${PROMOTE_BRANCH}" --form "variables[VERSION_PARAM]=${CI_COMMIT_REF_NAME}" "${CI_API_V4_URL}/projects/fex%2Ffex-depot/trigger/pipeline"
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+24 -3
View File
@@ -172,6 +172,13 @@ set(TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set(OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version")
set(OVERRIDE_HASH "detect" CACHE STRING "Override the FEX git hash")
get_property(IS_MULTI_CONFIG GLOBAL PROPERTY GENERATOR_IS_MULTI_CONFIG)
if (NOT IS_MULTI_CONFIG AND NOT CMAKE_BUILD_TYPE)
set(CMAKE_BUILD_TYPE Release
CACHE STRING "Choose the type of build." FORCE)
message(STATUS "No build type set, defaulting to a Release build")
endif()
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
set(ENABLE_ASSERTIONS TRUE)
@@ -187,6 +194,8 @@ if (ENABLE_GDB_SYMBOLS)
add_compile_definitions(GDB_SYMBOLS_ENABLED=1)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -244,8 +253,15 @@ endif()
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
execute_process(COMMAND "${CCACHE_PROGRAM}" --print-version
OUTPUT_VARIABLE CCACHE_VERSION OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "Enabling ccache ${CCACHE_VERSION}")
if (CCACHE_VERSION VERSION_GREATER_EQUAL "4.8")
# Set sloppiness to enable caching even for files that use __DATE__/__TIME__ macros
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM} sloppiness=time_macros")
else()
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
endif()
endif()
endif()
@@ -446,6 +462,11 @@ if(ENUM_ENUM_WARNING)
add_compile_options(-Wno-deprecated-enum-enum-conversion)
endif()
# GCC enables -Wchanges-meaning by default and treats some cases as an error
if(CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
add_compile_options(-Wno-error=changes-meaning)
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
@@ -681,6 +702,6 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
if (NOT MINGW AND BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
-132
View File
@@ -1,132 +0,0 @@
{
"environments": [
{
"BuildPath": "${projectDir}\\out\\build\\${name}",
"InstallPath": "${projectDir}\\out\\install\\${name}",
"clangcl": "clang-cl.exe",
"cc": "clang",
"cxx": "clang++"
}
],
"configurations": [
{
"name": "WSL-Clang-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "WSL-Clang-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "x86-Clang-Cross-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "x64-Clang-Cross-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "Linux-Clang-Remote-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"cmakeExecutable": "/usr/bin/cmake",
"remoteCopySourcesExclusionList": [ ".vs", ".vscode", ".git", ".github", "build", "out", "bin" ],
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "linux_clang_x64" ],
"remoteMachineName": "${env.fexremote}",
"remoteCMakeListsRoot": "$HOME/projects/.vs/${projectDirName}/src",
"remoteBuildRoot": "$HOME/projects/.vs/${projectDirName}/build/${name}",
"remoteInstallRoot": "$HOME/projects/.vs/${projectDirName}/install/${name}",
"remoteCopySources": true,
"rsyncCommandArgs": "-t --delete --delete-excluded",
"remoteCopyBuildOutput": false,
"remoteCopySourcesMethod": "rsync",
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": []
}
]
}
+1
View File
@@ -0,0 +1 @@
No AI/ML/LLM/etc code contributions.
+2 -2
View File
@@ -311,7 +311,7 @@ class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {.ExtendedType {
, MetaType {.Extended {
.Header = {.MemType = TYPE_EXTENDED},
.rm = rm,
.Option = Option,
@@ -340,7 +340,7 @@ public:
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
} Extended;
struct {
HeaderStruct Header;
IndexType Index;
+50 -50
View File
@@ -3627,8 +3627,8 @@ public:
void strb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3650,8 +3650,8 @@ public:
}
void ldrb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3673,8 +3673,8 @@ public:
}
void ldrsb(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3696,8 +3696,8 @@ public:
}
void ldrsb(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3719,8 +3719,8 @@ public:
}
void strh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -3742,8 +3742,8 @@ public:
}
void ldrh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -3765,8 +3765,8 @@ public:
}
void ldrsh(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3788,8 +3788,8 @@ public:
}
void ldrsh(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3811,8 +3811,8 @@ public:
}
void str(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3834,8 +3834,8 @@ public:
}
void ldr(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3857,8 +3857,8 @@ public:
}
void ldrsw(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsw(rt, MemSrc.rn);
} else {
@@ -3880,8 +3880,8 @@ public:
}
void str(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3903,8 +3903,8 @@ public:
}
void ldr(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3926,8 +3926,8 @@ public:
}
void prfm(ARMEmitter::Prefetch prfop, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
prfm(prfop, MemSrc.rn);
} else {
@@ -3946,9 +3946,9 @@ public:
void strb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3970,9 +3970,9 @@ public:
}
void ldrb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3994,8 +3994,8 @@ public:
}
void strh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -4017,8 +4017,8 @@ public:
}
void ldrh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -4040,8 +4040,8 @@ public:
}
void str(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4063,8 +4063,8 @@ public:
}
void ldr(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4086,8 +4086,8 @@ public:
}
void str(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4109,8 +4109,8 @@ public:
}
void ldr(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4132,8 +4132,8 @@ public:
}
void str(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4155,8 +4155,8 @@ public:
}
void ldr(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
+7
View File
@@ -46,6 +46,13 @@
"@PREFIX_LIB@/libwayland-client.so.0",
"@PREFIX_LIB@/libwayland-client.so.0.20.0"
]
},
"cuda": {
"Library" : "libcuda-guest.so",
"Overlay": [
"@PREFIX_LIB@/libcuda.so",
"@PREFIX_LIB@/libcuda.so.1"
]
}
}
}
+141 -90
View File
@@ -1,32 +1,37 @@
#
# This file is autogenerated by pip-compile with Python 3.13
# This file is autogenerated by pip-compile with Python 3.14
# by the following command:
#
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
black==26.3.1 \
--hash=sha256:0126ae5b7c09957da2bdbd91a9ba1207453feada9e9fe51992848658c6c8e01c \
--hash=sha256:0f76ff19ec5297dd8e66eb64deda23631e642c9393ab592826fd4bdc97a4bce7 \
--hash=sha256:28ef38aee69e4b12fda8dba75e21f9b4f979b490c8ac0baa7cb505369ac9e1ff \
--hash=sha256:2bd5aa94fc267d38bb21a70d7410a89f1a1d318841855f698746f8e7f51acd1b \
--hash=sha256:2c50f5063a9641c7eed7795014ba37b0f5fa227f3d408b968936e24bc0566b07 \
--hash=sha256:2d6bfaf7fd0993b420bed691f20f9492d53ce9a2bcccea4b797d34e947318a78 \
--hash=sha256:41cd2012d35b47d589cb8a16faf8a32ef7a336f56356babd9fcf70939ad1897f \
--hash=sha256:474c27574d6d7037c1bc875a81d9be0a9a4f9ee95e62800dab3cfaadbf75acd5 \
--hash=sha256:5602bdb96d52d2d0672f24f6ffe5218795736dd34807fd0fd55ccd6bf206168b \
--hash=sha256:5e9d0d86df21f2e1677cc4bd090cd0e446278bcbbe49bf3659c308c3e402843e \
--hash=sha256:5ed0ca58586c8d9a487352a96b15272b7fa55d139fc8496b519e78023a8dab0a \
--hash=sha256:6c54a4a82e291a1fee5137371ab488866b7c86a3305af4026bdd4dc78642e1ac \
--hash=sha256:6e131579c243c98f35bce64a7e08e87fb2d610544754675d4a0e73a070a5aa3a \
--hash=sha256:855822d90f884905362f602880ed8b5df1b7e3ee7d0db2502d4388a954cc8c54 \
--hash=sha256:86a8b5035fce64f5dcd1b794cf8ec4d31fe458cf6ce3986a30deb434df82a1d2 \
--hash=sha256:8a33d657f3276328ce00e4d37fe70361e1ec7614da5d7b6e78de5426cb56332f \
--hash=sha256:92c0ec1f2cc149551a2b7b47efc32c866406b6891b0ee4625e95967c8f4acfb1 \
--hash=sha256:9a5e9f45e5d5e1c5b5c29b3bd4265dcc90e8b92cf4534520896ed77f791f4da5 \
--hash=sha256:afc622538b430aa4c8c853f7f63bc582b3b8030fd8c80b70fb5fa5b834e575c2 \
--hash=sha256:b07fc0dab849d24a80a29cfab8d8a19187d1c4685d8a5e6385a5ce323c1f015f \
--hash=sha256:b5e6f89631eb88a7302d416594a32faeee9fb8fb848290da9d0a5f2903519fc1 \
--hash=sha256:bf9bf162ed91a26f1adba8efda0b573bc6924ec1408a52cc6f82cb73ec2b142c \
--hash=sha256:c7e72339f841b5a237ff14f7d3880ddd0fc7f98a1199e8c4327f9a4f478c1839 \
--hash=sha256:ddb113db38838eb9f043623ba274cfaf7d51d5b0c22ecb30afe58b1bb8322983 \
--hash=sha256:dfdd51fc3e64ea4f35873d1b3fb25326773d55d2329ff8449139ebaad7357efb \
--hash=sha256:f1cd08e99d2f9317292a311dfe578fd2a24b15dbce97792f9c4d752275c1fa56 \
--hash=sha256:f89f2ab047c76a9c03f78d0d66ca519e389519902fa27e7a91117ef7611c0568
# via
# -r requirements_formatting.txt.in
# darker
@@ -205,56 +210,56 @@ click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==46.0.5 \
--hash=sha256:02f547fce831f5096c9a567fd41bc12ca8f11df260959ecc7c3202555cc47a72 \
--hash=sha256:039917b0dc418bb9f6edce8a906572d69e74bd330b0b3fea4f79dab7f8ddd235 \
--hash=sha256:1abfdb89b41c3be0365328a410baa9df3ff8a9110fb75e7b52e66803ddabc9a9 \
--hash=sha256:2ae6971afd6246710480e3f15824ed3029a60fc16991db250034efd0b9fb4356 \
--hash=sha256:2b7a67c9cd56372f3249b39699f2ad479f6991e62ea15800973b956f4b73e257 \
--hash=sha256:351695ada9ea9618b3500b490ad54c739860883df6c1f555e088eaf25b1bbaad \
--hash=sha256:38946c54b16c885c72c4f59846be9743d699eee2b69b6988e0a00a01f46a61a4 \
--hash=sha256:3b4995dc971c9fb83c25aa44cf45f02ba86f71ee600d81091c2f0cbae116b06c \
--hash=sha256:3ce58ba46e1bc2aac4f7d9290223cead56743fa6ab94a5d53292ffaac6a91614 \
--hash=sha256:3ee190460e2fbe447175cda91b88b84ae8322a104fc27766ad09428754a618ed \
--hash=sha256:4108d4c09fbbf2789d0c926eb4152ae1760d5a2d97612b92d508d96c861e4d31 \
--hash=sha256:420d0e909050490d04359e7fdb5ed7e667ca5c3c402b809ae2563d7e66a92229 \
--hash=sha256:47fb8a66058b80e509c47118ef8a75d14c455e81ac369050f20ba0d23e77fee0 \
--hash=sha256:4c3341037c136030cb46e4b1e17b7418ea4cbd9dd207e4a6f3b2b24e0d4ac731 \
--hash=sha256:4d7e3d356b8cd4ea5aff04f129d5f66ebdc7b6f8eae802b93739ed520c47c79b \
--hash=sha256:4d8ae8659ab18c65ced284993c2265910f6c9e650189d4e3f68445ef82a810e4 \
--hash=sha256:4e817a8920bfbcff8940ecfd60f23d01836408242b30f1a708d93198393a80b4 \
--hash=sha256:50bfb6925eff619c9c023b967d5b77a54e04256c4281b0e21336a130cd7fc263 \
--hash=sha256:556e106ee01aa13484ce9b0239bca667be5004efb0aabbed28d353df86445595 \
--hash=sha256:582f5fcd2afa31622f317f80426a027f30dc792e9c80ffee87b993200ea115f1 \
--hash=sha256:5be7bf2fb40769e05739dd0046e7b26f9d4670badc7b032d6ce4db64dddc0678 \
--hash=sha256:60ee7e19e95104d4c03871d7d7dfb3d22ef8a9b9c6778c94e1c8fcc8365afd48 \
--hash=sha256:61aa400dce22cb001a98014f647dc21cda08f7915ceb95df0c9eaf84b4b6af76 \
--hash=sha256:68f68d13f2e1cb95163fa3b4db4bf9a159a418f5f6e7242564fc75fcae667fd0 \
--hash=sha256:7d1f30a86d2757199cb2d56e48cce14deddf1f9c95f1ef1b64ee91ea43fe2e18 \
--hash=sha256:7d731d4b107030987fd61a7f8ab512b25b53cef8f233a97379ede116f30eb67d \
--hash=sha256:803812e111e75d1aa73690d2facc295eaefd4439be1023fefc4995eaea2af90d \
--hash=sha256:80a8d7bfdf38f87ca30a5391c0c9ce4ed2926918e017c29ddf643d0ed2778ea1 \
--hash=sha256:8293f3dea7fc929ef7240796ba231413afa7b68ce38fd21da2995549f5961981 \
--hash=sha256:8456928655f856c6e1533ff59d5be76578a7157224dbd9ce6872f25055ab9ab7 \
--hash=sha256:890bcb4abd5a2d3f852196437129eb3667d62630333aacc13dfd470fad3aaa82 \
--hash=sha256:94a76daa32eb78d61339aff7952ea819b1734b46f73646a07decb40e5b3448e2 \
--hash=sha256:9f16fbdf4da055efb21c22d81b89f155f02ba420558db21288b3d0035bafd5f4 \
--hash=sha256:a3d1fae9863299076f05cb8a778c467578262fae09f9dc0ee9b12eb4268ce663 \
--hash=sha256:a3d507bb6a513ca96ba84443226af944b0f7f47dcc9a399d110cd6146481d24c \
--hash=sha256:abace499247268e3757271b2f1e244b36b06f8515cf27c4d49468fc9eb16e93d \
--hash=sha256:ba2a27ff02f48193fc4daeadf8ad2590516fa3d0adeeb34336b96f7fa64c1e3a \
--hash=sha256:bc84e875994c3b445871ea7181d424588171efec3e185dced958dad9e001950a \
--hash=sha256:bfd56bb4b37ed4f330b82402f6f435845a5f5648edf1ad497da51a8452d5d62d \
--hash=sha256:c18ff11e86df2e28854939acde2d003f7984f721eba450b56a200ad90eeb0e6b \
--hash=sha256:c3bcce8521d785d510b2aad26ae2c966092b7daa8f45dd8f44734a104dc0bc1a \
--hash=sha256:c4143987a42a2397f2fc3b4d7e3a7d313fbe684f67ff443999e803dd75a76826 \
--hash=sha256:c69fd885df7d089548a42d5ec05be26050ebcd2283d89b3d30676eb32ff87dee \
--hash=sha256:ced80795227d70549a411a4ab66e8ce307899fad2220ce5ab2f296e687eacde9 \
--hash=sha256:d66e421495fdb797610a08f43b05269e0a5ea7f5e652a89bfd5a7d3c1dee3648 \
--hash=sha256:d861ee9e76ace6cf36a6a89b959ec08e7bc2493ee39d07ffe5acb23ef46d27da \
--hash=sha256:e9251e3be159d1020c4030bd2e5f84d6a43fe54b6c19c12f51cde9542a2817b2 \
--hash=sha256:f145bba11b878005c496e93e257c1e88f154d278d2638e6450d17e0f31e558d2 \
--hash=sha256:fe346b143ff9685e40192a4960938545c699054ba11d4f9029f94751e3f71d87
cryptography==48.0.0 \
--hash=sha256:0890f502ddf7d9c6426129c3f49f5c0a39278ed7cd6322c8755ffca6ee675a13 \
--hash=sha256:0c558d2cdffd8f4bbb30fc7134c74d2ca9a476f830bb053074498fbc86f41ed6 \
--hash=sha256:16cd65b9330583e4619939b3a3843eec1e6e789744bb01e7c7e2e62e33c239c8 \
--hash=sha256:18349bbc56f4743c8b12dc32e2bccb2cf83ee8b69a3bba74ef8ae857e26b3d25 \
--hash=sha256:1e2d54c8be6152856a36f0882ab231e70f8ec7f14e93cf87db8a2ed056bf160c \
--hash=sha256:22a5cb272895dce158b2cacdfdc3debd299019659f42947dbdac6f32d68fe832 \
--hash=sha256:27241b1dc9962e056062a8eef1991d02c3a24569c95975bd2322a8a52c6e5e12 \
--hash=sha256:2b4d59804e8408e2fea7d1fbaf218e5ec984325221db76e6a241a9abd6cdd95c \
--hash=sha256:2eb992bbd4661238c5a397594c83f5b4dc2bc5b848c365c8f991b6780efcc5c7 \
--hash=sha256:369a6348999f94bbd53435c894377b20ab95f25a9065c283570e70150d8abc3c \
--hash=sha256:3cb07a3ed6431663cd321ea8a000a1314c74211f823e4177fefa2255e057d1ec \
--hash=sha256:40ba1f85eaa6959837b1d51c9767e230e14612eea4ef110ee8854ada22da1bf5 \
--hash=sha256:4defde8685ae324a9eb9d818717e93b4638ef67070ac9bc15b8ca85f63048355 \
--hash=sha256:55b7718303bf06a5753dcdccf2f3945cf18ad7bffde41b61226e4db31ab89a9c \
--hash=sha256:561215ea3879cb1cbbf272867e2efda62476f240fb58c64de6b393ae19246741 \
--hash=sha256:58d00498e8933e4a194f3076aee1b4a97dfec1a6da444535755822fe5d8b0b86 \
--hash=sha256:59baa2cb386c4f0b9905bd6eb4c2a79a69a128408fd31d32ca4d7102d4156321 \
--hash=sha256:5a5ed8fde7a1d09376ca0b40e68cd59c69fe23b1f9768bd5824f54681626032a \
--hash=sha256:5b012212e08b8dd5edc78ef54da83dd9892fd9105323b3993eff6bea65dc21d7 \
--hash=sha256:5c3932f4436d1cccb036cb0eaef46e6e2db91035166f1ad6505c3c9d5a635920 \
--hash=sha256:614d0949f4790582d2cc25553abd09dd723025f0c0e7c67376a1d77196743d6e \
--hash=sha256:76341972e1eff8b4bea859f09c0d3e64b96ce931b084f9b9b7db8ef364c30eff \
--hash=sha256:77a2ccbbe917f6710e05ba9adaa25fb5075620bf3ea6fb751997875aff4ae4bd \
--hash=sha256:7995ef305d7165c3f11ae07f2517e5a4f1d5c18da1376a0a9ed496336b69e5f3 \
--hash=sha256:7ce4bfae76319a532a2dc68f82cc32f5676ee792a983187dac07183690e5c66f \
--hash=sha256:7e8eac43dfca5c4cccc6dad9a80504436fca53bb9bc3100a2386d730fbe6b602 \
--hash=sha256:84cf79f0dc8b36ac5da873481716e87aef31fcfa0444f9e1d8b4b2cece142855 \
--hash=sha256:8c7378637d7d88016fa6791c159f698b3d3eed28ebf844ac36b9dc04a14dae18 \
--hash=sha256:8cd666227ef7af430aa5914a9910e0ddd703e75f039cef0825cd0da71b6b711a \
--hash=sha256:906cbf0670286c6e0044156bc7d4af9cbb0ef6db9f73e52c3ec56ba6bdde5336 \
--hash=sha256:9071196d81abc88b3516ac8cdfad32e2b66dd4a5393a8e68a961e9161ddc6239 \
--hash=sha256:9249e3cd978541d665967ac2cb2787fd6a62bddf1e75b3e347a594d7dacf4f74 \
--hash=sha256:984a20b0f62a26f48a3396c72e4bc34c66e356d356bf370053066b3b6d54634a \
--hash=sha256:9be5aafa5736574f8f15f262adc81b2a9869e2cfe9014d52a44633905b40d52c \
--hash=sha256:9c459db21422be75e2809370b829a87eb37f74cd785fc4aa9ea1e5f43b47cda4 \
--hash=sha256:9ccdac7d40688ecb5a3b4a604b8a88c8002e3442d6c60aead1db2a89a041560c \
--hash=sha256:a0e692c683f4df67815a2d258b324e66f4738bd7a96a218c826dce4f4bd05d8f \
--hash=sha256:a5da777e32ffed6f85a7b2b3f7c5cbc88c146bfcd0a1d7baf5fcc6c52ee35dd4 \
--hash=sha256:a64697c641c7b1b2178e573cbc31c7c6684cd56883a478d75143dbb7118036db \
--hash=sha256:ad64688338ed4bc1a6618076ba75fd7194a5f1797ac60b47afe926285adb3166 \
--hash=sha256:bd72e68b06bb1e96913f97dd4901119bc17f39d4586a5adf2d3e47bc2b9d58b5 \
--hash=sha256:c17dfe85494deaeddc5ce251aebd1d60bbe6afc8b62071bb0b469431a000124f \
--hash=sha256:c18684a7f0cc9a3cb60328f496b8e3372def7c5d2df39ac267878b05565aaaae \
--hash=sha256:cc90c0b39b2e3c65ef52c804b72e3c58f8a04ab2a1871272798e5f9572c17d20 \
--hash=sha256:db63bf618e5dea46c07de12e900fe1cdd2541e6dc9dbae772a70b7d4d4765f6a \
--hash=sha256:ea8990436d914540a40ab24b6a77c0969695ed52f4a4874c5137ccf7045a7057 \
--hash=sha256:ecde28a596bead48b0cfd2a1b4416c3d43074c2d785e3a398d7ec1fc4d0f7fbb \
--hash=sha256:f5333311663ea94f75dd408665686aaf426563556bb5283554a3539177e03b8c \
--hash=sha256:fdfef35d751d510fcef5252703621574364fec16418c4a1e5e1055248401054b
# via
# -r requirements_formatting.txt.in
# pyjwt
@@ -276,9 +281,9 @@ graylint==1.1.1 \
--hash=sha256:0fd8e02972ca03d0ef2bf0adea76b5343efcd492d7afb5f658f3e3a724f55a36 \
--hash=sha256:b7e0eab6c159684dbf5ef84e942c3340f6a6549b02a3d11b1a1763cc4f8f0593
# via darker
idna==3.10 \
--hash=sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9 \
--hash=sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3
idna==3.16 \
--hash=sha256:cc246e3a3f89580c3a951b5ad298ca4638078b2cdd4f115654332b5c26daded5 \
--hash=sha256:d7a6da03db833450fca25d2358ac9ff06cd624577a4aea3a596d5c0f77b8e03d
# via
# -r requirements_formatting.txt.in
# requests
@@ -290,9 +295,9 @@ packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
pathspec==1.0.4 \
--hash=sha256:0210e2ae8a21a9137c0d470578cb0e595af87edaa6ebf12ff176f14a02e0e645 \
--hash=sha256:fb6ae2fd4e7c921a165808a552060e722767cfa526f99ca5156ed2ce45a5c723
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
@@ -306,10 +311,12 @@ pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pyjwt==2.12.1 \
--hash=sha256:28ca37c070cad8ba8cd9790cd940535d40274d22f80ab87f3ac6a713e6e8454c \
--hash=sha256:c74a7a2adf861c04d002db713dd85f84beb242228e671280bf709d765b03672b
# via
# -r requirements_formatting.txt.in
# pygithub
pynacl==1.6.2 \
--hash=sha256:018494d6d696ae03c7e656e5e74cdfd8ea1326962cc401bcf018f1ed8436811c \
--hash=sha256:04316d1fc625d860b6c162fff704eb8426b1a8bcd3abacea11142cbd99a6b574 \
@@ -339,9 +346,53 @@ pynacl==1.6.2 \
# via
# -r requirements_formatting.txt.in
# pygithub
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
pytokens==0.4.1 \
--hash=sha256:0fc71786e629cef478cbf29d7ea1923299181d0699dbe7c3c0f4a583811d9fc1 \
--hash=sha256:11edda0942da80ff58c4408407616a310adecae1ddd22eef8c692fe266fa5009 \
--hash=sha256:140709331e846b728475786df8aeb27d24f48cbcf7bcd449f8de75cae7a45083 \
--hash=sha256:24afde1f53d95348b5a0eb19488661147285ca4dd7ed752bbc3e1c6242a304d1 \
--hash=sha256:26cef14744a8385f35d0e095dc8b3a7583f6c953c2e3d269c7f82484bf5ad2de \
--hash=sha256:27b83ad28825978742beef057bfe406ad6ed524b2d28c252c5de7b4a6dd48fa2 \
--hash=sha256:292052fe80923aae2260c073f822ceba21f3872ced9a68bb7953b348e561179a \
--hash=sha256:29d1d8fb1030af4d231789959f21821ab6325e463f0503a61d204343c9b355d1 \
--hash=sha256:2a44ed93ea23415c54f3face3b65ef2b844d96aeb3455b8a69b3df6beab6acc5 \
--hash=sha256:30f51edd9bb7f85c748979384165601d028b84f7bd13fe14d3e065304093916a \
--hash=sha256:34bcc734bd2f2d5fe3b34e7b3c0116bfb2397f2d9666139988e7a3eb5f7400e3 \
--hash=sha256:3ad72b851e781478366288743198101e5eb34a414f1d5627cdd585ca3b25f1db \
--hash=sha256:3f901fe783e06e48e8cbdc82d631fca8f118333798193e026a50ce1b3757ea68 \
--hash=sha256:42f144f3aafa5d92bad964d471a581651e28b24434d184871bd02e3a0d956037 \
--hash=sha256:4a14d5f5fc78ce85e426aa159489e2d5961acf0e47575e08f35584009178e321 \
--hash=sha256:4a58d057208cb9075c144950d789511220b07636dd2e4708d5645d24de666bdc \
--hash=sha256:4e691d7f5186bd2842c14813f79f8884bb03f5995f0575272009982c5ac6c0f7 \
--hash=sha256:5502408cab1cb18e128570f8d598981c68a50d0cbd7c61312a90507cd3a1276f \
--hash=sha256:584c80c24b078eec1e227079d56dc22ff755e0ba8654d8383b2c549107528918 \
--hash=sha256:5ad948d085ed6c16413eb5fec6b3e02fa00dc29a2534f088d3302c47eb59adf9 \
--hash=sha256:670d286910b531c7b7e3c0b453fd8156f250adb140146d234a82219459b9640c \
--hash=sha256:682fa37ff4d8e95f7df6fe6fe6a431e8ed8e788023c6bcc0f0880a12eab80ad1 \
--hash=sha256:6d6c4268598f762bc8e91f5dbf2ab2f61f7b95bdc07953b602db879b3c8c18e1 \
--hash=sha256:79fc6b8699564e1f9b521582c35435f1bd32dd06822322ec44afdeba666d8cb3 \
--hash=sha256:8bdb9d0ce90cbf99c525e75a2fa415144fd570a1ba987380190e8b786bc6ef9b \
--hash=sha256:8fcb9ba3709ff77e77f1c7022ff11d13553f3c30299a9fe246a166903e9091eb \
--hash=sha256:941d4343bf27b605e9213b26bfa1c4bf197c9c599a9627eb7305b0defcfe40c1 \
--hash=sha256:967cf6e3fd4adf7de8fc73cd3043754ae79c36475c1c11d514fc72cf5490094a \
--hash=sha256:970b08dd6b86058b6dc07efe9e98414f5102974716232d10f32ff39701e841c4 \
--hash=sha256:97f50fd18543be72da51dd505e2ed20d2228c74e0464e4262e4899797803d7fa \
--hash=sha256:9bd7d7f544d362576be74f9d5901a22f317efc20046efe2034dced238cbbfe78 \
--hash=sha256:add8bf86b71a5d9fb5b89f023a80b791e04fba57960aa790cc6125f7f1d39dfe \
--hash=sha256:b35d7e5ad269804f6697727702da3c517bb8a5228afa450ab0fa787732055fc9 \
--hash=sha256:b49750419d300e2b5a3813cf229d4e5a4c728dae470bcc89867a9ad6f25a722d \
--hash=sha256:d31b97b3de0f61571a124a00ffe9a81fb9939146c122c11060725bd5aea79975 \
--hash=sha256:d70e77c55ae8380c91c0c18dea05951482e263982911fc7410b1ffd1dadd3440 \
--hash=sha256:d9907d61f15bf7261d7e775bd5d7ee4d2930e04424bab1972591918497623a16 \
--hash=sha256:da5baeaf7116dced9c6bb76dc31ba04a2dc3695f3d9f74741d7910122b456edc \
--hash=sha256:dc74c035f9bfca0255c1af77ddd2d6ae8419012805453e4b0e7513e17904545d \
--hash=sha256:dcafc12c30dbaf1e2af0490978352e0c4041a7cde31f4f81435c2a5e8b9cabb6 \
--hash=sha256:ee44d0f85b803321710f9239f335aafe16553b39106384cef8e6de40cb4ef2f6 \
--hash=sha256:f66a6bbe741bd431f6d741e617e0f39ec7257ca1f89089593479347cc4d13324
# via black
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# -r requirements_formatting.txt.in
# pygithub
@@ -355,9 +406,9 @@ typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.6.3 \
--hash=sha256:1b62b6884944a57dbe321509ab94fd4d3b307075e0c2eae991ac71ee15ad38ed \
--hash=sha256:bf272323e553dfb2e87d9bfd225ca7b0f467b919d7bbd355436d3fd37cb0acd4
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via
# -r requirements_formatting.txt.in
# pygithub
+6 -5
View File
@@ -1,9 +1,10 @@
black~=25.1
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=46.0.5
urllib3>=2.6.3
requests>=2.32.4
idna>=3.7
cryptography>=46.0.7
urllib3>=2.7.0
requests>=2.33.0
idna>=3.15
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.12.1
+7 -1
View File
@@ -6,7 +6,8 @@ set(FEXCORE_BASE_SRCS
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/SpinWaitLock.cpp)
Utils/SpinWaitLock.cpp
Utils/WildcardMatcher.cpp)
if (NOT MINGW)
list(APPEND FEXCORE_BASE_SRCS
@@ -123,6 +124,11 @@ else()
endif()
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# GCC requires libatomic to use 128-bit atomics
list(APPEND LIBS atomic)
endif()
# Generate config
configure_file(${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
${CMAKE_BINARY_DIR}/generated/Config/Config.json)
+19
View File
@@ -260,6 +260,10 @@ struct FEX_PACKED X80SoftFloat {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
// Inf/NaN pass through unchanged in the significand slot.
if (lhs.Top.Exponent == 0x7FFF) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Top.Exponent = 0x3FFF;
Tmp.Top.Sign = lhs.Top.Sign;
@@ -288,6 +292,14 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
// +/-Inf returns +Inf in the exponent slot; NaN propagates.
if (lhs.Top.Exponent == 0x7FFF) {
if ((lhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
X80SoftFloat Result(0, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
return lhs;
}
int32_t TrueExp = lhs.Top.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
@@ -324,6 +336,13 @@ struct FEX_PACKED X80SoftFloat {
#else
extFloat80_t Zero {0, 0};
if (extF80_eq(state, lhs, Zero)) {
// FSCALE(0, +Inf) is 0 * Inf, which is invalid. FSCALE(0, anything
// else) is still 0.
if (rhs.Top.Exponent == 0x7FFF && rhs.Top.Sign == 0 && (rhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
X80SoftFloat QNaN(0, 0x7FFFUL, 0xC000000000000000ULL);
return QNaN;
}
return lhs;
}
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
+4
View File
@@ -16,7 +16,11 @@ struct VectorScalarF64Pair {
#ifdef ARCHITECTURE_arm64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
#ifdef __clang__
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
#else
using VectorRegType = __attribute__((vector_size(16))) uint8_t;
#endif
struct VectorRegPairType {
VectorRegType val[2];
};
+14 -4
View File
@@ -23,6 +23,13 @@
"Enable the code caching subsystem"
]
},
"EnableLazyCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable lazy loading of chunks in code caches"
]
},
"EnableCodeCacheValidation": {
"Type": "bool",
"Default": "false",
@@ -75,7 +82,9 @@
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
"DISABLESSE4A": "disablesse4a",
"ENABLEMOPS": "enablemops",
"DISABLEMOPS": "disablemops"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -99,7 +108,8 @@
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it",
"\t{enable,disable}mops: Will force enable or disable FEAT_MOPS even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -192,7 +202,7 @@
},
"DisableL2Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
@@ -200,7 +210,7 @@
},
"DynamicL1Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
+1 -1
View File
@@ -53,6 +53,6 @@ FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionN
}
bool FEXCore::Context::ContextImpl::IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const {
return Thread->CPUBackend->IsAddressInCodeBuffer(Address);
return Thread->CPUBackend->IsAddressInCodeBuffer(Address) || CodeCache.IsAddressInMappedCodeBuffer(Address);
}
} // namespace FEXCore::Context
+92 -4
View File
@@ -64,6 +64,8 @@ struct CustomIRResult {
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
constexpr static bool BLOCK_DEBUGGING = false;
class CodeCache : public AbstractCodeCache {
public:
CodeCache(ContextImpl&);
@@ -76,11 +78,17 @@ public:
bool IsGeneratingCache = false;
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
FEX_CONFIG_OPT(EnableLazyCodeCaching, ENABLELAZYCODECACHINGWIP);
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
fextl::unique_ptr<MappedCodeCacheFile> LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo&, uint64_t FileStartVA) override;
bool EnableLoadedSection(Core::InternalThreadState*, MappedCodeCacheFile&, const ExecutableFileSectionInfo&) override;
void FinalizeCodePages(MappedCodeCacheFile&, std::span<std::byte> CodeRange) override;
/**
* Performs expensive extra validation on the loaded code cache data.
@@ -112,12 +120,14 @@ public:
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param RelocationOffset Offset to subtract from relocation target offsets
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations,
uint32_t RelocationOffset, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -127,6 +137,7 @@ public:
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState&, uint64_t GuestRIP, uint64_t MaxInst) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
@@ -208,7 +219,7 @@ public:
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -232,6 +243,83 @@ public:
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
// Manual debugging tooling which is useful for developers.
struct TrackingEmpty {
// RIP stepping handling
virtual void AddSingleStepTarget(uint64_t GuestRIP) {}
virtual void AllTargetSingleStep() {}
virtual void RemoveSingleStepTarget(uint64_t GuestRIP) {}
virtual bool IsSingleStepTarget(uint64_t GuestRIP) {
return false;
}
// Watchpoints
virtual void AddWriteWatchPoint(uint64_t Ptr) {}
virtual void AddReadWatchPoint(uint64_t Ptr) {}
virtual bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
virtual bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
};
struct TrackingPossible final : public TrackingEmpty {
void AddSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.emplace(GuestRIP);
}
void RemoveSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.erase(GuestRIP);
}
void AllTargetSingleStep() override {
SingleStepEverything = true;
}
bool IsSingleStepTarget(uint64_t GuestRIP) override {
return SingleStepEverything || SingleStepTargets.contains(GuestRIP);
}
void AddWriteWatchPoint(uint64_t Ptr) override {
WatchWriteTargets.emplace(Ptr);
}
void AddReadWatchPoint(uint64_t Ptr) override {
WatchReadTargets.emplace(Ptr);
}
bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchWriteTargets, Ptr, Size);
}
bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchReadTargets, Ptr, Size);
}
private:
bool SingleStepEverything {};
fextl::set<uint64_t> SingleStepTargets {};
fextl::set<uint64_t> WatchWriteTargets {};
fextl::set<uint64_t> WatchReadTargets {};
static bool ContainsRange(const fextl::set<uint64_t>& Set, uint64_t Ptr, size_t Size) {
for (auto it = Set.lower_bound(Ptr); it != Set.end(); --it) {
auto Watch = *it;
if (Watch < Ptr) {
break;
}
if (Watch >= Ptr && Watch < (Ptr + Size)) {
return true;
}
}
return false;
}
};
using TrackingStructure = std::conditional<BLOCK_DEBUGGING, TrackingPossible, TrackingEmpty>::type;
TrackingStructure BlockDebuggerTracker {};
public:
struct {
uint64_t VirtualMemSize {1ULL << 36};
@@ -262,7 +350,7 @@ public:
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::Utils::WritePriorityMutex::Mutex CodeInvalidationMutex {};
uint32_t StrictSplitLockMutex {};
@@ -448,8 +448,10 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
return;
}
if ((Constant >> 32) == 0) {
if ((Constant >> 32) == 0 && !NOPPad) {
// If the upper 32-bits is all zero, we can now switch to a 32-bit move.
// NOTE: The NOP padding code does not appropriately adjust to this yet,
// so we skip this optimization in that case
s = ARMEmitter::Size::i32Bit;
Is64Bit = false;
Segments = std::min(Segments, 2);
@@ -584,8 +586,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
}};
for (auto& RegPair : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, -16);
for (const auto& [rt, rt2] : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(rt, rt2, ARMEmitter::Reg::rsp, -16);
}
// Additionally we need to store the lower 64bits of v8-v15
@@ -602,9 +604,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We just saved x19 so it is safe
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r19, ARMEmitter::Reg::rsp, 0);
for (auto& RegQuad : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::r19, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::r19, 32);
}
}
@@ -614,9 +615,8 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
}};
for (auto& RegQuad : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::rsp, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::rsp, 32);
}
constexpr static std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
@@ -628,8 +628,8 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
}};
for (auto& RegPair : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, 16);
for (const auto& [rt, rt2] : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(rt, rt2, ARMEmitter::Reg::rsp, 16);
}
}
@@ -659,7 +659,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
#endif
if (SetPredRegs && (EmitterCTX->HostFeatures.SupportsSVE256 || EmitterCTX->HostFeatures.SupportsSVE128)) {
if (SetPredRegs && EmitterCTX->HostFeatures.SupportsSVE()) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
@@ -677,7 +677,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask, bool NZCV) {
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
@@ -698,7 +698,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
#endif
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're spilling, we need to spill NZCV since it
// is always static and almost certainly clobbered by the subsequent code.
//
@@ -710,25 +710,25 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
unsigned PFAFSpillMask = Options.GPRSpillMask & PFAFMask;
Options.GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRSpillMask) && ((1U << Reg2.Idx()) & GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg1.Idx()) & GPRSpillMask)) {
str(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg2.Idx()) & GPRSpillMask)) {
str(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRSpillMask) && ((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg1.Idx()) & Options.GPRSpillMask)) {
str(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
str(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFSpillMask) {
if (Options.NZCV && PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
@@ -737,21 +737,21 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
stp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), PFOffset);
}
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TmpReg);
}
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
if (Options.GPRSpillMask && Options.FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -764,12 +764,12 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRSpillMask) && ((1U << Reg2.Idx()) & FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRSpillMask) && ((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -777,8 +777,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2, bool NZCV) {
void Arm64Emitter::FillStaticRegs(FillStaticRegOptions Options) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
@@ -789,20 +788,21 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
LOGMAN_THROW_A_FMT(Options.GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = Options.GPRFillMask;
if (!Options.OptionalReg.has_value()) {
Options.OptionalReg = FindTempReg(&TempGPRFillMask);
}
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
if (!Options.OptionalReg2.has_value()) {
Options.OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
LOGMAN_THROW_A_FMT(Options.OptionalReg.has_value() && Options.OptionalReg2.has_value(), "Didn't have an SRA register to use as a "
"temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
auto TmpReg = *Options.OptionalReg;
auto TmpReg2 = *Options.OptionalReg2;
#ifdef ARCHITECTURE_arm64ec
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
@@ -812,7 +812,7 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
@@ -822,23 +822,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
}
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
FillSpecialRegs(TmpReg, TmpReg2, true, Options.FPRs);
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TmpReg);
}
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
if (Options.GPRFillMask && Options.FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -851,12 +851,12 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRFillMask) && ((1U << Reg2.Idx()) & FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRFillMask) && ((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -865,23 +865,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
uint32_t PFAFFillMask = GPRFillMask & PFAFMask;
GPRFillMask &= ~PFAFMask;
uint32_t PFAFFillMask = Options.GPRFillMask & PFAFMask;
Options.GPRFillMask &= ~PFAFMask;
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRFillMask) && ((1U << Reg2.Idx()) & GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg1.Idx()) & GPRFillMask) {
ldr(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg2.Idx()) & GPRFillMask) {
ldr(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRFillMask) && ((1U << Reg2.Idx()) & Options.GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg1.Idx()) & Options.GPRFillMask) {
ldr(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg2.Idx()) & Options.GPRFillMask) {
ldr(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFFillMask) {
if (Options.NZCV && PFAFFillMask) {
LOGMAN_THROW_A_FMT(PFAFFillMask == PFAFMask, "PF/AF not filled together");
ldp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw));
@@ -1056,7 +1056,10 @@ size_t Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, boo
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
SpillStaticRegs(TmpReg, {
.GPRSpillMask = PreserveSRAMask,
.FPRSpillMask = PreserveSRAFPRMask,
});
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -1103,7 +1106,11 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
}
// Fill the static registers.
FillStaticRegs(FPRs, PreserveSRAMask, PreserveSRAFPRMask);
FillStaticRegs({
.GPRFillMask = PreserveSRAMask,
.FPRFillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
// Pop the vector registers.
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
@@ -135,10 +135,35 @@ protected:
// Returning REG_INVALID if there was no mapping.
FEXCore::X86State::X86Reg GetX86RegRelationToARMReg(ARMEmitter::Register Reg);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U, bool NZCV = true);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt, bool NZCV = true);
struct SpillStaticRegOptions final {
uint32_t GPRSpillMask {~0U};
uint32_t FPRSpillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
struct FillStaticRegOptions final {
std::optional<ARMEmitter::Register> OptionalReg {std::nullopt};
std::optional<ARMEmitter::Register> OptionalReg2 {std::nullopt};
uint32_t GPRFillMask {~0U};
uint32_t FPRFillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
void SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options);
void FillStaticRegs(FillStaticRegOptions Options);
void SpillStaticRegs(ARMEmitter::Register TmpReg) {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
SpillStaticRegs(TmpReg, {});
}
void FillStaticRegs() {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
FillStaticRegs({});
}
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -178,7 +203,9 @@ protected:
if (SupportsPreserveAllABI) {
return SpillForPreserveAllABICall(TmpReg, FPRs);
} else {
SpillStaticRegs(TmpReg, FPRs);
SpillStaticRegs(TmpReg, {
.FPRs = FPRs,
});
return PushDynamicRegs(TmpReg);
}
}
@@ -188,7 +215,7 @@ protected:
FillForPreserveAllABICall(FPRs);
} else {
PopDynamicRegs();
FillStaticRegs(FPRs);
FillStaticRegs({.FPRs = FPRs});
}
}
+5 -1
View File
@@ -12,7 +12,6 @@
#include <cstdint>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -44,6 +43,8 @@ namespace CPU {
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB_UPPER
{0x0706'0504'0302'0100ULL, 0x1716'1514'1312'1110ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP
{0x0F0E'0D0C'0B0A'0908ULL, 0x1F1E'1D1C'1B1A'1918ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_ONE
{0xD49A'784B'CD1B'8AFEULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_LOG2_10
{0xB8AA'3B29'5C17'F0BCULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_LOG2_E
@@ -363,6 +364,9 @@ namespace CPU {
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<void*>(Ptr), Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
+116 -4
View File
@@ -92,6 +92,7 @@ namespace ProductNames {
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_ORYON_3[] = "Oryon-3";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
@@ -141,7 +142,7 @@ constexpr uint32_t FAMILY_IDENTIFIER = GenerateFamily(CPUFamily {
#endif
#ifdef ARCHITECTURE_arm64
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -186,8 +187,9 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 67> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 68> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x002, 1, ProductNames::ARM_ORYON_3}, // Qualcomm Oryon-3
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x039, 1, ProductNames::ARM_Avalanche_M2Max}, // Apple Avalanche (M2 Max)
@@ -408,7 +410,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
}
#else
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
return 0;
}
@@ -756,6 +758,95 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 29) | // Arch capabilities - Speculative side channel mitigations
(0 << 30) | // Arch capabilities - MSR module specific
(0 << 31); // SSBD - Speculative Store Bypass Disable
} else if (Leaf == 1) {
Res.eax = (0U << 0) | // SHA512
(0U << 1) | // SM3
(0U << 2) | // SM4
(0U << 3) | // RAO_INT
(0U << 4) | // AVX_VNNI
(0U << 5) | // AVX512_BF16
(0U << 6) | // LASS (Linear Address Space Separation)
(0U << 7) | // CMPCCXADD
(0U << 8) | // ARCH_PERFMON_EXT
(0U << 9) | // Reserved
(0U << 10) | // FAST_REP_MOVSB
(0U << 11) | // FAST_REP_STOSB
(0U << 12) | // FAST_REP_CMPSB_SCASB
(0U << 13) | // Reserved
(0U << 14) | // Reserved
(0U << 15) | // Reserved
(0U << 16) | // Reserved
(0U << 17) | // FRED (Flexible Return and Event Delivery)
(0U << 18) | // LKGS (Load into Kernel GS Base)
(0U << 19) | // WRMSRNS
(0U << 20) | // NMI_SRC
(0U << 21) | // AMX_FP16
(0U << 22) | // HRESET
(0U << 23) | // AVX_IFMA
(0U << 24) | // Reserved
(0U << 25) | // Reserved
(0U << 26) | // LAM (Linear Address Masking)
(0U << 27) | // MSRLIST
(0U << 28) | // Reserved
(0U << 29) | // Reserved
(0U << 30) | // INVD_DISABLE_POST_BIOS_DONE
(0U << 31); // MOVRS
// Bits 4-31 currently reserved.
Res.ebx = (0U << 0) | // PPIN
(0U << 1) | // PBNDKB
(0U << 2) | // Reserved
(0U << 3); // CPUIDMAXVAL_LIM_RMV
// Bits 6-31 also reserved.
Res.ecx = (0U << 0) | // RDT_M_ASYM
(0U << 1) | // RDT_A_ASYM
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // Reserved
(0U << 5); // MSR_IMM
// Bits 25-31 also reserved.
Res.edx = (0U << 0) | // Reserved
(0U << 1) | // Reserved
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // AVX_VNNI_INT8
(0U << 5) | // AVX_NE_CONVERT
(0U << 6) | // Reserved
(0U << 7) | // Reserved
(0U << 8) | // AMX_COMPLEX
(0U << 9) | // Reserved
(0U << 10) | // AVX_VNNI_INT16
(0U << 11) | // Reserved
(0U << 12) | // Reserved
(0U << 13) | // UTMR (User-timer events)
(0U << 14) | // PREFETCHI
(0U << 15) | // USER_MSR
(0U << 16) | // Reserved
(0U << 17) | // UIRET_UIF
(0U << 18) | // CET_SSS
(0U << 19) | // AVX10
(0U << 20) | // Reserved
(0U << 21) | // APX_F
(0U << 22) | // SEC-TEE_ATTESTATION
(0U << 23) | // MWAIT
(0U << 24); // SLSM (Static LSM)
} else if (Leaf == 2) {
// All bits are reserved except for EDX
Res.eax = 0;
Res.ebx = 0;
Res.ecx = 0;
// Bits 8-31 are reserved.
Res.edx = (0U << 0) | // PSFD
(0U << 1) | // IPRED_CTRL
(0U << 2) | // RRSBA_CTRL
(0U << 3) | // DDPD_U
(0U << 4) | // BHI_CTRL
(0U << 5) | // MCDT_NO
(0U << 6) | // UC_LOCK_DISABLE
(0U << 7); // MONITOR_MITG_NO
}
return Res;
@@ -813,7 +904,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
uint64_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1U << CTX->Config.TSCScale;
@@ -834,6 +925,27 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) const {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_24h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
// EAX indicates the maximum number of subleaves.
Res.eax = 0;
// Bits 19-31 reserved
// NOTE: We return all zero here until we have a CPU with AVX10
// even if some of the fields otherwise have fixed values.
Res.ebx = (0U << 0) | // (bits 0-7 specify the vector ISA version)
(0U << 16); // Defined as always 0b111
// All bits reserved
Res.ecx = 0;
Res.edx = 0;
}
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
+84 -2
View File
@@ -14,7 +14,7 @@ namespace Context {
class ContextImpl;
}
uint32_t GetCycleCounterFrequency();
uint64_t GetCycleCounterFrequency();
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
@@ -176,6 +176,7 @@ private:
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_24h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf) const;
@@ -200,7 +201,7 @@ private:
void SetupHostHybridFlag();
void SetupFeatures();
static constexpr size_t PRIMARY_FUNCTION_COUNT = 27;
static constexpr size_t PRIMARY_FUNCTION_COUNT = 37;
static constexpr size_t HYPERVISOR_FUNCTION_COUNT = 2;
static constexpr size_t EXTENDED_FUNCTION_COUNT = 32;
static constexpr std::array<FunctionHandler, PRIMARY_FUNCTION_COUNT> Primary = {
@@ -268,7 +269,48 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
&CPUIDEmu::Function_1Ah,
// 0x1B: PCONFIG info
&CPUIDEmu::Function_Reserved,
// 0x1C: Last Branch Records (LBR) info
&CPUIDEmu::Function_Reserved,
// 0x1D: Tile info
&CPUIDEmu::Function_Reserved,
// 0x1E: TMUL info
&CPUIDEmu::Function_Reserved,
// 0x1F: V2 Extended topology
&CPUIDEmu::Function_Reserved,
// 0x20: Processor History Reset info
&CPUIDEmu::Function_Reserved,
// 0x21: Unimplemented
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Architectural Performance Monitoring Extended
&CPUIDEmu::Function_Reserved,
// 0x24: Converged Vector ISA
&CPUIDEmu::Function_24h,
#else
// 0x1A: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1B: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1C: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1D: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1E: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1F: Reserved
&CPUIDEmu::Function_Reserved,
// 0x20: Reserved
&CPUIDEmu::Function_Reserved,
// 0x21: Reserved
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Reserved
&CPUIDEmu::Function_Reserved,
// 0x24: Reserved
&CPUIDEmu::Function_Reserved,
#endif
};
@@ -340,9 +382,49 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: PCONFIG info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Last Branch Records (LBR) info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Tile info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: TMUL info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: V2 Extended topology
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Processor History Reset info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Unimplemented/Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Architectural Performance Monitoring Extended
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Converged Vector ISA
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 0x1A: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
}};
+403 -186
View File
@@ -1,5 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include "FEXCore/Utils/LogManager.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Utils/TypeDefines.h"
#include "FEXCore/fextl/memory.h"
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SpinWaitLock.h>
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
@@ -16,10 +21,14 @@
#include <FEXHeaderUtils/Filesystem.h>
#include <algorithm>
#include <git_version.h>
#include <span>
#include <xxhash.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <fstream>
namespace FEXCore {
@@ -32,6 +41,38 @@ ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
MappedCodeCacheFile::~MappedCodeCacheFile() {
if (CacheManager) {
CacheManager->UnregisterMappedCodeBuffer(*this);
}
#ifndef _WIN32
if (!CodeBuffer.empty()) {
FEXCore::Allocator::munmap(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#endif
}
void AbstractCodeCache::RegisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
MappedCodeBuffers.push_back(Code.CodeBuffer);
// Unregister on destruction of Code
Code.CacheManager = this;
}
void AbstractCodeCache::UnregisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
std::erase_if(MappedCodeBuffers, [&](const auto& Elem) { return Elem.data() == Code.CodeBuffer.data(); });
}
bool AbstractCodeCache::IsAddressInMappedCodeBuffer(uintptr_t Address) const {
for (const auto& Range : MappedCodeBuffers) {
auto Start = reinterpret_cast<uintptr_t>(Range.data());
if (Address >= Start && Address < Start + Range.size_bytes()) {
return true;
}
}
return false;
}
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
@@ -233,7 +274,10 @@ uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
struct CodeCacheHeader {
std::array<char, 4> Magic = ExpectedMagic;
uint32_t FormatVersion = 1;
// Version history:
// 1: Initial version
// 2: Padding code buffer data to enable direct mapping
uint32_t FormatVersion = 2;
uint8_t FEXVersion[20] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
@@ -260,7 +304,7 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
std::ranges::copy(GIT_HASH, header.FEXVersion);
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.CodeBufferSize = FEXCore::AlignUp(CTX.LatestOffset, Utils::FEX_PAGE_SIZE);
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
@@ -300,21 +344,25 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
{
auto AlignedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, AlignedSize);
lseek(fd, AlignedSize, SEEK_SET);
}
// Dump the host code (relocated for position-independent serialization)
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, 0, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Pad to next page in file for mmap
{
auto PaddedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, PaddedSize);
lseek(fd, PaddedSize, SEEK_SET);
}
// Dump code pages
static_assert(OrderedContainer<decltype(LookupCache.CodePages)>, "Non-deterministic data source");
@@ -332,178 +380,6 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
return true;
}
bool CodeCache::LoadData(Core::InternalThreadState* Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
// Read file header
CodeCacheHeader header {};
::memcpy(&header, MappedCacheFile, sizeof(header));
MappedCacheFile += sizeof(header);
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", header.NumBlocks, BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (!ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return false;
}
if (!ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return false;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return false;
}
// Read guest<->host block mappings
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(header.NumBlocks);
{
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, MappedCacheFile, sizeof(BlockPtr.first));
MappedCacheFile += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, MappedCacheFile, sizeof(BlockPtr.second.HostCode));
MappedCacheFile += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, MappedCacheFile, sizeof(NumGuestPages));
MappedCacheFile += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), MappedCacheFile, std::span {BlockPtr.second.CodePages}.size_bytes());
MappedCacheFile += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Consistency check: VMA regions at the top and end should belong to the same file
auto [min_val, max_val] = ranges::minmax_element(BlockList, std::less {}, &decltype(BlockList)::value_type::first);
auto MinBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, min_val->first + BinarySection.FileStartVA);
auto MaxBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, max_val->first + BinarySection.FileStartVA);
if (&MinBound->FileInfo != &BinarySection.FileInfo || &MaxBound->FileInfo != &BinarySection.FileInfo) {
ERROR_AND_DIE_FMT("Cached blocks offsets {:#x}-{:#x} out of bounds for guest library {} ({:016x} @ {:#x}) while trying to load "
"section {:#x}-{:#x}!",
min_val->first, max_val->first, BinarySection.FileInfo.Filename, BinarySection.FileInfo.FileId,
BinarySection.FileStartVA, BinarySection.BeginVA, BinarySection.EndVA);
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
if (BlockList.empty()) {
// Not an error since there is just no data to load
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
}
// Read relocations
fextl::vector<FEXCore::CPU::Relocation> Relocations(header.NumRelocations, FEXCore::CPU::Relocation::Default());
::memcpy(Relocations.data(), MappedCacheFile, Relocations.size() * sizeof(Relocations[0]));
MappedCacheFile += Relocations.size() * sizeof(Relocations[0]);
// Pad to next page in file, which contains CodeBuffer data
MappedCacheFile = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(MappedCacheFile), Utils::FEX_PAGE_SIZE));
// Prepare CodeBuffer: Page aligned and big enough to hold all cached data
auto Lock = std::unique_lock {CTX.CodeBufferWriteMutex};
if (Thread) {
if (auto Prev = Thread->CPUBackend->CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = Thread->LookupCache->AcquireWriteLock();
Thread->LookupCache->ChangeGuestToHostMapping(*Prev, *CTX.GetLatest()->LookupCache, lk);
}
}
auto CodeBuffer = CTX.GetLatest();
LOGMAN_THROW_A_FMT(reinterpret_cast<uintptr_t>(CodeBuffer->Ptr) % 0x1000 == 0, "Expected CodeBuffer base to be page-aligned");
const auto Delta = AlignUp(CTX.LatestOffset, 0x1000) - CTX.LatestOffset;
CTX.LatestOffset += Delta;
while (CTX.LatestOffset + header.CodeBufferSize > CodeBuffer->UsableSize()) {
if (Thread) {
CTX.ClearCodeCache(Thread);
CodeBuffer = CTX.GetLatest();
LogMan::Msg::IFmt("Increased code buffer size to {} MiB for cache load", CodeBuffer->AllocatedSize / 1024 / 1024);
} else {
ERROR_AND_DIE_FMT("Cannot extend codebuffer without thread!");
}
}
// Read CodeBuffer data from file. Make sure the destination is page-aligned.
// TODO: Only load the data needed for the selected section
auto CodeBufferRange =
std::as_writable_bytes(std::span {CodeBuffer->Ptr, CodeBuffer->UsableSize()}).subspan(CTX.LatestOffset, header.CodeBufferSize);
::memcpy(CodeBufferRange.data(), MappedCacheFile, header.CodeBufferSize);
MappedCacheFile += header.CodeBufferSize;
CTX.LatestOffset += header.CodeBufferSize;
// Apply FEX relocations
auto Ret = ApplyCodeRelocations(BinarySection.FileStartVA, CodeBufferRange, Relocations, false);
LOGMAN_THROW_A_FMT(Ret == true, "Failed to apply code cache relocations");
{
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
// Register blocks to LookupCache
for (auto& [Guest, Host] : BlockList) {
for (auto& CodePage : Host.CodePages) {
CodePage += BinarySection.FileStartVA;
}
auto HostCode = reinterpret_cast<void*>(Host.HostCode + reinterpret_cast<uintptr_t>(CodeBufferRange.data()));
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Host.CodePages), HostCode, WriteLock);
}
// Register loaded code ranges
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < header.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, MappedCacheFile, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
MappedCacheFile += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, MappedCacheFile, sizeof(NumEntrypoints));
MappedCacheFile += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), MappedCacheFile, NumEntrypoints * sizeof(Entrypoints[0]));
MappedCacheFile += NumEntrypoints * sizeof(Entrypoints[0]);
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, CodeBufferRange);
}
return true;
}
void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<uint64_t> GuestBlocks, const fextl::set<uint64_t>& HostBlocks,
std::span<std::byte> CachedCode) {
LOGMAN_THROW_A_FMT(!HostBlocks.empty(), "Tried to validate without any host blocks");
@@ -558,7 +434,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
NewRelocations.erase(std::remove_if(NewRelocations.begin(), NewRelocations.end(), [](const CPU::Relocation& Reloc) {
return Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL && Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
}));
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, false);
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, 0, false);
if (ValidationCTX->LatestOffset <= CodeBufferRangeRef.size()) {
// Reference compilation produced fewer bytes than our cache, so validation is going to fail.
@@ -616,15 +492,17 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
ValidationThread->LookupCache->ClearCache(ValidationThread->LookupCache->AcquireWriteLock());
ValidationCTX->LatestOffset = 0;
LogMan::Msg::IFmt("\tSuccessfully validated cache");
LogMan::Msg::IFmt(" successfully validated cache");
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
std::span<const FEXCore::CPU::Relocation> EntryRelocations, uint32_t RelocationOffset, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
LOGMAN_THROW_A_FMT(Reloc.Header.Offset >= RelocationOffset, "Invalid relocation offset");
LOGMAN_THROW_A_FMT(Reloc.Header.Offset - RelocationOffset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Header.Offset - RelocationOffset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
@@ -663,4 +541,343 @@ bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> C
return true;
}
fextl::unique_ptr<MappedCodeCacheFile>
CodeCache::LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo& FileInfo, uint64_t FileStartVA) {
if (!EnableCodeCaching) {
return nullptr;
}
FEXCORE_PROFILE_SCOPED("LoadCache");
// Read file header
CodeCacheHeader header {};
::memcpy(&header, CacheFile.data(), sizeof(header));
if (!std::ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return nullptr;
}
if (!std::ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return nullptr;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return nullptr;
}
// Skip over BlockEntry data since it won't be used until EnableLoadedSection
// TODO: Store direct offset to relocations in the header
auto* BlockListStart = CacheFile.data() + sizeof(header);
auto* Cursor = BlockListStart;
for (uint32_t i = 0; i < header.NumBlocks; ++i) {
Cursor += sizeof(uint64_t); // guest address
Cursor += sizeof(uint64_t); // host code address
uint64_t NumGuestCodePages;
::memcpy(&NumGuestCodePages, Cursor, sizeof(NumGuestCodePages));
Cursor += sizeof(NumGuestCodePages);
Cursor += NumGuestCodePages * sizeof(uint64_t);
}
auto Relocations = std::span {reinterpret_cast<const FEXCore::CPU::Relocation*>(Cursor), header.NumRelocations};
Cursor += Relocations.size_bytes();
// Pad to next page to get the code buffer data
Cursor = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(Cursor), Utils::FEX_PAGE_SIZE));
auto CodeDataInFile = std::span {Cursor, header.CodeBufferSize};
#ifndef _WIN32
// Allocate target memory for post-relocation code. This is PROT_NONE until
// the first execution, so that contents can be lazily populated in a
// frontend-provided segfault handler.
void* CodeBufferAllocation = Allocator::mmap(nullptr, header.CodeBufferSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (CodeBufferAllocation == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to reserve target memory for code cache");
return nullptr;
}
auto CodeBuffer = std::span {static_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#else
// TODO: Implement lazy mapping on Windows
auto CodeBuffer = CodeDataInFile;
#endif
// Group relocations by page
size_t NumPages = header.CodeBufferSize / Utils::FEX_PAGE_SIZE;
fextl::vector<MappedCodeCacheFile::PageRelocationRange> PageRelocationRanges(NumPages, {0, 0});
auto RelocBaseOffset = std::as_bytes(Relocations).data() - CacheFile.data();
auto RelocIt = Relocations.begin();
for (size_t Page = 0; Page < NumPages; ++Page) {
auto EndRelocIt = std::upper_bound(RelocIt, Relocations.end(), Page,
[](auto& Page, auto& Reloc) { return Page < Reloc.Header.Offset / Utils::FEX_PAGE_SIZE; });
PageRelocationRanges.at(Page) = {static_cast<uint32_t>(RelocBaseOffset + (RelocIt - Relocations.begin()) * sizeof(CPU::Relocation)),
static_cast<uint32_t>(EndRelocIt - RelocIt)};
RelocIt = EndRelocIt;
}
auto Storage = FEXCore::Allocator::aligned_alloc(alignof(MappedCodeCacheFile), sizeof(MappedCodeCacheFile));
return fextl::unique_ptr<MappedCodeCacheFile>(
new (Storage) MappedCodeCacheFile {this, CacheFile, CodeDataInFile, CodeBuffer, BlockListStart, header.NumBlocks, header.NumCodePages,
std::move(PageRelocationRanges), fextl::vector<bool>(NumPages), FileStartVA});
}
bool CodeCache::EnableLoadedSection(Core::InternalThreadState* Thread, MappedCodeCacheFile& Code, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
FEXCORE_PROFILE_SCOPED("EnableLoadedSection");
// Read block list from cache file
// TODO: Store section-ized BlockLists in cache file
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(Code.NumBlocks);
{
auto* Cursor = Code.BlockListInFile;
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, Cursor, sizeof(BlockPtr.first));
Cursor += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, Cursor, sizeof(BlockPtr.second.HostCode));
Cursor += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, Cursor, sizeof(NumGuestPages));
Cursor += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), Cursor, std::span {BlockPtr.second.CodePages}.size_bytes());
Cursor += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
if (begin == end) {
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", BlockList.size(), BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (EnableLazyCodeCaching) {
LogMan::Msg::IFmt(" lazy mapping: base={:#14x} -> host={}; cache_source={}", BinarySection.FileStartVA,
fmt::ptr(Code.CodeBuffer.data()), fmt::ptr(Code.MappedFile.data()));
}
// Register blocks to LookupCache.
// The host addresses will point into the protected code buffer, so that FEX
// can lazily apply relocations on first execution of each page.
auto CodeBuffer = CTX.GetLatest();
{
FEXCORE_PROFILE_SCOPED("Decode");
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
for (auto& [Guest, Block] : BlockList) {
for (auto& CodePage : Block.CodePages) {
CodePage += BinarySection.FileStartVA;
}
LOGMAN_THROW_A_FMT(Block.HostCode < Code.CodeBuffer.size_bytes(), "Host offset {:#x} out of range ({:#x})", Block.HostCode,
Code.CodeBuffer.size_bytes());
auto HostCode = &Code.CodeBuffer[Block.HostCode];
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Block.CodePages), HostCode, WriteLock);
}
// Guest code pages
auto* Cursor = Code.CodeBufferInFile.data() + Code.CodeBufferInFile.size_bytes();
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < Code.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, Cursor, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
Cursor += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, Cursor, sizeof(NumEntrypoints));
Cursor += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), Cursor, std::span {Entrypoints}.size_bytes());
Cursor += std::span {Entrypoints}.size_bytes();
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
#ifndef _WIN32
if (!EnableLazyCodeCaching || EnableCodeCacheValidation) {
#else
// TODO: Implement lazy mapping on Windows
if (true) {
#endif
auto Range = SelectCodeRangeToFinalize(Code, 0, Code.CodeBuffer.size_bytes() / Utils::FEX_PAGE_SIZE);
FinalizeCodePages(Code, Range);
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, Code.CodeBuffer);
}
return true;
}
} // namespace FEXCore::Context
namespace FEXCore {
static std::span<CPU::Relocation> SpanPageRelocations(const MappedCodeCacheFile& Code, size_t PageIndex) {
auto [Offset, Count] = Code.PageRelocationRanges.at(PageIndex);
return std::span {reinterpret_cast<FEXCore::CPU::Relocation*>(Code.MappedFile.data() + Offset), Count};
}
std::span<std::byte> AbstractCodeCache::SelectCodeRangeToFinalize(MappedCodeCacheFile& Code, size_t StartPage, size_t EndPage) {
// First, check if we were racing another thread in loading this range
if (std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, false) == Code.LoadedPages.begin() + EndPage) {
return {};
}
LOGMAN_THROW_A_FMT(StartPage < EndPage, "Invalid page range [{}, {})", StartPage, EndPage);
LOGMAN_THROW_A_FMT(EndPage <= Code.NumPages(), "End page {} out of range ({})", EndPage, Code.NumPages());
// Include any pages that have relocations or block link records crossing
// into the current page range. This ensures we don't attempt to finalize
// any page twice, partially apply FEX relocations, or trigger page loads
// during block linking.
while (EndPage < Code.NumPages()) {
auto PageRelocs = SpanPageRelocations(Code, EndPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > EndPage * Utils::FEX_PAGE_SIZE) {
++EndPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, EndPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < EndPage * Utils::FEX_PAGE_SIZE + 0x18) {
++EndPage;
continue;
}
}
break;
};
while (StartPage != 0) {
auto PageRelocs = SpanPageRelocations(Code, StartPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > StartPage * Utils::FEX_PAGE_SIZE) {
--StartPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, StartPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < StartPage * Utils::FEX_PAGE_SIZE + 0x18) {
--StartPage;
continue;
}
}
break;
};
return Code.CodeBuffer.subspan(StartPage * Utils::FEX_PAGE_SIZE, (EndPage - StartPage) * Utils::FEX_PAGE_SIZE);
}
} // namespace FEXCore
namespace FEXCore::Context {
void CodeCache::FinalizeCodePages(MappedCodeCacheFile& Code, std::span<std::byte> CodeRange) {
const size_t StartOffset = CodeRange.data() - Code.CodeBuffer.data();
const auto StartPage = StartOffset / Utils::FEX_PAGE_SIZE;
const auto EndPage = StartPage + CodeRange.size_bytes() / Utils::FEX_PAGE_SIZE;
const size_t Size = CodeRange.size_bytes();
// None of the selected pages should be loaded at all; otherwise, SelectCodeRangeToFinalize returned inconsistent ranges
LOGMAN_THROW_A_FMT(std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, true) == Code.LoadedPages.begin() + EndPage,
"Inconsistent page load state");
FEXCORE_PROFILE_SCOPED("FinalizeCodePages");
#ifndef _WIN32
// Atomicity is critical when making the finalized code data visible.
// We ensure this by remapping a temporary buffer onto the PROT_NONE
// placeholder page in CodeBuffer. Some constraints to keep in mind are:
// 1. Pages can't be write-only (readability is implicitly added), so
// we can't change CodeBuffer from PROT_NONE to PROT_WRITE even for just
// a short duration
// 2. Naive mremap from CodeBufferInFile to CodeBuffer would leave a gap in
// the former, which would make cleanup overly complicated
//
// Due to (1), we can't apply relocations in place (CodeBufferInFile); at
// least a secondary buffer is needed for execution (CodeBuffer).
// Due to (2), a third buffer is temporarily allocated here and freed on
// completion. The final code data is computed here and then the memory
// is remapped onto CodeBuffer.
auto* Staging = reinterpret_cast<std::byte*>(Allocator::VirtualAlloc(nullptr, Size, true));
if (!Staging) {
ERROR_AND_DIE_FMT("Failed to allocate {} bytes of staging memory for code-cache finalization", Size);
}
// Copy code from the cache file to the staging buffer
memcpy(Staging, Code.CodeBufferInFile.data() + StartOffset, Size);
// Apply relocations
auto StagingSpan = std::span {Staging, Size};
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, StagingSpan, PageRelocations, static_cast<uint32_t>(StartOffset), false);
Code.LoadedPages[i] = true;
}
// Atomically make the finalized code data visible by remapping the staging
// buffer onto the requested CodeBuffer window. MREMAP_DONTUNMAP is used to
// leave the old VA range reserved so that we can cleanly deallocate it
// through Allocator.
void* RemapResult = ::mremap(Staging, Size, Size, MREMAP_FIXED | MREMAP_MAYMOVE | MREMAP_DONTUNMAP, CodeRange.data());
if (RemapResult == MAP_FAILED) {
ERROR_AND_DIE_FMT("{}: mremap failed: {}", __FUNCTION__, errno);
}
Allocator::VirtualFree(Staging, Size);
// Release resident file pages that will no longer be needed. The VA range is left allocated to allow cleanup with a single VirtualFree.
Allocator::VirtualDontNeed(Code.CodeBufferInFile.data() + StartOffset, Size);
#else
// TODO: Implement lazy mapping on Windows
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, Code.CodeBuffer, PageRelocations, 0, false);
Code.LoadedPages[i] = true;
}
#endif
ARMEmitter::Emitter::ClearICache(CodeRange.data(), Size);
}
} // namespace FEXCore::Context
+35 -6
View File
@@ -30,7 +30,7 @@ $end_info$
#include "Interface/IR/RegisterAllocationData.h"
#include "Utils/Allocator.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include "Utils/variable_length_integer.h"
#include <FEXCore/Config/Config.h>
@@ -358,6 +358,16 @@ bool ContextImpl::InitCore() {
Config.NeedsPendingInterruptFaultCheck = true;
}
if constexpr (BLOCK_DEBUGGING) {
// If the developer wants to do any single-stepping points or watch points.
// Add them here.
//
// eg:
// BlockDebuggerTracker.AllTargetSingleStep();
// BlockDebuggerTracker.AddSingleStepTarget(0x14000'0000ULL);
// BlockDebuggerTracker.AddWriteWatchPoint(0x420BA5ED);
}
return true;
}
@@ -376,7 +386,7 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this, Thread);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
@@ -456,7 +466,6 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
if (Config.StrictInProcessSplitLocks) {
FEXCore::Utils::SpinWaitLock::unlock(&StrictSplitLockMutex);
}
return;
}
}
@@ -502,11 +511,14 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
bool ContextImpl::CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) {
return Thread.FrontendDecoder->CheckIfCacheable(Thread, reinterpret_cast<const uint8_t*>(GuestRIP), GuestRIP, MaxInst);
}
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
uint64_t TotalInstructions {0};
@@ -638,6 +650,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->StartNewBlock();
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
@@ -645,6 +658,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
Thread->OpDispatcher->StartNewBlock();
}
if (TableInfo && TableInfo->OpcodeDispatcher.OpDispatch) {
@@ -692,6 +706,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST ||
Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::BAD_RELOCATION) {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
} else if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::UNIMPLEMENTED_INST) {
Thread->OpDispatcher->UnimplementedOp(DecodedInfo);
} else {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
}
@@ -706,8 +722,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {{}, 0, 0, 0, 0};
Thread->OpDispatcher->DelayedDisownBuffer();
return {std::nullopt, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
@@ -774,6 +790,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
// OpDispatcher IR already released in this case.
return {{}, nullptr, 0, 0, false};
}
@@ -784,6 +801,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
// Raced to compile, release the OpDispatcher IR.
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
@@ -813,6 +831,17 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst) {
if constexpr (BLOCK_DEBUGGING) {
// Block debugging logic is hand-written and needs to be handled with care.
// Force MaxInst to only be one in this case.
MaxInst = 1;
// If the entrypoint is part of the single step targets then single step it.
if (BlockDebuggerTracker.IsSingleStepTarget(GuestRIP)) {
return CompileSingleStep(Frame, GuestRIP);
}
}
auto Thread = Frame->Thread;
FEXCORE_PROFILE_SCOPED("CompileBlock");
FEXCORE_PROFILE_ACCUMULATION(Thread, AccumulatedJITTime);
File diff suppressed because it is too large. Load diff
@@ -28,6 +28,10 @@ class ContextImpl;
namespace FEXCore::CPU {
#define STATE_PTR(STATE_TYPE, FIELD) STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
#define STATE_PTR_IDX(STATE_TYPE, FIELD, INDEX) STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::STATE_TYPE, FIELD, INDEX)
#define FALLBACK_HANDLER_OFFSET(INDEX, FIELD) \
STATE.R(), \
(ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.FallbackHandlerPointers, INDEX) + offsetof(FEXCore::Core::FallbackABIInfo, FIELD))
class Dispatcher final : public Arm64Emitter {
public:
@@ -95,6 +99,18 @@ private:
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
// F64 reduced-precision shared handlers
uint64_t F64SinHandlerAddress {};
uint64_t F64CosHandlerAddress {};
uint64_t F64TanHandlerAddress {};
uint64_t F64F2XM1HandlerAddress {};
uint64_t F64ScaleHandlerAddress {};
uint64_t F64AtanHandlerAddress {};
uint64_t F64FYL2XHandlerAddress {};
uint64_t F64FYL2XP1HandlerAddress {};
uint64_t F64FPREMHandlerAddress {};
uint64_t F64FPREM1HandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
@@ -105,6 +121,26 @@ private:
void EmitF32ToExtF80();
void EmitF64ToExtF80();
// Shared label set for the LUT-based F64 log2 path used by both FYL2X and
// FYL2XP1. The pool is emitted once via EmitF64Log2Constants.
struct F64Log2Constants {
ARMEmitter::ForwardLabel One;
ARMEmitter::ForwardLabel A0, A1, A2, A3, A4, A5, A6, A7;
ARMEmitter::ForwardLabel Table;
};
void EmitF64Sin();
void EmitF64Cos();
void EmitF64Tan();
void EmitF64F2XM1();
void EmitF64Scale();
void EmitF64Atan();
void EmitF64FYL2X(F64Log2Constants& C);
void EmitF64FYL2XP1(F64Log2Constants& C);
void EmitF64Log2Constants(F64Log2Constants& C);
void EmitF64FPREM();
void EmitF64FPREM1();
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
+76 -46
View File
@@ -124,9 +124,9 @@ uint8_t Decoder::ReadByte() {
}
std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream + InstructionSize + Offset);
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize + Offset);
if (CheckRangeExecutable(ByteAddress, 1)) {
return InstStream[InstructionSize + Offset];
return InstStream.AdjustedInstStream[InstructionSize + Offset];
} else {
return std::nullopt;
}
@@ -136,9 +136,9 @@ std::pair<uint64_t, bool> Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
uint64_t Address = reinterpret_cast<uint64_t>(InstStream + InstructionSize);
uint64_t Address = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize);
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream[InstructionSize], Size);
std::memcpy(&Res, &InstStream.AdjustedInstStream[InstructionSize], Size);
} else {
HitNonExecutableRange = true;
// See PeekByte, this specific case may cause some executable memory to read as 0 but it doesn't matter as the entire instruction will be rolled back anyway.
@@ -342,7 +342,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
}
bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
Decoder::DecodedBlockStatus Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
// TODO: Move this in to `NormalOpHeader`, Dispatch tables have a bug currently where some subtables don't inherit flags correctly.
@@ -354,11 +354,16 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SUPPORTS_LOCK) && (DecodeInst->Flags & DecodeFlags::FLAG_LOCK)) {
// Instruction has lock prefix but doesn't support lock.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P), "Group Ops "
@@ -390,15 +395,15 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
if (Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
const bool UseVEXL = Options.L && !(Info->Flags & InstFlags::FLAGS_VEX_L_IGNORE);
@@ -507,7 +512,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -576,7 +581,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const auto VEXOperand = Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_SRC_MASK;
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_NO_OPERAND && Options.vvvv) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) {
@@ -594,11 +599,11 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM) {
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SF_MOD_DST) {
if (!ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
} else {
if (!ModRMOperand(DecodeInst->Dest, DecodeInst->Src[CurrentSrc], HasXMMDst, HasXMMSrc, HasMMDst, HasMMSrc, Is8BitDest, Is8BitSrc)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
++CurrentSrc;
@@ -660,22 +665,27 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
Bytes = 0;
}
if ((DecodeInst->Flags & DecodeFlags::FLAG_LOCK) && DecodeInst->Dest.IsGPR()) {
// Instruction has lock prefix, but the destination isn't memory, this is invalid.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
return DecodedBlockStatus::SUCCESS;
}
bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
Decoder::DecodedBlockStatus Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
DecodeInst->OPRaw = DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
@@ -732,7 +742,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
};
uint8_t Field = RegToField[ModRM.reg];
if (Field == 255) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LocalOp = (Field << 3) | ModRM.rm;
@@ -751,7 +761,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
if (!VEXTable) {
// AVX not enabled.
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint16_t map_select = 1;
@@ -761,7 +771,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
if ((Byte1 & 0b10000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
@@ -772,7 +782,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte1 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.L = (Byte1 & 0b100) != 0;
@@ -783,14 +793,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte2 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
@@ -801,7 +811,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -831,14 +841,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
FEXCORE_TELEMETRY_SET(TYPE_USES_EVEX_OPS, 1);
// EVEX unsupported
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_MSG_A_FMT("Invalid instruction decoding type");
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
Decoder::DecodedBlockStatus Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
LastEscapePrefix = 0;
Instruction.fill(0);
@@ -849,7 +859,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
for (;;) {
if (InstructionSize >= MAX_INST_SIZE) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint8_t Op = ReadByte();
switch (Op) {
@@ -1035,10 +1045,10 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
}
if (DecodeInst->Dest.IsGPR()) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
return true;
return DecodedBlockStatus::SUCCESS;
}
void Decoder::DecodeREXIfValid(int8_t ExpectedOffset) {
@@ -1076,16 +1086,16 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
HitBadRelocation = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
auto ErrorDuringDecoding = DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
if (ErrorDuringDecoding != DecodedBlockStatus::SUCCESS || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
auto Result = ErrorDuringDecoding != DecodedBlockStatus::SUCCESS ? ErrorDuringDecoding :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
DecodeInst->InstSize = 0;
return Result;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
@@ -1331,7 +1341,7 @@ void Decoder::AddBranchTarget(uint64_t Target) {
}
}
const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
const Decoder::DecodeStream Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
@@ -1342,10 +1352,23 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
// Offset 0x400: vtime
// Offset 0x800: vgetcpu
uint64_t Offset = RIP - VSyscall_Base;
return VSyscallData + Offset;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = VSyscallData + Offset,
};
}
return _InstStream - EntryPoint + RIP;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = _InstStream - EntryPoint + RIP,
};
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
DecodeInstructionsAtEntry(&Thread, InstStream, PC, MaxInst);
bool Uncacheable = HitBadRelocation;
DelayedDisownBuffer();
return !Uncacheable;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
@@ -1366,7 +1389,6 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
@@ -1465,6 +1487,13 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
}
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
if (HitBadRelocation) {
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks = {*BlockIt};
BlockInfo.EntryPoints.clear();
BlockInfo.CodePages.clear();
return;
}
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
@@ -1483,7 +1512,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
// Can not continue this block at all on invalid instruction
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
if (!EntryBlock && BlockIt->BlockStatus != DecodedBlockStatus::BAD_RELOCATION) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
// Just need to undo additions that this block decoding has caused.
@@ -1493,10 +1522,11 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
"PartialDecode",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
BlockIt->BlockStatus == DecodedBlockStatus::UNIMPLEMENTED_INST ? "Unimplemented" :
"PartialDecode",
OpAddress);
}
break;
+28 -5
View File
@@ -32,6 +32,7 @@ public:
NOEXEC_INST,
PARTIAL_DECODE_INST,
BAD_RELOCATION,
UNIMPLEMENTED_INST,
};
// New Frontend decoding
@@ -54,6 +55,8 @@ public:
};
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
@@ -90,7 +93,7 @@ private:
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
@@ -109,8 +112,8 @@ private:
InstructionSize += Size;
}
bool NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
DecodedBlockStatus NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
DecodedBlockStatus NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
void DecodeREXIfValid(int8_t ExpectedOffset = -1);
@@ -125,7 +128,27 @@ private:
bool HitNonExecutableRange {};
bool HitBadRelocation {};
const uint8_t* InstStream {};
struct DecodeStream {
// Original instruction stream RIP location.
const uint8_t* InstStream;
// Adjusted location for FEX actually decodes from.
const uint8_t* AdjustedInstStream;
DecodeStream& operator-=(size_t offset) noexcept {
InstStream -= offset;
AdjustedInstStream -= offset;
return *this;
}
DecodeStream& operator+=(size_t offset) noexcept {
InstStream += offset;
AdjustedInstStream += offset;
return *this;
}
};
DecodeStream InstStream;
IR::OpSize GetGPROpSize() const {
return BlockInfo.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
@@ -167,6 +190,6 @@ private:
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_TABLE_SIZE>* VEXTable {};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_GROUP_TABLE_SIZE>* VEXTableGroup {};
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
const DecodeStream AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
};
} // namespace FEXCore::Frontend
@@ -302,6 +302,16 @@ struct OpHandlers<IR::OP_F80FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame, true};
const X80SoftFloat One {&State.State, 1.0};
return X80SoftFloat::FYL2X(&State.State, X80SoftFloat::FADD(&State.State, Src1, One), Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -417,6 +427,14 @@ struct OpHandlers<IR::OP_F64FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(1.0 + src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -72,6 +72,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80DIV>::handle)};
Info[Core::OPINDEX_F80FYL2X] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2X>::handle)};
Info[Core::OPINDEX_F80FYL2XP1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2XP1>::handle)};
Info[Core::OPINDEX_F80ATAN] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80ATAN>::handle)};
Info[Core::OPINDEX_F80FPREM1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
@@ -97,6 +99,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle)};
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle)};
Info[Core::OPINDEX_F64FYL2XP1] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2XP1>::handle)};
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle)};
@@ -254,6 +258,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
COMMON_BINARY_X87_OP(MUL)
COMMON_BINARY_X87_OP(DIV)
COMMON_BINARY_X87_OP(FYL2X)
COMMON_BINARY_X87_OP(FYL2XP1)
COMMON_BINARY_X87_OP(ATAN)
COMMON_BINARY_X87_OP(FPREM1)
COMMON_BINARY_X87_OP(FPREM)
@@ -268,6 +273,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
// Double Precision Binary
COMMON_BINARY_F64_OP(FYL2X)
COMMON_BINARY_F64_OP(FYL2XP1)
COMMON_BINARY_F64_OP(ATAN)
COMMON_BINARY_F64_OP(FPREM1)
COMMON_BINARY_F64_OP(FPREM)
+3 -3
View File
@@ -274,7 +274,7 @@ DEF_OP(CmpPairZ) {
// Restore NzCV
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 0, 0xb /* NzCV */);
rmif(TMP1, 28, 0xb /* NzCV */);
} else {
cset(ARMEmitter::Size::i32Bit, TMP2, ARMEmitter::Condition::CC_EQ);
bfi(ARMEmitter::Size::i32Bit, TMP1, TMP2, 30 /* lsb: Z */, 1);
@@ -523,7 +523,7 @@ DEF_OP(AndWithFlags) {
}
DEF_OP(AndShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
auto Op = IROp->C<IR::IROp_AndShift>();
and_(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1), GetReg(Op->Src2), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
@@ -721,7 +721,7 @@ DEF_OP(Extr) {
}
DEF_OP(PDep) {
auto Op = IROp->C<IR::IROp_PExt>();
auto Op = IROp->C<IR::IROp_PDep>();
const auto EmitSize = ConvertSize48(IROp);
const auto Dest = GetReg(Node);
@@ -329,7 +329,7 @@ DEF_OP(TelemetrySetValue) {
auto Op = IROp->C<IR::IROp_TelemetrySetValue>();
auto Src = GetReg(Op->Value);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.TelemetryValueAddresses[Op->TelemetryValueIndex]));
ldr(TMP2, STATE_PTR_IDX(CpuStateFrame, Pointers.TelemetryValueAddresses, Op->TelemetryValueIndex));
// Cortex fuses cmp+cset.
cmp(ARMEmitter::Size::i32Bit, Src, 0);
@@ -55,6 +55,29 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if constexpr (Context::BLOCK_DEBUGGING) {
// Skip block linking when BLOCK_DEBUGGING as it adds overhead and is unncessary.
// This is a debug only feature and doesn't need caching help.
bool IsInlineRIP = IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP);
ARMEmitter::ForwardLabel l_ExitLink;
if (IsInlineRIP) {
ldr(TMP1, &l_ExitLink);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
} else {
auto RipReg = GetReg(Op->NewRIP);
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
}
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
br(TMP2);
if (IsInlineRIP) {
BindOrRestart(&l_ExitLink);
dc64(NewRIP);
}
return;
}
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef ARCHITECTURE_arm64ec
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
@@ -265,7 +288,10 @@ DEF_OP(Syscall) {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = GPRSpillMask,
.FPRSpillMask = FPRSpillMask,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -299,7 +325,12 @@ DEF_OP(Syscall) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r1,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = GPRSpillMask,
.FPRFillMask = FPRSpillMask,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -322,7 +353,10 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(TMP1, true, ~0U, ~0U, false); // spill to ctx before ra64 spill
// spill to ctx before ra64 spill
SpillStaticRegs(TMP1, {
.NZCV = false,
});
PushDynamicRegs(TMP1);
@@ -337,7 +371,10 @@ DEF_OP(Thunk) {
PopDynamicRegs();
FillStaticRegs(true, ~0U, ~0U, std::nullopt, std::nullopt, false); // load from ctx after ra64 refill
// load from ctx after ra64 refill
FillStaticRegs({
.NZCV = false,
});
}
DEF_OP(ValidateCode) {
+50 -36
View File
@@ -68,6 +68,10 @@ PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintMsg(const char* Value) {
LogMan::Msg::DFmt("{}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
@@ -133,8 +137,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.S(), Src1.S());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -151,8 +155,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -176,8 +180,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(ARMEmitter::Size::i32Bit, TMP2, Src1);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -194,8 +198,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -212,8 +216,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -230,8 +234,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -254,8 +258,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -276,8 +280,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -294,8 +298,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -312,8 +316,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -330,8 +334,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -351,8 +355,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -369,8 +373,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -394,8 +398,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -416,8 +420,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -434,8 +438,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
// tmp2 (x1/x11): source 2
// tmp3 (x2/x12): source 3
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
stp<ARMEmitter::IndexType::PRE>(TMP1, ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
@@ -476,8 +480,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP2.Q(), Src2.Q());
movz(ARMEmitter::Size::i32Bit, TMP1, Control);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP2, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP2);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -636,6 +640,8 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Ptrs.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Ptrs.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Ptrs.PrintMsgValue = reinterpret_cast<uint64_t>(PrintMsg);
Ptrs.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Ptrs.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Ptrs.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
@@ -770,8 +776,15 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
constexpr size_t InterruptPageOffset =
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState);
if constexpr (InterruptPageOffset <= 32760) {
str(ARMEmitter::XReg::zr, STATE, InterruptPageOffset);
} else {
// Need to use vector 128-bit store for this range.
// Doesn't matter which register we use to store.
str(ARMEmitter::QReg::q0, STATE, InterruptPageOffset);
}
}
#ifdef ARCHITECTURE_arm64ec
@@ -841,6 +854,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CallReturnTargets.clear();
PendingJumpThunks.clear();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
Relocations.resize(PrevNumAllocations, FEXCore::CPU::Relocation::Default()); // Discard any relocations generated from a previous attempt
CodeData.EntryPoints.clear();
+159 -84
View File
@@ -563,12 +563,12 @@ DEF_OP(LoadDF) {
auto Flag = X86State::RFLAG_DF_RAW_LOC;
// DF needs sign extension to turn 0x1/0xFF into 1/-1
ldrsb(Dst.X(), STATE, offsetof(FEXCore::Core::CPUState, flags[Flag]));
ldrsb(Dst.X(), STATE, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, Flag));
}
DEF_OP(ContextClear) {
auto Op = IROp->C<IR::IROp_ContextClear>();
if (CTX->HostFeatures.SupportsCLZERO) {
if (CTX->HostFeatures.PreferZVAForVZero) {
// We can use CLZero directly when hardware supports it.
// Provides a fairly generous speed-up on Ampere1A hardware.
// TODO: When FEAT_MOPS hardware ships, test memset using MOPS.
@@ -1849,13 +1849,6 @@ DEF_OP(StoreMemTSO) {
}
DEF_OP(MemSet) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic forward path directly matches ARM's SETP/SETM/SETE instruction,
// while the backward version needs some fixup to convert it to a forward direction.
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
// Additionally: This is commonly used as a memset to zero. If we know up-front with an inline constant
// that the value is zero, we can optimize any operation larger than 8-bit down to 8-bit to use the MOPS implementation.
const auto Op = IROp->C<IR::IROp_MemSet>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -1933,8 +1926,30 @@ DEF_OP(MemSet) {
ARMEmitter::SubRegSize::i8Bit;
auto EmitMemset = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
// Sets the result to the final address written depending on
// whether or not the memset is forwards or backwards.
const auto MakeFinalAddress = [&] {
if (IsBackwards) {
switch (Size) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
} else {
switch (Size) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -1943,12 +1958,56 @@ DEF_OP(MemSet) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
const bool Is8Bit = SubRegSize == ARMEmitter::SubRegSize::i8Bit;
// We can handle 8-bit memsets and any other size that happens
// to be using an inlined zero value (resulting in the use of ZR).
//
// NOTE:
// Strictly speaking, this can also be trivially expanded to handle other sizes
// that happen to use any value that could fit inside a byte if the need
// arises. This does increase branching and code generation, however, since
// we'd still need to emit the fallback in the event a value for a larger size
// falls outside the range of a byte instead of only generating the MOPS code.
if (Is8Bit || Value == ARMEmitter::Reg::zr) {
// If we're performing a non-byte-sized zeroing operation then we need to
// scale the counter accordingly. (e.g. a 64-bit memset of size 2 needs to
// be turned into an 8-bit memset of size 16)
if (!Is8Bit) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ToUnderlying(SubRegSize));
}
// If backwards, then we need to adjust the starting address because
// set{p, m, e} memset forwards, so we need to slide this bad boy
// back like: (address - count) + 1.
//
// This lets us offset the address such that we can treat a backwards
// memset as if it were a forwards one.
if (IsBackwards) {
sub(TMP2, TMP2, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
}
// Unfortunately set operations fiddle with NZCV, so we need to preserve it.
mrs(TMP3, ARMEmitter::SystemRegister::NZCV);
setp(TMP2, TMP1, Value.X());
setm(TMP2, TMP1, Value.X());
sete(TMP2, TMP1, Value.X());
msr(ARMEmitter::SystemRegister::NZCV, TMP3);
MakeFinalAddress();
(void)Bind(&DoneInternal);
return;
}
}
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::BackwardLabel AgainInternal256 {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
ARMEmitter::BackwardLabel AgainInternal128 {};
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
@@ -1986,39 +2045,23 @@ DEF_OP(MemSet) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
}
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
MemStoreTSO(Value, Size, SizeDirection);
} else {
MemStore(Value, OpSize, SizeDirection);
MemStore(Value, Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
MakeFinalAddress();
};
if (DirectionIsInline) {
@@ -2041,10 +2084,6 @@ DEF_OP(MemSet) {
}
DEF_OP(MemCpy) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic path directly matches ARM's CPYP/CPYM/CPYE instruction,
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
const auto Op = IROp->C<IR::IROp_MemCpy>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -2175,8 +2214,40 @@ DEF_OP(MemCpy) {
};
auto EmitMemcpy = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
const auto FinalizeAddresses = [&] {
if (IsBackwards) {
switch (Size) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
} else {
switch (Size) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -2185,6 +2256,48 @@ DEF_OP(MemCpy) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
// In the event we have an overlap (gross), we need to fall back
// to the non-mops copy handler. Since the overlap check needs to
// make use of NZCV, we need to save it. This can be avoided with
// ARMv9.6+'s FEAT_CMPBR, but alas, we don't have access to that right now.
//
// NOTE: That we need to temporarily trash TMP1 and restore it after the
// comparison.
ARMEmitter::ForwardLabel OverlapCase;
mrs(TMP4, ARMEmitter::SystemRegister::NZCV);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP2, TMP3);
cmp(ARMEmitter::Size::i64Bit, TMP1, Length.X());
mov(TMP1, Length.X());
(void)bc(ARMEmitter::Condition::CC_LT, &OverlapCase);
// If doing something larger than a byte copy, then we need to scale
// the counter value accordingly to convert it to bytes.
if (Size > 1) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ilog2(Size));
}
// Adjust addresses so that we treat the backward copy as a forward copy
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, Size);
}
// Unfortunately copy operations fiddle with NZCV, so we need to preserve it.
cpyfp(TMP2, TMP3, TMP1);
cpyfm(TMP2, TMP3, TMP1);
cpyfe(TMP2, TMP3, TMP1);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
(void)b(&DoneInternal);
// Turns out we overlap and need to fall back. Make sure to restore NZCV.
(void)Bind(&OverlapCase);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
}
ARMEmitter::ForwardLabel AbsPos {};
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
@@ -2198,7 +2311,7 @@ DEF_OP(MemCpy) {
sub(ARMEmitter::Size::i64Bit, TMP4, TMP4, 32);
(void)tbnz(TMP4, 63, &AgainInternal);
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2233,7 +2346,7 @@ DEF_OP(MemCpy) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2241,9 +2354,9 @@ DEF_OP(MemCpy) {
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
MemCpyTSO(Size, SizeDirection);
} else {
MemCpy(OpSize, SizeDirection);
MemCpy(Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
@@ -2255,54 +2368,14 @@ DEF_OP(MemCpy) {
mov(TMP2, MemRegSrc.X());
mov(TMP3, Length.X());
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
FinalizeAddresses();
};
if (DirectionIsInline) {
LOGMAN_THROW_A_FMT(DirectionConstant == 1 || DirectionConstant == -1, "unexpected direction");
EmitMemcpy(DirectionConstant);
} else {
// Emit forward direction memset then backward direction memset.
// Emit forward direction memcpy then backward direction memcpy.
for (int32_t Direction : {1, -1}) {
EmitMemcpy(Direction);
if (Direction == 1) {
@@ -2327,11 +2400,12 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CIVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
@@ -2355,11 +2429,12 @@ DEF_OP(CacheLineClean) {
auto MemReg = GetReg(Op->Addr);
// Clean dcache only
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
+32 -4
View File
@@ -73,8 +73,8 @@ DEF_OP(Break) {
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Constant);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case Core::FAULT_SIGILL:
@@ -210,6 +210,25 @@ DEF_OP(Print) {
PopDynamicRegs();
}
DEF_OP(PrintMsg) {
auto Op = IROp->C<IR::IROp_PrintMsg>();
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, reinterpret_cast<uintptr_t>(Op->Value));
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintMsgValue));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillStaticRegs();
PopDynamicRegs();
}
DEF_OP(ProcessorID) {
if (CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
mrs(GetReg(Node), ARMEmitter::SystemRegister::TPIDRRO_EL0);
@@ -227,7 +246,10 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(TMP1, false, SpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = SpillMask,
.FPRs = false,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -264,7 +286,13 @@ DEF_OP(ProcessorID) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r8,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = SpillMask,
.FPRs = false,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
+245 -56
View File
@@ -977,7 +977,7 @@ DEF_OP(LoadNamedVectorConstant) {
}
// Load the pointer.
auto GenerateMemOperand = [this](IR::OpSize OpSize, uint32_t NamedConstant, ARMEmitter::Register Base) {
const auto ConstantOffset = offsetof(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants[NamedConstant]);
const auto ConstantOffset = ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants, NamedConstant);
if (ConstantOffset <= 255 || // Unscaled 9-bit signed
((ConstantOffset & (IR::OpSizeToSize(OpSize) - 1)) == 0 &&
@@ -985,13 +985,13 @@ DEF_OP(LoadNamedVectorConstant) {
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, ConstantOffset);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[NamedConstant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, NamedConstant));
return ARMEmitter::ExtendedMemOperand(TMP1, ARMEmitter::IndexType::OFFSET, 0);
};
if (OpSize == IR::OpSize::i256Bit) {
// Handle SVE 32-byte variant upfront.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, Op->Constant));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Dst.Z(), PRED_TMP_32B.Zeroing(), TMP1, 0);
return;
}
@@ -1013,7 +1013,7 @@ DEF_OP(LoadNamedVectorIndexedConstant) {
const auto Dst = GetVReg(Node);
// Load the pointer.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers, Op->Constant));
switch (OpSize) {
case IR::OpSize::i8Bit: ldrb(Dst, TMP1, Op->Index); break;
@@ -1036,24 +1036,28 @@ DEF_OP(VMov) {
const auto Dst = GetVReg(Node);
const auto Source = GetVReg(Op->Source);
const auto Sub64BitHandler = [&](ARMEmitter::SubRegSize InsertSize) {
if (Dst != Source) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
ins(InsertSize, Dst, 0, Source, 0);
} else {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(InsertSize, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
}
};
switch (OpSize) {
case IR::OpSize::i8Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i8Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i8Bit);
break;
}
case IR::OpSize::i16Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i16Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i16Bit);
break;
}
case IR::OpSize::i32Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i32Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i32Bit);
break;
}
case IR::OpSize::i64Bit: {
@@ -1095,16 +1099,21 @@ DEF_OP(VAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE ADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE ADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of ADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
addp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
addp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1298,16 +1307,21 @@ DEF_OP(VFAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE FADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE FADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of FADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
faddp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
faddp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1434,8 +1448,8 @@ DEF_OP(VFMin) {
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on false.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector1.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
@@ -1466,7 +1480,8 @@ DEF_OP(VFMax) {
const auto Mask = PRED_TMP_32B;
const auto ComparePred = ARMEmitter::PReg::p0;
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector2.Z(), Vector1.Z());
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector1.Z(), Vector2.Z());
not_(ComparePred, Mask.Zeroing(), ComparePred);
if (Dst == Vector1) {
// Trivial case where Vector1 is also the destination.
@@ -1488,17 +1503,17 @@ DEF_OP(VFMax) {
if (Dst == Vector1) {
// Destination is already Vector1, need to insert Vector2 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
}
}
}
@@ -1525,9 +1540,14 @@ DEF_OP(VFRecp) {
return;
}
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
if (Dst != Vector) {
fmov(SubRegSize.Vector, Dst.Z(), 1.0);
fdiv(SubRegSize.Vector, Dst.Z(), Pred, Dst.Z(), Vector.Z());
} else {
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
}
} else {
if (IsScalar) {
if (ElementSize == IR::OpSize::i32Bit && HostSupportsRPRES) {
@@ -1779,10 +1799,14 @@ DEF_OP(VUMin) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1828,10 +1852,14 @@ DEF_OP(VSMin) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1877,10 +1905,14 @@ DEF_OP(VUMax) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1926,10 +1958,14 @@ DEF_OP(VSMax) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -2979,9 +3015,14 @@ DEF_OP(VInsElement) {
auto Reg = GetVReg(Op->DestVector);
if (HostSupportsSVE256 && Is256Bit) {
// Broadcast our source value across a temporary,
// then combine with the destination.
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
// Broadcast our source value across a temporary, then combine
// with the destination.
//
// We don't need to perform the dup if we're just merging a 128-bit vector into
// into an equivalent position since we have a predicate set up already.
if (!(ElementSize == IR::OpSize::i128Bit && SrcIdx == DestIdx)) {
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
}
// We don't need to move the data unnecessarily if
// DestVector just so happens to also be the IR op
@@ -2994,10 +3035,12 @@ DEF_OP(VInsElement) {
if (ElementSize == IR::OpSize::i128Bit) {
if (DestIdx == 0) {
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), VTMP2.Z());
const auto Source = SrcIdx == 0 ? SrcVector : VTMP2;
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), Source.Z());
} else {
const auto Source = SrcIdx == 1 ? SrcVector : VTMP2;
not_(Predicate, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), VTMP2.Z());
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), Source.Z());
}
} else {
const auto UpperBound = 16 >> FEXCore::ilog2(IR::OpSizeToSize(ElementSize));
@@ -4434,7 +4477,7 @@ DEF_OP(VFNMLA) {
// - SVE - FMLS
// - ASIMD - FMLS
// - Scalar - FMSUB
const auto Op = IROp->C<IR::IROp_VFMLA>();
const auto Op = IROp->C<IR::IROp_VFNMLA>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4502,7 +4545,7 @@ DEF_OP(VFNMLS) {
// - ASIMD - FMLS (With Negated addend)
// - Scalar - FNMADD
const auto Op = IROp->C<IR::IROp_VFMLS>();
const auto Op = IROp->C<IR::IROp_VFNMLS>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4611,4 +4654,150 @@ DEF_OP(VFCopySign) {
}
}
DEF_OP(F64FPREM) {
const auto Op = IROp->C<IR::IROp_F64FPREM>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREMHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64FPREM1) {
const auto Op = IROp->C<IR::IROp_F64FPREM1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SIN) {
const auto Op = IROp->C<IR::IROp_F64SIN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64SinHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64COS) {
const auto Op = IROp->C<IR::IROp_F64COS>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64CosHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64TAN) {
const auto Op = IROp->C<IR::IROp_F64TAN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64TanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src1=y(ST1), Src2=x(ST0). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64ATAN) {
const auto Op = IROp->C<IR::IROp_F64ATAN>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64AtanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2X) {
const auto Op = IROp->C<IR::IROp_F64FYL2X>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2XP1) {
const auto Op = IROp->C<IR::IROp_F64FYL2XP1>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XP1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SCALE) {
const auto Op = IROp->C<IR::IROp_F64SCALE>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64ScaleHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64F2XM1) {
const auto Op = IROp->C<IR::IROp_F64F2XM1>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64F2XM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
} // namespace FEXCore::CPU
@@ -41,6 +41,9 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
// Disable THP on the Lookup cache.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<const void*>(PagePointer), TotalCacheSize, FEXCore::Allocator::THPControl::Disable);
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -84,8 +87,11 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
}
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// TODO: Preserve code cache entries?
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
// TODO: Rename this member to avoid confusion with code caching
CachedCodePages.clear();
}
+1 -1
View File
@@ -3,7 +3,7 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/Utils/WritePriorityMutex.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -502,7 +502,7 @@ void OpDispatchBuilder::LEAVEOp(OpcodeArgs) {
auto NewGPR = Pop(OperandSize, SP);
// Store the new stack pointer
StoreGPRRegister(X86State::REG_RSP, SP, OperandSize);
StoreGPRRegister(X86State::REG_RSP, SP, GPRSize);
// Store what we loaded to RBP
StoreGPRRegister(X86State::REG_RBP, NewGPR, OperandSize);
@@ -2427,9 +2427,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
auto BitSelect = (Size == (LshrSize * 8)) ? Src : Src.And(Mask);
auto LshrOpSize = IR::SizeToOpSize(LshrSize);
// OF/SF/AF/PF undefined. ZF must be preserved. We choose to preserve OF/SF
// too since we just use an rmif to insert into CF directly. We could
// optimize perhaps.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
//
// Set CF before the action to save a move, except for complements where we
// can reuse the invert.
@@ -2547,7 +2548,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
Value = _Lshr(std::max(OpSize::i32Bit, GetOpSize(Value)), Value, BitSelect.Ref());
}
// OF/SF/ZF/AF/PF undefined.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
SetCFDirect(Value, 0, true);
}
}
@@ -2643,7 +2647,10 @@ void OpDispatchBuilder::IMULOp(OpcodeArgs) {
}
// 64-bit special cased to save a move
Ref Result = Size < OpSize::i64Bit ? _Mul(OpSize::i64Bit, Src1, Src2) : nullptr;
Ref Result {};
if (Size < OpSize::i64Bit) {
Result = _Mul(OpSize::i64Bit, Src1, Src2);
}
Ref ResultHigh {};
if (Size == OpSize::i8Bit) {
// Result is stored in AX
@@ -3446,30 +3453,38 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
}
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("LODSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
StoreResultGPR(Op, Src);
// Offset the pointer
Ref TailDest_RSI = LoadGPRRegister(X86State::REG_RSI);
StoreGPRRegister(X86State::REG_RSI, OffsetByDir(TailDest_RSI, IR::OpSizeToSize(Size)));
Ref TailDest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t PtrDir) {
// XXX: Theoretically LODS could be optimized to
// RSI += {-}(RCX * Size)
// RAX = [RSI - Size]
@@ -3497,7 +3512,8 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
@@ -3513,8 +3529,13 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RSI = Add(OpSize::i64Bit, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
TailDest_RSI = Add(AddrSize, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
// Jump back to the start, we have more work to do
Jump(LoopStart);
@@ -4213,6 +4234,94 @@ void OpDispatchBuilder::UpdatePrefixFromSegment(Ref Segment, uint32_t SegmentReg
}
}
uint64_t OpDispatchBuilder::CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad) {
if constexpr (!Context::BLOCK_DEBUGGING) {
LOGMAN_MSG_A_FMT("Tried to calculate address without block debugging enabled!");
FEX_UNREACHABLE;
}
const auto GPRSize = GetGPROpSize();
const auto GPRMask = GPRSize == OpSize::i64Bit ? ~0ULL : ~0U;
// This makes the assumption that InternalThreadState is synchronized at the point of call!
uint64_t Ptr {};
if (Operand.IsLiteral()) {
Ptr = Operand.Literal();
if (Operand.Data.Literal.Size != 8 && IsLoad) {
// zero extend
uint64_t width = Operand.Data.Literal.Size * 8;
Ptr &= ((1ULL << width) - 1);
}
} else if (Operand.IsGPR()) {
// Not a memory source.
return ~0ULL;
} else if (Operand.IsGPRDirect()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
} else if (Operand.IsGPRIndirect() || Operand.IsGPRIndirectRelocation()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
Ptr += static_cast<int32_t>(Operand.Data.GPRIndirect.Displacement);
} else if (Operand.IsRIPRelative() || Operand.IsRIPRelativeRelocation()) {
// 64-bit is RIP relative, while 32-bit is absolute.
if (Is64BitMode) {
Ptr = Op->PC + Op->InstSize + static_cast<int32_t>(Operand.Data.RIPLiteral.Value) - Entry;
} else {
Ptr = Operand.Data.RIPLiteral.Value;
}
} else if (Operand.IsSIB() || Operand.IsSIBRelocation()) {
const bool IsVSIB = IsLoad && ((Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0);
if (IsVSIB) {
// TODO: Unhandled.
return ~0ULL;
}
if (Operand.Data.SIB.Base != FEXCore::X86State::REG_INVALID) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Base] & GPRMask;
}
if (Operand.Data.SIB.Index != FEXCore::X86State::REG_INVALID) {
Ptr += (Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Index] * Operand.Data.SIB.Scale) & GPRMask;
}
Ptr += static_cast<int32_t>(Operand.Data.SIB.Offset);
}
auto AppendSegment = [&](uint64_t Ptr, uint32_t Flags, uint32_t DefaultPrefix = FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX,
bool Override = false) -> uint64_t {
uint32_t Prefix = Flags & FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS;
if (Is64BitMode) {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.fs_cached;
} else if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.gs_cached;
}
// If there was any other segment in 64bit then it is ignored
} else {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX || Override) {
// If there was no prefix then use the default one if available
// Or the argument only uses a specific prefix (with override set)
Prefix = DefaultPrefix;
}
// With the segment register optimization we store the GDT bases directly in the segment register to remove indexed loads
switch (Prefix) {
[[likely]] case FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX:
return Ptr;
case FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX: return Ptr + Thread->CurrentFrame->State.es_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX: return Ptr + Thread->CurrentFrame->State.cs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX: return Ptr + Thread->CurrentFrame->State.ss_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX: return Ptr + Thread->CurrentFrame->State.ds_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX: return Ptr + Thread->CurrentFrame->State.fs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX: return Ptr + Thread->CurrentFrame->State.gs_cached;
default: FEX_UNREACHABLE;
}
}
return Ptr;
};
return AppendSegment(Ptr, Op->Flags);
};
AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand,
MemoryAccessType AccessType, bool IsLoad) {
const auto GPRSize = GetGPROpSize();
@@ -4303,6 +4412,7 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
auto [Align, LoadData, ForceLoad, AccessType, AllowUpperGarbage] = Options;
AddressMode A = DecodeAddress(Op, Operand, AccessType, true /* IsLoad */);
Ref Result {};
if (Operand.IsGPR()) {
const auto gpr = Operand.Data.GPR.GPR;
const auto highIndex = Operand.Data.GPR.HighBits ? 1 : 0;
@@ -4332,22 +4442,35 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
}
}
if ((IsOperandMem(Operand, true) && LoadData) || ForceLoad) {
const bool ShouldLoad = (IsOperandMem(Operand, true) && LoadData) || ForceLoad;
if (ShouldLoad) {
if (OpSize == OpSize::f80Bit) {
Ref MemSrc = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
return _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
if (CTX->HostFeatures.SupportsSVE()) {
Result = _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
} else {
// For X87 extended doubles, Split the load.
auto Res = _LoadMem(Class, OpSize::i64Bit, MemSrc, Align == OpSize::iInvalid ? OpSize : Align);
return _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
Result = _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
}
} else {
Result = _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
}
} else {
Result = LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (ShouldLoad && CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, true);
if (CTX->BlockDebuggerTracker.ContainsReadWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit read watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
return _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
} else {
return LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
return Result;
}
Ref OpDispatchBuilder::LoadGPRRegister(uint32_t GPR, IR::OpSize Size, uint8_t Offset, bool AllowUpperGarbage) {
@@ -4462,7 +4585,7 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
if (OpSize == OpSize::f80Bit) {
Ref MemStoreDst = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
if (CTX->HostFeatures.SupportsSVE()) {
_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, Src, MemStoreDst);
} else {
// For X87 extended doubles, split before storing
@@ -4473,6 +4596,16 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
} else {
_StoreMemAutoTSO(Class, OpSize, A, Src, Align == OpSize::iInvalid ? OpSize : Align);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, false);
if (CTX->BlockDebuggerTracker.ContainsWriteWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit write watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
}
}
void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src,
@@ -4484,11 +4617,10 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
StoreResult(Class, Op, Op->Dest, Src, Align, AccessType);
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
, CTX {ctx} {
ResetWorkingList();
, CTX {ctx}
, Thread {Thread} {
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
SaveAVXStateFunc = &OpDispatchBuilder::SaveAVXState;
RestoreAVXStateFunc = &OpDispatchBuilder::RestoreAVXState;
@@ -4501,7 +4633,8 @@ OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
}
void OpDispatchBuilder::ResetWorkingList() {
IREmitter::ResetWorkingList();
IREmitter::ReownOrClaimBuffer();
JumpTargets.clear();
BlockSetRIP = false;
DecodeFailure = false;
@@ -4645,35 +4778,36 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
case 0xCD: { // INT imm8
uint8_t Literal = Op->Src[0].Literal();
#ifndef _WIN32
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Linux) {
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
} else if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Wow64 ||
CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
#else
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
#endif
#ifdef ARCHITECTURE_arm64ec
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
}
}
}
#endif
Reason.ErrorRegister = Literal << 3 | (0b010);
Reason.Signal = Core::FAULT_SIGSEGV;
@@ -4887,6 +5021,11 @@ void OpDispatchBuilder::CLZeroOp(OpcodeArgs) {
}
void OpDispatchBuilder::Prefetch(OpcodeArgs, bool ForStore, bool Stream, uint8_t Level) {
if (Op->Src[0].IsGPR()) {
// NOP instance.
return;
}
Ref DestMem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
_Prefetch(ForStore, Stream, Level, DestMem, Invalid(), MemOffsetType::SXTX, 1);
}
@@ -4918,6 +5057,7 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
return;
}
const auto GPRSize = GetGPROpSize();
const auto SrcSize = OpSizeFromSrc(Op);
// Destination GPR size is always 4 or 8 bytes depending on widening
const auto DstSize = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REX_WIDENING ? OpSize::i64Bit : OpSize::i32Bit;
@@ -4926,16 +5066,15 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
// Incoming memory is 8, 16, 32, or 64
Ref Src {};
if (Op->Src[0].IsGPR()) {
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], GPRSize, Op->Flags);
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
} else {
Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
}
auto Result = _CRC32(Dest, Src, OpSizeFromSrc(Op));
auto Result = _CRC32(Dest, Src, SrcSize);
StoreResultGPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
template<bool Reseed>
void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
void OpDispatchBuilder::RDRANDOp(OpcodeArgs, bool Reseed) {
if (!CTX->HostFeatures.SupportsRAND) {
UnimplementedOp(Op);
return;
@@ -4960,9 +5099,6 @@ void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
}
}
template void OpDispatchBuilder::RDRANDOp<true>(OpcodeArgs);
template void OpDispatchBuilder::RDRANDOp<false>(OpcodeArgs);
void OpDispatchBuilder::BreakOp(OpcodeArgs, FEXCore::IR::BreakDefinition BreakDefinition) {
const auto GPRSize = GetGPROpSize();
+77 -137
View File
@@ -303,9 +303,11 @@ public:
StartNewBlock();
}
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
// Should only be called at the start of IR Emission.
void ResetWorkingList();
void ResetDecodeFailure() {
NeedsBlockEnd = DecodeFailure = false;
}
@@ -358,7 +360,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorAlignedOp(OpcodeArgs);
void MOVVectorUnalignedOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs, bool IsAVX);
void ALUOp(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, unsigned SrcIdx);
void LSLOp(OpcodeArgs);
void INTOp(OpcodeArgs);
@@ -468,8 +470,7 @@ public:
void AAMOp(OpcodeArgs);
void AADOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
void RDRANDOp(OpcodeArgs, bool Reseed);
enum class Segment {
FS,
@@ -499,8 +500,7 @@ public:
void VectorALUROp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VectorUnaryOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void RSqrt3DNowOp(OpcodeArgs, bool Duplicate);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorUnaryDuplicateOp(OpcodeArgs);
void VectorUnaryDuplicateOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void MOVQOp(OpcodeArgs, VectorOpType VectorType);
void MOVQMMXOp(OpcodeArgs);
@@ -522,36 +522,24 @@ public:
void PSLLDQ(OpcodeArgs);
void PSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void MOVDDUPOp(OpcodeArgs);
template<IR::OpSize DstElementSize>
void CVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void CVTFPR_To_GPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool Widen>
void Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void Scalar_CVT_Float_To_Float(OpcodeArgs);
void CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void Vector_CVT_Int_To_Float(OpcodeArgs, IR::OpSize SrcElementSize, bool Widen, bool IsAVX);
void Vector_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize, bool IsAVX);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void Vector_CVT_Float_To_Int(OpcodeArgs);
void Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode, bool IsAVX);
void MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs);
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void MASKMOVOp(OpcodeArgs);
void MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType);
void TZCNT(OpcodeArgs);
void LZCNT(OpcodeArgs);
template<IR::OpSize ElementSize>
void VFCMPOp(OpcodeArgs);
void VFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void SHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PINSROp(OpcodeArgs);
void PINSROp(OpcodeArgs, IR::OpSize ElementSize);
void InsertPSOp(OpcodeArgs);
void PExtrOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PSIGN(OpcodeArgs);
template<IR::OpSize ElementSize>
void VPSIGN(OpcodeArgs);
void PSIGN(OpcodeArgs, IR::OpSize ElementSize);
void VPSIGN(OpcodeArgs, IR::OpSize ElementSize);
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
@@ -574,53 +562,32 @@ public:
// AVX Ops
void AVXVectorXOROp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXVectorRound(OpcodeArgs);
void AVXVectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXScalar_CVT_Float_To_Float(OpcodeArgs);
void VectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarUnaryInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs);
void VectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void InsertMMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize>
void InsertCVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize DstElementSize>
void AVXInsertCVTGPR_To_FPR(OpcodeArgs);
void InsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
void AVXInsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void InsertScalar_CVT_Float_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs);
void InsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
RoundMode TranslateRoundType(uint8_t Mode);
template<IR::OpSize ElementSize>
void InsertScalarRound(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarRound(OpcodeArgs);
void InsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void InsertScalarFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarFCMPOp(OpcodeArgs);
void InsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize>
void AVXCVTGPR_To_FPR(OpcodeArgs);
void AVXVFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVXVFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void VADDSUBPOp(OpcodeArgs);
void VADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VAESDecOp(OpcodeArgs);
void VAESDecLastOp(OpcodeArgs);
@@ -629,34 +596,31 @@ public:
void VANDNOp(OpcodeArgs);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, Ref ZeroRegister, uint64_t Selector);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint64_t Selector);
void VBLENDPDOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPBLENDWOp(OpcodeArgs);
void VBROADCASTOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VDPPOp(OpcodeArgs);
void VDPPOp(OpcodeArgs, IR::OpSize ElementSize);
void VEXTRACT128Op(OpcodeArgs);
template<IROps IROp, IR::OpSize ElementSize>
void VHADDPOp(OpcodeArgs);
void VHADDPOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VHSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VINSERTOp(OpcodeArgs);
void VINSERTPSOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool IsStore>
void VMASKMOVOp(OpcodeArgs);
void VMASKMOVOp(OpcodeArgs, IR::OpSize ElementSize, bool IsStore);
void VMOVHPOp(OpcodeArgs);
void VMOVLPOp(OpcodeArgs);
void VMOVDDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSLDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSDOp(OpcodeArgs);
void VMOVSSOp(OpcodeArgs);
@@ -667,15 +631,14 @@ public:
void VMPSADBWOp(OpcodeArgs);
void VPACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPALIGNROp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs);
void VPCMPESTRMOp(OpcodeArgs);
void VPCMPISTRIOp(OpcodeArgs);
void VPCMPISTRMOp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPESTRMOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRMOp(OpcodeArgs, bool IsAVX);
void VCVTPH2PSOp(OpcodeArgs);
void VCVTPS2PHOp(OpcodeArgs);
@@ -688,36 +651,28 @@ public:
void VPERMILImmOp(OpcodeArgs, IR::OpSize ElementSize);
Ref VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize, Ref Src, Ref Indices);
template<IR::OpSize ElementSize>
void VPERMILRegOp(OpcodeArgs);
void VPERMILRegOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHADDSWOp(OpcodeArgs);
void VPHSUBOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHSUBSWOp(OpcodeArgs);
void VPINSRBOp(OpcodeArgs);
void VPINSRBWOp(OpcodeArgs, IR::OpSize ElementSize);
void VPINSRDQOp(OpcodeArgs);
void VPINSRWOp(OpcodeArgs);
void VPMADDUBSWOp(OpcodeArgs);
void VPMADDWDOp(OpcodeArgs);
template<bool IsStore>
void VPMASKMOVOp(OpcodeArgs);
void VPMASKMOVOp(OpcodeArgs, bool IsStore);
void VPMULHRSWOp(OpcodeArgs);
template<bool Signed>
void VPMULHWOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool Signed>
void VPMULLOp(OpcodeArgs);
void VPMULHWOp(OpcodeArgs, bool Signed);
void VPMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
void VPSADBWOp(OpcodeArgs);
void VPSHUFBOp(OpcodeArgs);
void VPSHUFWOp(OpcodeArgs, IR::OpSize ElementSize, bool Low);
void VPSLLOp(OpcodeArgs, IR::OpSize ElementSize);
@@ -726,7 +681,6 @@ public:
void VPSLLVOp(OpcodeArgs);
void VPSRAOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAVDOp(OpcodeArgs);
@@ -734,17 +688,14 @@ public:
void VPSRLDOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLDQOp(OpcodeArgs);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKHOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKLOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VSHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VTESTPOp(OpcodeArgs);
void VTESTPOp(OpcodeArgs, IR::OpSize ElementSize);
void VZEROOp(OpcodeArgs);
@@ -828,32 +779,24 @@ public:
void XSaveOp(OpcodeArgs);
void PAlignrOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void UCOMISxOp(OpcodeArgs);
void UCOMISxOp(OpcodeArgs, IR::OpSize ElementSize);
void LDMXCSR(OpcodeArgs);
void STMXCSR(OpcodeArgs);
template<IR::OpSize ElementSize>
void PACKUSOp(OpcodeArgs);
void PACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void PACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PACKSSOp(OpcodeArgs);
void PMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
template<IR::OpSize ElementSize, bool Signed>
void PMULLOp(OpcodeArgs);
void MOVQ2DQ(OpcodeArgs, bool ToXMM);
template<bool ToXMM>
void MOVQ2DQ(OpcodeArgs);
template<IR::OpSize ElementSize>
void ADDSUBPOp(OpcodeArgs);
void ADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void PFNACCOp(OpcodeArgs);
void PFPNACCOp(OpcodeArgs);
void PSWAPDOp(OpcodeArgs);
template<uint8_t CompType>
void VPFCMPOp(OpcodeArgs);
void VPFCMPOp(OpcodeArgs, uint8_t CompType);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
@@ -862,16 +805,12 @@ public:
void PMADDWD(OpcodeArgs);
void PMADDUBSW(OpcodeArgs);
template<bool Signed>
void PMULHW(OpcodeArgs);
void PMULHW(OpcodeArgs, bool Signed);
void PMULHRSW(OpcodeArgs);
void MOVBEOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void HSUBP(OpcodeArgs);
template<IR::OpSize ElementSize>
void PHSUB(OpcodeArgs);
void HSUBP(OpcodeArgs, IR::OpSize ElementSize);
void PHSUB(OpcodeArgs, IR::OpSize ElementSize);
void PHADDS(OpcodeArgs);
void PHSUBS(OpcodeArgs);
@@ -918,25 +857,24 @@ public:
};
RefVSIB LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags);
template<OpSize AddrElementSize>
void VPGATHER(OpcodeArgs);
void VPGATHER(OpcodeArgs, OpSize AddrElementSize);
template<IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed>
void ExtendVectorElements(OpcodeArgs);
template<IR::OpSize ElementSize>
void VectorRound(OpcodeArgs);
void AVXExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
void ExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
Ref VectorBlend(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VectorBlend(OpcodeArgs);
Ref VectorBlendImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorBlend(OpcodeArgs, IR::OpSize ElementSize);
void VectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
void PTestOpImpl(OpSize Size, Ref Dest, Ref Src);
void PTestOp(OpcodeArgs);
void AVXPHMINPOSUWOp(OpcodeArgs);
void PHMINPOSUWOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void DPPOp(OpcodeArgs);
void DPPOp(OpcodeArgs, IR::OpSize ElementSize);
void MPSADBWOp(OpcodeArgs);
void PCLMULQDQOp(OpcodeArgs);
@@ -1374,6 +1312,7 @@ private:
};
FEXCore::Context::ContextImpl* CTX {};
FEXCore::Core::InternalThreadState* Thread;
constexpr static unsigned FullNZCVMask = (1U << FEXCore::X86State::RFLAG_CF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_ZF_RAW_LOC) |
(1U << FEXCore::X86State::RFLAG_SF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_OF_RAW_LOC);
@@ -1441,7 +1380,7 @@ private:
Ref PALIGNROpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1, const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm, bool IsAVX);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask, bool IsAVX);
Ref PHADDSOpImpl(OpSize Size, Ref Src1, Ref Src2);
@@ -1479,7 +1418,7 @@ private:
Ref PSRLDOpImpl(OpcodeArgs, IR::OpSize ElementSize, Ref Src, Ref ShiftVec);
Ref SHUFOpImpl(OpcodeArgs, IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
Ref SHUFOpImpl(IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
void VMASKMOVOpImpl(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DataSize, bool IsStore, const X86Tables::DecodedOperand& MaskOp,
const X86Tables::DecodedOperand& DataOp);
@@ -1589,6 +1528,7 @@ private:
}
AddressMode DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, MemoryAccessType AccessType, bool IsLoad);
uint64_t CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad);
Ref LoadSource(RegClass Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {});
@@ -1666,7 +1606,7 @@ private:
[[nodiscard]]
static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
return static_cast<uint32_t>(ARRAY_OFFSETOF(Core::CPUState, gregs, reg));
}
[[nodiscard]]
@@ -1885,15 +1825,15 @@ private:
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF =
_Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
} else {
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, Value, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
}
}
@@ -1948,8 +1888,8 @@ private:
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
case MM0Index ... MM7Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, mm, Index - MM0Index);
case AVXHigh0Index ... AVXHigh15Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, avx_high, Index - AVXHigh0Index);
default: return ~0U;
}
}
@@ -2149,7 +2089,7 @@ private:
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContextGPR(OpSize::i8Bit, offsetof(Core::CPUState, flags[BitOffset]));
return _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(Core::CPUState, flags, BitOffset));
}
}
@@ -630,7 +630,7 @@ void OpDispatchBuilder::AVX128_VPSIGN(OpcodeArgs, IR::OpSize ElementSize) {
}
void OpDispatchBuilder::AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : ElementSize;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false);
@@ -1260,26 +1260,26 @@ void OpDispatchBuilder::AVX128_VAESKeyGenAssist(OpcodeArgs) {
}
void OpDispatchBuilder::AVX128_VPCMPESTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, false);
PCMPXSTRXOpImpl(Op, true, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPESTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, true);
PCMPXSTRXOpImpl(Op, true, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
}
void OpDispatchBuilder::AVX128_VPCMPISTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, false);
PCMPXSTRXOpImpl(Op, false, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPISTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, true);
PCMPXSTRXOpImpl(Op, false, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
@@ -1399,13 +1399,13 @@ void OpDispatchBuilder::AVX128_VSHUF(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
Result.Low = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
const uint8_t ShiftAmount = ElementSize == OpSize::i32Bit ? 0 : 2;
Result.High = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
Result.High = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -1484,12 +1484,12 @@ void OpDispatchBuilder::AVX128_VBLEND(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = VectorBlend(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
Result.Low = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
if (Is128Bit) {
Result = AVX128_Zext(Result.Low);
} else {
Result.High = VectorBlend(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
Result.High = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -2293,9 +2293,8 @@ void OpDispatchBuilder::AVX128_VCVTPS2PH(OpcodeArgs) {
_PopRoundingMode(OldFPCR);
}
// We need to eliminate upper junk if we're storing into a register with
// a 256-bit source (VCVTPS2PH's destination for registers is an XMM).
if (Op->Src[0].IsGPR() && SrcSize == OpSize::i256Bit) {
// We need to zero the upper 128 bits if we're storing into a register
if (Op->Dest.IsGPR()) {
Result = AVX128_Zext(Result.Low);
}
@@ -5,9 +5,9 @@
namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x0D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
{0x1D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x1D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, false>},
@@ -15,15 +15,15 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x8A, 1, &OpDispatchBuilder::PFNACCOp},
{0x8E, 1, &OpDispatchBuilder::PFPNACCOp},
{0x90, 1, &OpDispatchBuilder::VPFCMPOp<1>},
{0x90, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 1>},
{0x94, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryDuplicateOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x97, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, true>},
{0x9A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x9E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0xA0, 1, &OpDispatchBuilder::VPFCMPOp<2>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 2>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
// Can be treated as a move
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -32,7 +32,7 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, OpSize::i32Bit>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0xB0, 1, &OpDispatchBuilder::VPFCMPOp<0>},
{0xB0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 0>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
// Can be treated as a move
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -20,18 +20,18 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x03), 1, &OpDispatchBuilder::PHADDS},
{OPD(PF_38_NONE, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_66, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_66, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x10), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i8Bit>},
@@ -44,22 +44,22 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i32Bit>},
@@ -9,13 +9,13 @@ namespace FEXCore::IR {
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr DispatchTableEntry Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::VectorBlend<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::VectorBlend<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::VectorBlend<OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i16Bit>},
{OPD(REX, PF_3A_NONE, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(REX, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
@@ -24,17 +24,17 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x17), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::PINSROp<OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x21), 1, &OpDispatchBuilder::InsertPSOp},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, false>},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, false>},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, false>},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, false>},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
@@ -65,7 +65,7 @@ constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i32Bit>},
};
#undef PF_3A_NONE
@@ -69,12 +69,12 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
@@ -6,8 +6,7 @@ namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x07, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x06, 4, &OpDispatchBuilder::PermissionRestrictedOp},
{0x0B, 1, &OpDispatchBuilder::INTOp},
{0x0E, 1, &OpDispatchBuilder::X87EMMS},
@@ -44,7 +43,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xBE, 2, &OpDispatchBuilder::MOVSXOp},
{0xC0, 2, &OpDispatchBuilder::XADDOp},
{0xC3, 1, &OpDispatchBuilder::MOVGPRNTOp},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC8, 8, &OpDispatchBuilder::BSWAPOp},
@@ -56,10 +55,10 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::InsertMMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
@@ -71,7 +70,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
@@ -79,15 +78,15 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x70, 1, &OpDispatchBuilder::PSHUFW8ByteOp},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i8Bit>},
@@ -95,7 +94,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x77, 1, &OpDispatchBuilder::X87EMMS},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i32Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i32Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i32Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
@@ -116,9 +115,9 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -131,7 +130,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
@@ -152,23 +151,23 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x12, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, false>},
{0x16, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, false>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x6F, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, false>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQOp, OpDispatchBuilder::VectorOpType::SSE>},
@@ -176,36 +175,36 @@ constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0xB8, 1, &OpDispatchBuilder::PopcountOp},
{0xBC, 1, &OpDispatchBuilder::TZCNT},
{0xBD, 1, &OpDispatchBuilder::LZCNT},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, false>},
};
constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
// x52 = Invalid
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x78, 1, &OpDispatchBuilder::Insertq_imm},
{0x79, 1, &OpDispatchBuilder::Insertq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<false>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, false>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, false>},
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
@@ -217,10 +216,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
@@ -231,7 +230,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
@@ -239,15 +238,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x6C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
{0x6D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i64Bit>},
{0x6E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
@@ -260,15 +259,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x78, 1, nullptr}, // GROUP 17
{0x79, 1, &OpDispatchBuilder::Extrq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
{0x7F, 1, &OpDispatchBuilder::MOVVectorAlignedOp},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i64Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i32Bit>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i64Bit>},
@@ -288,10 +287,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, false>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -304,7 +303,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
File diff suppressed because it is too large. Load diff
@@ -163,7 +163,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
SubWithFlags(OpSize::i64Bit, Exponent, 0x7fff);
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
@@ -178,7 +178,8 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -623,13 +624,10 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
_F80FYL2XP1Stack();
} else {
_F80FYL2XStack();
}
_F80FYL2XStack();
}
void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
@@ -106,12 +106,24 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
bool CanUseFloatReg = Size == OpSize::i64Bit;
if (CanUseFloatReg) {
// If possible, it's faster to keep the data in an FPR than doing a GPR transfer.
if (Truncate) {
data = _Vector_FToZS(OpSize::i128Bit, OpSize::i64Bit, data);
} else {
data = _Vector_FToS(OpSize::i128Bit, OpSize::i64Bit, data);
}
StoreResultFPR_WithOpSize(Op, Op->Dest, data, OpSize::i64Bit, OpSize::i8Bit);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -370,6 +382,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// if val is +/-Inf, then sig = val, exp = +inf
// if val is NaN, then sig = val, exp = val
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
@@ -379,6 +393,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// Inf/NaN case
Ref ExpInfOnlyV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x7ff0'0000'0000'0000UL));
Ref ExpNanV = Node;
Ref SigInfV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
@@ -388,12 +407,24 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
// Mantissa non-zero => NaN (exp result = input); else Inf (exp result = +Inf)
Ref Mantissa = _And(OpSize::i64Bit, Gpr, Constant(0x000f'ffff'ffff'ffffULL));
_TestNZ(OpSize::i64Bit, Mantissa, Constant(~0ULL));
Ref ExpInfV = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfOnlyV, ExpNanV);
// Biased exponent == 0x7ff => Inf/NaN path, else non-zero-case.
Ref BiasedExp = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
SubWithFlags(OpSize::i64Bit, BiasedExp, 0x7ff);
Ref ExpNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfV, ExpNZV);
Ref SigNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigInfV, SigNZV);
// Zero folds on top.
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZOrInf);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZOrInf);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::iInvalid);
@@ -200,8 +200,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0}},
// Instructions
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -210,16 +210,16 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x06, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_06] }}},
{0x07, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_07] }}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x0E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_0E] }}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -227,8 +227,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x16, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_16] }}},
{0x17, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_17] }}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -236,24 +236,24 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x1E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1E] }}},
{0x1F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1F] }}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x27, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_27] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x2F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_2F] }}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -310,8 +310,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -34,7 +34,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> H0F3A_ArchSelect_LUT = {{
// ENTRY_1_3A_66_22
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::PINSROp<IR::OpSize::i64Bit> }},
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PINSROp, IR::OpSize::i64Bit> }},
},
}};
@@ -28,31 +28,31 @@ enum PrimaryGroup_LUT {
constexpr std::array<X86InstInfo[2], ENTRY_MAX> PrimaryGroup_ArchSelect_LUT = {{
{
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
@@ -66,23 +66,23 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr U16U8InfoStruct PrimaryGroupOpTable[] = {
// GROUP_1 | 0x80 | reg
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
// Duplicates the 0x80 opcode group
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_0] }}},
@@ -94,14 +94,14 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 6), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_6] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 7), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_7] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 2
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
@@ -161,8 +161,8 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
// GROUP 3
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 0), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 1), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 4), 1, X86InstInfo{"MUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 5), 1, X86InstInfo{"IMUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 1, X86InstInfo{"DIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -170,21 +170,21 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 0), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 1), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 4), 1, X86InstInfo{"MUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 5), 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 1, X86InstInfo{"DIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 7), 1, X86InstInfo{"IDIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// GROUP 4
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 2), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 5
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END | FLAGS_CALL , 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0}},
@@ -50,7 +50,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> SecondGroup_ArchSelect_LUT = {{
},
}};
constexpr auto SecondInstGroupOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
@@ -139,37 +139,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_8, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 9
@@ -179,7 +179,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
// CMPXCHG8B/16B works with all prefixes
// Tooling fails to decode CMPXCHG with prefix
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -188,7 +188,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -197,7 +197,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -206,7 +206,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -402,37 +402,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
// GROUP 16
// AMD documentation claims again that this entire group is n/a to prefix
// Tooling once again fails to disassemble oens with the prefix. Disable until proven otherwise
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
@@ -31,19 +31,19 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
{
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
}};
@@ -61,8 +61,8 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0x05, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_05] }}},
{0x06, 1, X86InstInfo{"CLTS", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x07, 1, X86InstInfo{"SYSRET", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x0A, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x0B, 1, X86InstInfo{"UD2", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY, 0}},
{0x0C, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
@@ -205,23 +205,23 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xA0, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A0] }}},
{0xA1, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A1] }}},
{0xA2, 1, X86InstInfo{"CPUID", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xA4, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xA5, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xA6, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xA8, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A8] }}},
{0xA9, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A9] }}},
{0xAA, 1, X86InstInfo{"RSM", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xAC, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xAD, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xAE, 1, X86InstInfo{"", TYPE_GROUP_15, FLAGS_NO_OVERLAY, 0}},
{0xAF, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB2, 1, X86InstInfo{"LSS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB4, 1, X86InstInfo{"LFS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB5, 1, X86InstInfo{"LGS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB6, 1, X86InstInfo{"MOVZX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
@@ -229,14 +229,14 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xB8, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{0xB9, 1, X86InstInfo{"", TYPE_GROUP_10, FLAGS_NO_OVERLAY, 0}},
{0xBA, 1, X86InstInfo{"", TYPE_GROUP_8, FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xBC, 1, X86InstInfo{"BSF", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBD, 1, X86InstInfo{"BSR", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBE, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xBF, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xC2, 1, X86InstInfo{"CMPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{0xC3, 1, X86InstInfo{"MOVNTI", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST, 0}},
{0xC4, 1, X86InstInfo{"PINSRW", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX | FLAGS_SF_SRC_GPR, 1}},
@@ -474,7 +474,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::VMOVSLDUPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, true>},
{OPD(1, 0b11, 0x12), 1, &OpDispatchBuilder::VMOVDDUPOp},
{OPD(1, 0b00, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
@@ -487,7 +487,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::VMOVSHDUPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, true>},
{OPD(1, 0b00, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
@@ -496,36 +496,36 @@ namespace AVX256 {
{OPD(1, 0b00, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b01, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i64Bit>},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{OPD(1, 0b01, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRECP, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{OPD(1, 0b01, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
@@ -541,42 +541,42 @@ namespace AVX256 {
{OPD(1, 0b00, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{OPD(1, 0b01, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{OPD(1, 0b01, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, true>},
{OPD(1, 0b01, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, true>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, true>},
{OPD(1, 0b00, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i8Bit>},
{OPD(1, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i16Bit>},
@@ -607,8 +607,8 @@ namespace AVX256 {
{OPD(1, 0b00, 0x77), 1, &OpDispatchBuilder::VZEROOp},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i32Bit>},
@@ -618,19 +618,19 @@ namespace AVX256 {
{OPD(1, 0b01, 0x7F), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x7F), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::VPINSRWOp},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xC5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(1, 0b00, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xD1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xD2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i32Bit>},
@@ -653,14 +653,14 @@ namespace AVX256 {
{OPD(1, 0b01, 0xE1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xE3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::VPMULHWOp<false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::VPMULHWOp<true>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, true>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, true>},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0xE8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{OPD(1, 0b01, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
@@ -671,11 +671,11 @@ namespace AVX256 {
{OPD(1, 0b01, 0xEE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSMAX, OpSize::i16Bit>},
{OPD(1, 0b01, 0xEF), 1, &OpDispatchBuilder::AVXVectorXOROp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b01, 0xF1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xF2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xF3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF5), 1, &OpDispatchBuilder::VPMADDWDOp},
{OPD(1, 0b01, 0xF6), 1, &OpDispatchBuilder::VPSADBWOp},
{OPD(1, 0b01, 0xF7), 1, &OpDispatchBuilder::MASKMOVOp},
@@ -689,8 +689,8 @@ namespace AVX256 {
{OPD(1, 0b01, 0xFE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
{OPD(2, 0b01, 0x00), 1, &OpDispatchBuilder::VPSHUFBOp},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x03), 1, &OpDispatchBuilder::VPHADDSWOp},
{OPD(2, 0b01, 0x04), 1, &OpDispatchBuilder::VPMADDUBSWOp},
@@ -698,14 +698,14 @@ namespace AVX256 {
{OPD(2, 0b01, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPHSUBOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x07), 1, &OpDispatchBuilder::VPHSUBSWOp},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::VPSIGN<OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::VPSIGN<OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::VPSIGN<OpSize::i32Bit>},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0B), 1, &OpDispatchBuilder::VPMULHRSWOp},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x13), 1, &OpDispatchBuilder::VCVTPH2PSOp},
{OPD(2, 0b01, 0x16), 1, &OpDispatchBuilder::VPERMDOp},
@@ -717,28 +717,28 @@ namespace AVX256 {
{OPD(2, 0b01, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(2, 0b01, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(2, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPACKUSOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x36), 1, &OpDispatchBuilder::VPERMDOp},
{OPD(2, 0b01, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
@@ -752,7 +752,7 @@ namespace AVX256 {
{OPD(2, 0b01, 0x3F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VUMAX, OpSize::i32Bit>},
{OPD(2, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::AVXPHMINPOSUWOp},
{OPD(2, 0b01, 0x45), 1, &OpDispatchBuilder::VPSRLVOp},
{OPD(2, 0b01, 0x46), 1, &OpDispatchBuilder::VPSRAVDOp},
{OPD(2, 0b01, 0x47), 1, &OpDispatchBuilder::VPSLLVOp},
@@ -764,13 +764,13 @@ namespace AVX256 {
{OPD(2, 0b01, 0x78), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i8Bit>},
{OPD(2, 0b01, 0x79), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i16Bit>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::VPMASKMOVOp<false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::VPMASKMOVOp<true>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, true>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x96), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, true, 1, 3, 2>}, // VFMADDSUB
{OPD(2, 0b01, 0x97), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, false, 1, 3, 2>}, // VFMSUBADD
@@ -820,10 +820,10 @@ namespace AVX256 {
{OPD(3, 0b01, 0x04), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x06), 1, &OpDispatchBuilder::VPERM2Op},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0C), 1, &OpDispatchBuilder::VPBLENDDOp},
{OPD(3, 0b01, 0x0D), 1, &OpDispatchBuilder::VBLENDPDOp},
{OPD(3, 0b01, 0x0E), 1, &OpDispatchBuilder::VPBLENDWOp},
@@ -837,15 +837,15 @@ namespace AVX256 {
{OPD(3, 0b01, 0x18), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x19), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x1D), 1, &OpDispatchBuilder::VCVTPS2PHOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::VPINSRBOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i8Bit>},
{OPD(3, 0b01, 0x21), 1, &OpDispatchBuilder::VINSERTPSOp},
{OPD(3, 0b01, 0x22), 1, &OpDispatchBuilder::VPINSRDQOp},
{OPD(3, 0b01, 0x38), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x39), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::VDPPOp<OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::VDPPOp<OpSize::i64Bit>},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x42), 1, &OpDispatchBuilder::VMPSADBWOp},
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::VPCLMULQDQOp},
@@ -855,10 +855,10 @@ namespace AVX256 {
{OPD(3, 0b01, 0x4B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i64Bit>},
{OPD(3, 0b01, 0x4C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i8Bit>},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, true>},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, true>},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, true>},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, true>},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
@@ -392,8 +392,11 @@ namespace InstFlags {
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SUPPORTS_LOCK = (1ULL << 31);
// Flags [57..32]: Undefined
// Flags [60..58]: Dst size
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
// Flags [63..61]: Src size
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
constexpr InstFlagType SIZE_MASK = 0b111;
+35 -10
View File
@@ -136,6 +136,7 @@
"u16": "uint16_t",
"u32": "uint32_t",
"u64": "uint64_t",
"c_str": "const char*",
"OpSize": "FEXCore::IR::OpSize",
"SSA": "OrderedNode*",
"GPR": "OrderedNode*",
@@ -240,6 +241,12 @@
"Desc": ["Debug operation that prints an SSA value to the console",
"May only print 64bits of the value"]
},
"PrintMsg c_str:$Value": {
"HasSideEffects": true,
"Desc": ["Debug operation that prints an string to the console.",
"This is for debug only! Will break code caching!"
]
},
"GPR = AllocateGPR i1:$ForPair": {
"Desc": ["Silly pseudo-instruction to allocate a register for a future destination",
"Note: if an instruction uses allocated destinations-as-sources,",
@@ -710,7 +717,7 @@
"HasSideEffects": true
},
"CacheLineClean GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline cleanat the address specified",
"Desc": ["Does a 64 byte cacheline clean at the address specified",
"Only cleans the data cachelines. Doesn't do any zeroing",
"Skips the invalidation step of the CacheLineClear operation"
],
@@ -2751,39 +2758,43 @@
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM1 FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2XP1 FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": true
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SIN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64COS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR:$Sin, FPR:$Cos = F64SINCOS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
@@ -3201,6 +3212,20 @@
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"FPR = F80FYL2XP1Stack": {
"Desc": [
"Computes ST1 * log2(1 + ST0)",
"Stores the result in ST1, and pops the top of the stack.",
"Returns the new value at the top of the stack, i.e. the result of the operation."
],
"HasSideEffects": true,
"DestSize": "OpSize::i128Bit",
"X87": true
},
"FPR = F80FYL2XP1 FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"F80VBSLStack OpSize:#RegisterSize, FPR:$VectorMask, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Does a vector bitwise select.",
+10 -2
View File
@@ -38,6 +38,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg)
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const char* const Arg) {
*out << fextl::fmt::format("'{}'", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
@@ -204,6 +208,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorCon
return "movmaskb";
case NamedVectorConstant::NAMED_VECTOR_MOVMASKB_UPPER:
return "movmaskb_upper";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP:
return "v256_mid_element_swap";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER:
return "v256_mid_element_swap_upper";
case NamedVectorConstant::NAMED_VECTOR_ZERO:
return "vectorzero";
case NamedVectorConstant::NAMED_VECTOR_X87_ONE:
@@ -356,8 +364,8 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
++CurrentIndent;
AddIndent();
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), HeaderOp->OriginalRIP, HeaderOp->BlockCount,
HeaderOp->NumHostInstructions);
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), +HeaderOp->OriginalRIP, +HeaderOp->BlockCount,
+HeaderOp->NumHostInstructions);
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
{
+7 -5
View File
@@ -21,15 +21,15 @@ class IREmitter {
public:
IREmitter(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, bool SupportsTSOImm9)
: DualListData {ThreadAllocator, 8 * 1024 * 1024}
, SupportsTSOImm9(SupportsTSOImm9) {
ReownOrClaimBuffer();
ResetWorkingList();
}
, SupportsTSOImm9(SupportsTSOImm9) {}
virtual ~IREmitter() = default;
void ReownOrClaimBuffer() {
DualListData.ReownOrClaimBuffer();
// Reset the working list on new buffer.
ResetWorkingList();
}
void DelayedDisownBuffer() {
@@ -39,7 +39,6 @@ public:
IRListView ViewIR() {
return IRListView(&DualListData);
}
void ResetWorkingList();
/**
* @name IR allocation routines
@@ -512,6 +511,9 @@ protected:
fextl::vector<Ref> CodeBlocks;
uint64_t Entry {};
bool SupportsTSOImm9 {};
private:
void ResetWorkingList();
};
} // namespace FEXCore::IR
@@ -119,11 +119,7 @@ class DualIntrusiveAllocatorThreadPool final : public DualIntrusiveAllocator {
public:
DualIntrusiveAllocatorThreadPool(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, size_t Size)
: DualIntrusiveAllocator {Size}
, PoolObject {ThreadAllocator, Size * 2} {
// Claim a buffer on allocation
PoolObject.ReownOrClaimBuffer();
}
, PoolObject {ThreadAllocator, Size * 2} {}
void ReownOrClaimBuffer() {
Data = PoolObject.ReownOrClaimBuffer();
List = Data + MemorySize;
@@ -190,12 +186,12 @@ public:
}
[[nodiscard]]
unsigned PostRA() const {
bool PostRA() const {
return GetHeader()->PostRA;
}
[[nodiscard]]
unsigned SpillSlots() const {
uint32_t SpillSlots() const {
return GetHeader()->SpillSlots;
}
+5 -10
View File
@@ -8,23 +8,18 @@ class CPUIDEmu;
struct HostFeatures;
} // namespace FEXCore
namespace FEXCore::Utils {
class IntrusivePooledAllocator;
}
namespace FEXCore::IR {
class Pass;
class RegisterAllocationPass;
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID);
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures&, OpSize GPROpSize);
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID);
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures&, OpSize GPROpSize);
namespace Validation {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
fextl::unique_ptr<Pass> CreateIRValidation();
} // namespace Validation
namespace Debug {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper();
fextl::unique_ptr<Pass> CreateIRDumper();
}
} // namespace FEXCore::IR
@@ -51,7 +51,7 @@ void IRDumper::Run(IREmitter* IREmit) {
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpToFile) {
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), +HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
FD = FEXCore::File::File(fileName.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
}
@@ -60,14 +60,14 @@ void IRDumper::Run(IREmitter* IREmit) {
fextl::stringstream out;
FEXCore::IR::Dump(&out, &IR);
if (FD.IsValid()) {
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
} else {
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
}
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper() {
fextl::unique_ptr<Pass> CreateIRDumper() {
return fextl::make_unique<IRDumper>();
}
} // namespace FEXCore::IR::Debug
@@ -271,7 +271,7 @@ void IRValidation::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation() {
fextl::unique_ptr<Pass> CreateIRValidation() {
return fextl::make_unique<IRValidation>();
}
} // namespace FEXCore::IR::Validation
@@ -747,7 +747,7 @@ void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination() {
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination() {
return fextl::make_unique<DeadFlagCalculationEliminination>();
}
@@ -781,7 +781,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
IR->GetHeader()->PostRA = true;
}
fextl::unique_ptr<IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID) {
fextl::unique_ptr<IR::Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID) {
return fextl::make_unique<ConstrainedRAPass>(CPUID);
}
} // namespace FEXCore::IR
@@ -188,7 +188,7 @@ private:
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
if (Features.SupportsSVE()) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
@@ -785,6 +785,12 @@ void X87StackOptimization::Run(IREmitter* Emit) {
break;
}
case OP_F80FYL2XP1STACK: {
HandleBinopStack(OP_F64FYL2XP1, false, OP_F80FYL2XP1, 1, 0, 1);
StackPop();
break;
}
case OP_F80ATANSTACK: {
HandleBinopStack(OP_F64ATAN, false, OP_F80ATAN, 1, 1, 0);
StackPop();
@@ -862,25 +868,20 @@ void X87StackOptimization::Run(IREmitter* Emit) {
Ref SinValue {};
Ref CosValue {};
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
}
#ifdef VIXL_SIMULATOR
if (DisableVixlIndirectCalls() == 0) {
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
} else {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
} else
else if (DisableVixlIndirectCalls() == 0) {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
#endif
{
else {
SinValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
CosValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
if (ReducedPrecisionMode) {
IREmit->_F64SINCOS(St0, SinValue, CosValue);
} else {
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
// Push values
@@ -1230,7 +1231,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
return;
}
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures& Features, OpSize GPROpSize) {
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures& Features, OpSize GPROpSize) {
return fextl::make_unique<X87StackOptimization>(Features, GPROpSize);
}
} // namespace FEXCore::IR
+36 -22
View File
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/Allocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
@@ -32,8 +33,8 @@ std::pmr::memory_resource* get_default_resource() {
}
} // namespace fextl::pmr
#ifndef _WIN32
namespace FEXCore::Allocator {
#ifndef _WIN32
MMAP_Hook mmap {::mmap};
MUNMAP_Hook munmap {::munmap};
@@ -113,27 +114,17 @@ FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
};
for (auto Bits : TLBSizes) {
uintptr_t Size = 1ULL << Bits;
// Just try allocating
// We can't actually determine VA size on ARM safely
auto Find = [](uintptr_t Size) -> bool {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void* Ptr = ::mmap(reinterpret_cast<void*>(Size - FEXCore::Utils::FEX_PAGE_SIZE * i), FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE,
MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
if (Ptr == (void*)(Size - FEXCore::Utils::FEX_PAGE_SIZE * i)) {
return true;
}
}
}
return false;
};
if (Find(Size)) {
HostVASize = Bits;
// We can't actually determine VA size on ARM safely.
// Instead, try allocating the page at the top of the range.
// If this succeeds OR the page is reported as already existing,
// we know we're in valid VA space. Otherwise, we must go lower.
void* Addr = reinterpret_cast<void*>((1ULL << Bits) - FEXCore::Utils::FEX_PAGE_SIZE);
void* Ptr = ::mmap(Addr, FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
}
if (Ptr != (void*)~0ULL || errno == EEXIST) {
return Bits;
}
}
@@ -261,9 +252,19 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
// Block remaining memory gaps
bool SupportsDontDump = true;
for (auto RegionIt = Regions.begin(); RegionIt != Regions.end(); ++RegionIt) {
auto Alloc = ::mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
if (SupportsDontDump) {
// Mark these regions as don't dump so that coredump doesn't try dumping large unmapped regions.
// Ideally coredump would be smart enough to only dump resident pages, but here we are.
auto Result = madvise(RegionIt->Ptr, RegionIt->Size, MADV_DONTDUMP);
if (Result == -1) {
SupportsDontDump = false;
}
}
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "StealMemoryRegion: mmap({}, {:x}) failed: {}", fmt::ptr(RegionIt->Ptr), RegionIt->Size, errno);
LogMan::Throw::AFmt(Alloc == RegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(RegionIt->Ptr));
}
@@ -304,5 +305,18 @@ void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child) {
Alloc64->UnlockAfterFork(Thread, Child);
}
}
} // namespace FEXCore::Allocator
#else
void VirtualNameNOP(const char*, const void*, size_t) {}
void VirtualTHPNOP(const void* Ptr, size_t Size, THPControl Control) {}
VirtualNamePtr VirtualName {VirtualNameNOP};
VirtualTHPPtr VirtualTHPControl {VirtualTHPNOP};
void SetupHooks(size_t PageSize, HookPtrs Ptrs) {
VirtualName = Ptrs.VirtualName;
VirtualTHPControl = Ptrs.VirtualTHPControl;
}
#endif
} // namespace FEXCore::Allocator
+2
View File
@@ -1,11 +1,13 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Allocator {
void InitializeAllocator(size_t PageSize);
void LockBeforeFork(FEXCore::Core::InternalThreadState* Thread);
void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child);
} // namespace FEXCore::Allocator
+3 -1
View File
@@ -2,7 +2,6 @@
#ifdef ENABLE_FEX_ALLOCATOR
#include <rpmalloc/rpmalloc.h>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#include <sys/mman.h>
#else
@@ -109,6 +108,9 @@ static void* FEX_rp_mmap(size_t size, size_t alignment, size_t* offset, size_t*
#define PR_SET_VMA_ANON_NAME 0
#endif
prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, ptr, map_size, global_config.page_name);
// Disable HUGEPAGE on allocation from rpmalloc.
madvise(ptr, map_size, MADV_NOHUGEPAGE);
}
if (ptr == nullptr) {
+55 -19
View File
@@ -2,7 +2,7 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Context/Context.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
@@ -343,11 +343,12 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
// 32bit
uint64_t Addr = GPRs[AddressReg];
// Lower register must be even, so only upper register can be 31.
uint32_t DesiredLower = GPRs[DesiredReg1];
uint32_t DesiredUpper = GPRs[DesiredReg2];
uint32_t DesiredUpper = DesiredReg2 == 31 ? 0 : GPRs[DesiredReg2];
uint32_t ExpectedLower = GPRs[ExpectedReg1];
uint32_t ExpectedUpper = GPRs[ExpectedReg2];
uint32_t ExpectedUpper = ExpectedReg2 == 31 ? 0 : GPRs[ExpectedReg2];
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
@@ -1352,7 +1353,9 @@ static std::optional<uint64_t> DoCAS(uint32_t Size, uint64_t Desired, uint64_t E
}
static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg, uint32_t* StrictSplitLockMutex) {
std::optional<uint64_t> Res = DoCAS(Size, GPRs[DesiredReg], GPRs[ExpectedReg], GPRs[AddressReg], StrictSplitLockMutex);
uint64_t Desired = DesiredReg == 31 ? 0 : GPRs[DesiredReg];
uint64_t Expected = ExpectedReg == 31 ? 0 : GPRs[ExpectedReg];
std::optional<uint64_t> Res = DoCAS(Size, Desired, Expected, GPRs[AddressReg], StrictSplitLockMutex);
if (!Res.has_value()) {
return false;
}
@@ -1384,6 +1387,8 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
uint8_t Op = (Instr >> 12) & 0xF;
uint64_t Source = SourceReg == 31 ? 0 : GPRs[SourceReg];
if (Size == 2) {
auto NOPExpected = [](uint16_t SrcVal, uint16_t) -> uint16_t {
return SrcVal;
@@ -1420,7 +1425,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS16<true>(GPRs[SourceReg],
auto Res = DoCAS16<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1465,7 +1470,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS32<true>(GPRs[SourceReg],
auto Res = DoCAS32<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1510,7 +1515,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS64<true>(GPRs[SourceReg],
auto Res = DoCAS64<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1572,9 +1577,11 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
uint64_t Addr = GPRs[AddressReg] + Offset;
constexpr bool DoRetry = false;
uint64_t Data = DataReg == 31 ? 0 : GPRs[DataReg];
if (Size == 2) {
DoCAS16<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint16_t SrcVal, uint16_t) -> uint16_t {
@@ -1589,7 +1596,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 4) {
DoCAS32<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint32_t SrcVal, uint32_t) -> uint32_t {
@@ -1604,7 +1611,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 8) {
DoCAS64<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint64_t SrcVal, uint64_t) -> uint64_t {
@@ -1834,6 +1841,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return Desired;
};
uint64_t Source = DataSourceReg == 31 ? 0 : GPRs[DataSourceReg];
if (Size == 2) {
using AtomicType = uint16_t;
CASDesiredFn<AtomicType> DesiredFunction {};
@@ -1852,7 +1860,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS16<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS16<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1879,7 +1887,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS32<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS32<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1906,7 +1914,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS64<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS64<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
if (AtomicFetch && ResultReg != 31) {
@@ -1947,8 +1955,24 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
uint32_t* StrictSplitLockMutex {CTX->Config.StrictInProcessSplitLocks ? &CTX->StrictSplitLockMutex : nullptr};
if (!IsJIT) [[unlikely]] {
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if ((Instr & ArchHelpers::Arm64::CASPAL_MASK) == ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (ArchHelpers::Arm64::HandleCASPAL(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & ArchHelpers::Arm64::CASAL_MASK) == ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (ArchHelpers::Arm64::HandleCASAL(GPRs, Instr, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0)) {
// Skip this instruction now
return 4;
@@ -1991,22 +2015,34 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
} else if ((Instr & ArchHelpers::Arm64::STLXR_MASK) == ArchHelpers::Arm64::STLXR_INST) { // STLXR*
uint32_t StatusReg = Instr << 11 >> 27;
// // Emulate exclusive store by validating the address and value against the last unaligned LDAXR*.
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || Size > Thread->ExclusiveStore.Size) {
uint32_t SizeBytes = 1u << Size;
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || SizeBytes > Thread->ExclusiveStore.Size) {
if (StatusReg != 31) {
GPRs[StatusReg] = 1;
}
return 4;
}
if (std::optional<uint64_t> Prev =
DoCAS(Size, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
DoCAS(SizeBytes, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
if (StatusReg != 31) {
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, Size);
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, SizeBytes);
}
Thread->ExclusiveStore.Size = 0;
return 4;
}
} else if ((Instr & ArchHelpers::Arm64::ATOMIC_MEM_MASK) == ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (ArchHelpers::Arm64::HandleAtomicMemOp(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: 0x{:x} Instruction: 0x{:08x}\n", Op, ProgramCounter, PC[0]);
return std::nullopt;
}
}
return 0;
LogMan::Msg::EFmt("Unhandled non-JIT atomic");
return std::nullopt;
}
const auto Frame = Thread->CurrentFrame;
@@ -32,7 +32,7 @@ public:
// Differs from Itanium specification
LOGMAN_THROW_A_FMT(PMF.adj == 0, "C++ Pointer-To-Member representation didn't have adj == 0. Are you trying to cast a virtual member?");
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
return PMF.ptr;
}
@@ -54,7 +54,7 @@ public:
"members.");
return PMF.ptr;
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
}
+3 -3
View File
@@ -1,11 +1,11 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
namespace FEXCore::Utils::SpinWaitLock {
#ifdef ARCHITECTURE_arm64
constexpr uint64_t NanosecondsInSecond = 1'000'000'000ULL;
static uint32_t GetCycleCounterFrequency() {
static uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -21,7 +21,7 @@ static uint64_t CalculateCyclesPerNanosecond() {
return NanosecondsInSecond / CounterFrequency;
}
uint32_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CyclesPerNanosecond = CalculateCyclesPerNanosecond();
#endif
} // namespace FEXCore::Utils::SpinWaitLock
+23
View File
@@ -0,0 +1,23 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/WildcardMatcher.h>
namespace FEXCore::Utils::Wildcard {
static bool matchHelper(std::string_view pattern, std::string_view text, size_t p_idx, size_t t_idx) {
if (p_idx == pattern.size()) {
// Pattern exhausted
return (t_idx == text.size());
} else if (pattern[p_idx] == '*') {
// Wildcard: Try matching zero characters, or one or more characters
return matchHelper(pattern, text, p_idx + 1, t_idx) || (t_idx < text.size() && matchHelper(pattern, text, p_idx, t_idx + 1));
} else {
// Match normally
return (t_idx < text.size() && pattern[p_idx] == text[t_idx] && matchHelper(pattern, text, p_idx + 1, t_idx + 1));
}
}
bool Matches(std::string_view pattern, std::string_view text) {
return matchHelper(pattern, text, 0, 0);
}
} // namespace FEXCore::Utils::Wildcard
+91 -8
View File
@@ -8,6 +8,7 @@
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <atomic>
#include <cstdint>
@@ -27,7 +28,12 @@ namespace HLE {
struct SourcecodeMap;
} // namespace HLE
enum class GuestRelocationType : uint32_t { Rel32, Rel64 };
enum class GuestRelocationType : uint32_t {
Rel32,
Rel64,
// Skip blocks containing this relocation
Skip,
};
// Generic information associated with an executable file.
struct ExecutableFileInfo {
@@ -173,7 +179,55 @@ private:
CodeMapOpener& FileOpener;
};
class AbstractCodeCache;
/**
* Manages runtime state associated with a mapped code cache file.
*
* The mapped file pointer is managed by the frontend and must be valid
* throughout the lifetime of this object.
*/
struct MappedCodeCacheFile {
// Calls UnregisterMappedCodeBuffer internally, see its docstring about synchronization requirements
~MappedCodeCacheFile();
// If not nullptr, the MappedCodeCacheFile will be unregistered from this on destruction
AbstractCodeCache* CacheManager;
std::span<std::byte> MappedFile; // Mapped data of the whole cache file
std::span<std::byte> CodeBufferInFile; // Subspan of cached ARM64 data within MappedFile (pre-relocation)
std::span<std::byte> CodeBuffer; // Cached ARM64 data used for execution (post-relocation; owned by MappedCodeCacheFile)
std::byte* BlockListInFile; // Pointer to BlockListEntry data within MappedFile
uint32_t NumBlocks; // Number of BlockListEntry objects
uint32_t NumCodePages; // Number of code page entrypoint mappings
struct PageRelocationRange {
uint32_t Offset; // In bytes from start of file
uint32_t Length; // Number of relocations
};
// List of relocation ranges in the mapped cache file, grouped by the code page they apply to.
// This vector is indexed by the relative page offset from the start of the ARM64 code data.
//
// For example PageRelocationRanges[1] == { 0x100, 0x20 } means:
// - there are 0x20 bytes of relocation data at offset 0x100 in the cache file
// - these 0x20 bytes of relocation data will patch data at CodeBuffer[0x1000..0x2000]
fextl::vector<PageRelocationRange> PageRelocationRanges;
fextl::vector<bool> LoadedPages;
uint64_t GuestBase {}; // Guest base address for relocation application
// Helper member to prevent moving/copying without disallowing aggregate-construction
std::atomic<int> disallow_copy_or_move;
size_t NumPages() const {
return CodeBuffer.size_bytes() / FEXCore::Utils::FEX_PAGE_SIZE;
}
};
class AbstractCodeCache {
fextl::vector<std::span<std::byte>> MappedCodeBuffers;
public:
virtual ~AbstractCodeCache() = default;
@@ -185,13 +239,6 @@ public:
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
* Returns false if the provided cache file is invalid, and true otherwise.
*/
virtual bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) = 0;
/**
* Bundles the current Core state (CodeBuffer, GuestToHostMapping, ...) to a code cache and writes it to the given file descriptor.
* Returns true on success.
@@ -202,6 +249,42 @@ public:
* Function to be called before compiling any code for caching purposes
*/
virtual void InitiateCacheGeneration() = 0;
/**
* Loads a code cache from mapped memory.
*
* Code sections must be enabled in a second step (see EnableLoadedSection).
* Afterwards, individual code pages must be finalized using FinalizeCodePages.
*
* On success, this returns a MappedCodeCacheFile that must be kept alive
* as long the cache is in use.
*/
virtual fextl::unique_ptr<MappedCodeCacheFile> LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo&, uint64_t FileStartVA) = 0;
/**
* Registers cached blocks for the given file section to the LookupCache.
*
* Also runs extended cache validation if enabled.
*/
virtual bool EnableLoadedSection(Core::InternalThreadState*, MappedCodeCacheFile&, const ExecutableFileSectionInfo&) = 0;
/**
* Extend the given code range so that it can be safely finalized.
*
* This is required for example to avoid dangling page-crossing FEX relocations on the edges
*
* StartPage and EndPage a 0-based relative page offsets into the cached code.
*/
static std::span<std::byte> SelectCodeRangeToFinalize(MappedCodeCacheFile&, size_t StartPage, size_t EndPage);
/**
* Finalize code pages in the given range (see SelectCodePagesToFinalize) for execution.
*/
virtual void FinalizeCodePages(MappedCodeCacheFile&, std::span<std::byte> CodeRange) = 0;
void RegisterMappedCodeBuffer(MappedCodeCacheFile&);
void UnregisterMappedCodeBuffer(MappedCodeCacheFile&);
bool IsAddressInMappedCodeBuffer(uintptr_t Address) const;
};
} // namespace FEXCore
+3 -2
View File
@@ -13,10 +13,10 @@
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/WritePriorityMutex.h>
namespace FEXCore {
struct HostFeatures;
class ForkableSharedMutex;
class ThunkHandler;
} // namespace FEXCore
@@ -73,6 +73,7 @@ public:
*/
FEX_DEFAULT_VISIBILITY virtual void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) = 0;
FEX_DEFAULT_VISIBILITY virtual bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) = 0;
FEX_DEFAULT_VISIBILITY virtual void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) = 0;
FEX_DEFAULT_VISIBILITY virtual void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) = 0;
@@ -143,7 +144,7 @@ public:
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) = 0;
+13
View File
@@ -302,6 +302,7 @@ enum FallbackHandlerIndex {
OPINDEX_F80MUL,
OPINDEX_F80DIV,
OPINDEX_F80FYL2X,
OPINDEX_F80FYL2XP1,
OPINDEX_F80ATAN,
OPINDEX_F80FPREM1,
OPINDEX_F80FPREM,
@@ -315,6 +316,7 @@ enum FallbackHandlerIndex {
OPINDEX_F64ATAN,
OPINDEX_F64F2XM1,
OPINDEX_F64FYL2X,
OPINDEX_F64FYL2XP1,
OPINDEX_F64FPREM,
OPINDEX_F64FPREM1,
OPINDEX_F64SCALE,
@@ -337,6 +339,7 @@ struct JITPointers {
// Process specific
uint64_t PrintValue {};
uint64_t PrintVectorValue {};
uint64_t PrintMsgValue {};
uint64_t ThreadRemoveCodeEntryFromJIT {};
uint64_t CPUIDObj {};
uint64_t CPUIDFunction {};
@@ -375,6 +378,16 @@ struct JITPointers {
uint64_t L2Pointer {};
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t F64SinHandler {};
uint64_t F64CosHandler {};
uint64_t F64TanHandler {};
uint64_t F64F2XM1Handler {};
uint64_t F64ScaleHandler {};
uint64_t F64AtanHandler {};
uint64_t F64FYL2XHandler {};
uint64_t F64FYL2XP1Handler {};
uint64_t F64FPREMHandler {};
uint64_t F64FPREM1Handler {};
/** @} */
// Copy of process-wide named vector constants data.
+24 -6
View File
@@ -5,13 +5,20 @@
#include <cstdint>
namespace FEXCore {
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
struct HostFeatures {
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
// Whether or not the host supports any kind of SVE implementation.
[[nodiscard]]
bool SupportsSVE() const {
return SupportsSVE128 || SupportsSVE256;
}
uint32_t DCacheLineSize {};
uint32_t ICacheLineSize {};
bool SupportsCacheMaintenanceOps {};
@@ -41,11 +48,22 @@ struct HostFeatures {
bool SupportsWFXT {};
bool Supports3DNow {};
bool SupportsSSE4a {};
bool SupportsMOPS {};
bool PreferZVAForVZero {};
// Float exception behaviour
bool SupportsAFP {};
bool SupportsFloatExceptions {};
// Changes code generation slightly.
enum class HostTypeEnum {
Unknown,
Linux,
Wow64,
Arm64ec,
};
HostTypeEnum HostType {};
// Flag if this is InstCountCI
bool IsInstCountCI {};
@@ -15,7 +15,6 @@
namespace FEXCore {
class LookupCache;
class CompileService;
struct JITSymbolBuffer;
} // namespace FEXCore
@@ -102,10 +101,6 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
NonMovableUniquePtr<FEXCore::IR::PassManager> PassManager;
NonMovableUniquePtr<JITSymbolBuffer> SymbolBuffer;
std::shared_ptr<FEXCore::CompileService> CompileService;
std::shared_mutex ObjectCacheRefCounter {};
// This pointer is owned by the frontend.
FEXCore::SHMStats::ThreadStats* ThreadStats {};
@@ -130,9 +125,9 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
alignas(FEXCore::Utils::FEX_PAGE_SIZE) uint8_t InterruptFaultPage[FEXCore::Utils::FEX_PAGE_SIZE];
};
static_assert(std::is_standard_layout_v<FEXCore::Core::InternalThreadState>);
static_assert((offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) <
FEXCore::Utils::FEX_PAGE_SIZE,
"Fault page is outside of immediate range from CPU state");
static_assert(sizeof(FEXCore::Core::InternalThreadState) == (FEXCore::Utils::FEX_PAGE_SIZE * 2));
// Maximum unsigned-offset store range for fault page.
static_assert(
(offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) <= 65520,
"Fault page is outside of immediate range from CPU state");
} // namespace FEXCore::Core
+6
View File
@@ -34,6 +34,12 @@ enum NamedVectorConstant : uint8_t {
NAMED_VECTOR_MOVMASKB,
NAMED_VECTOR_MOVMASKB_UPPER,
// Used to swap [0, 1, 2, 3] into [0, 2, 1, 3] in lieu
// of Q operations introduced in SVE2.1. Can be removed when
// such operations become available.
NAMED_VECTOR_256_MID_ELEMENT_SWAP,
NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER,
NAMED_VECTOR_X87_ONE,
NAMED_VECTOR_X87_LOG2_10,
NAMED_VECTOR_X87_LOG2_E,
@@ -12,9 +12,6 @@ struct InternalThreadState;
}
namespace FEXCore::Allocator {
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize);
FEX_DEFAULT_VISIBILITY void ClearHooks();
FEX_DEFAULT_VISIBILITY size_t DetermineVASize();
#ifdef GLIBC_ALLOCATOR_FAULT
+24 -3
View File
@@ -27,6 +27,24 @@ enum class ProtectOptions : uint32_t {
};
FEX_DEF_NUM_OPS(ProtectOptions)
enum class THPControl {
Enable,
Disable,
};
#ifndef _WIN32
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize);
#else
using VirtualNamePtr = void (*)(const char*, const void*, size_t);
using VirtualTHPPtr = void (*)(const void*, size_t, THPControl);
struct HookPtrs {
VirtualNamePtr VirtualName;
VirtualTHPPtr VirtualTHPControl;
};
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize, HookPtrs Ptrs);
#endif
FEX_DEFAULT_VISIBILITY void ClearHooks();
#ifdef _WIN32
inline void* VirtualAlloc(void* Base, size_t Size, bool Execute = false, bool Commit = true) {
// Allocate top-down to avoid polluting the lower VA space, as even on 64-bit some programs (i.e. LuaJIT) require allocations below 4GB.
@@ -82,8 +100,8 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
return ::VirtualProtect(Ptr, Size, prot, nullptr) == 0;
}
inline void VirtualName(const char*, void*, size_t) {}
FEX_DEFAULT_VISIBILITY extern VirtualNamePtr VirtualName;
FEX_DEFAULT_VISIBILITY extern VirtualTHPPtr VirtualTHPControl;
#else
using MMAP_Hook = void* (*)(void*, size_t, int, int, int, off_t);
using MUNMAP_Hook = int (*)(void*, size_t);
@@ -123,6 +141,10 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
return ::mprotect(Ptr, Size, prot) == 0;
}
inline void VirtualTHPControl(const void* Ptr, size_t Size, THPControl Control) {
::madvise(const_cast<void*>(Ptr), Size, Control == THPControl::Enable ? MADV_HUGEPAGE : MADV_NOHUGEPAGE);
}
#endif
// Memory allocation routines to be defined externally.
@@ -142,7 +164,6 @@ void aligned_free(void* ptr);
FEX_DEFAULT_VISIBILITY extern void InitializeThread();
#ifndef _WIN32
void InitializeAllocator(size_t PageSize);
void SetupAllocatorHooks(void* (*)(void* addr, size_t length, int prot, int flags, int fd, off_t offset), int (*)(void* addr, size_t length));
#endif
@@ -28,6 +28,9 @@
// then program behavior is undefined.
#define FEX_UNREACHABLE __builtin_unreachable()
// Like offsetof but for array members with a dynamic element index
#define ARRAY_OFFSETOF(Type, ArrayMember, Index) (offsetof(Type, ArrayMember) + sizeof(Type::ArrayMember[0]) * (Index))
namespace FEXCore::Assert {
// This function can not be inlined
[[noreturn]]
@@ -2,7 +2,6 @@
#pragma once
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/mman.h>
#include <sys/user.h>
#include <sys/prctl.h>
@@ -35,6 +35,9 @@ public:
const auto Result = pthread_mutex_lock(&Mutex);
LOGMAN_THROW_A_FMT(Result == 0, "{} failed to lock with {}", __func__, Result);
}
bool try_lock() {
return pthread_mutex_trylock(&Mutex) == 0;
}
void unlock() {
const auto Result = pthread_mutex_unlock(&Mutex);
LOGMAN_THROW_A_FMT(Result == 0, "{} failed to unlock with {}", __func__, Result);
@@ -50,8 +50,9 @@ namespace FEXCore::Utils::SpinWaitLock {
#define SPINLOOP_32BIT SPINLOOP_BODY(ldar, w)
#define SPINLOOP_64BIT SPINLOOP_BODY(ldar, x)
extern uint32_t CycleCounterFrequency;
extern uint64_t CyclesPerNanosecond;
FEX_DEFAULT_VISIBILITY extern uint64_t CycleCounterFrequency;
FEX_DEFAULT_VISIBILITY extern uint64_t CyclesPerNanosecond;
///< Get the raw cycle counter which is synchronizing.
/// `CNTVCTSS_EL0` also does the same thing, but requires the FEAT_ECV feature.
@@ -0,0 +1,8 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <string_view>
namespace FEXCore::Utils::Wildcard {
bool Matches(std::string_view pattern, std::string_view text);
} // namespace FEXCore::Utils::Wildcard
@@ -9,11 +9,15 @@
#include <unistd.h>
#else
#include <synchapi.h>
// Don't pull in all WIN32 headers for INFINITE. Causes too many problems.
#ifndef INFINITE
#define INFINITE 0xffffffff
#endif
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
namespace FEXCore::Utils::WritePriorityMutex {
+1 -1
View File
@@ -1,5 +1,5 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <catch2/catch_test_macros.hpp>
#include <chrono>
#include <thread>
+5 -5
View File
@@ -361,6 +361,11 @@ def IsSupportedKernel():
return version_check(GetKernelVersion()) >= version_check("5.15")
def main():
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
# Only run on supported arch
if not IsSupportedArch():
print ( "{} is not a supported architecture".format(GetArch()))
@@ -371,11 +376,6 @@ def main():
print ( "Kernel {} is too old. FEX needs 5.15 minimum".format(GetKernelVersion()))
ExitWithStatus(-1)
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
if GetDistro()[0] == "ubuntu":
print ("Getting PPA status: {}".format(("NotInstalled", "Installed")[GetPPAStatus()]))
+2
View File
@@ -59,6 +59,7 @@ class HostFeatures(Flag) :
FEATURE_LRCPC = (1 << 14)
FEATURE_LRCPC2 = (1 << 15)
FEATURE_FRINTTS = (1 << 16)
FEATURE_MOPS = (1 << 17)
HostFeaturesLookup = {
"SVE128" : HostFeatures.FEATURE_SVE128,
@@ -78,6 +79,7 @@ HostFeaturesLookup = {
"LRCPC" : HostFeatures.FEATURE_LRCPC,
"LRCPC2" : HostFeatures.FEATURE_LRCPC2,
"FRINTTS" : HostFeatures.FEATURE_FRINTTS,
"MOPS" : HostFeatures.FEATURE_MOPS,
}
def GetHostFeatures(data):
+58 -37
View File
@@ -7,11 +7,12 @@
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/WildcardMatcher.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/SymlinkChecks.h>
#include <cstring>
#include <fmt/format.h>
#include <functional>
@@ -28,7 +29,8 @@
namespace FEX::Config {
namespace JSON {
static void LoadJSonConfig(const fextl::string& Config, std::function<void(const char* Name, const char* ConfigSring)> Func) {
static void LoadJSonConfig(const fextl::string& Config, std::optional<fextl::string> AppName,
std::function<void(const char* Name, const char* ConfigString)> Func) {
fextl::vector<char> Data;
if (!FEXCore::FileLoading::LoadFile(Data, Config)) {
return;
@@ -48,21 +50,35 @@ namespace JSON {
return;
}
for (const json_t* ConfigItem = json_getChild(ConfigList); ConfigItem != nullptr; ConfigItem = json_getSibling(ConfigItem)) {
const char* ConfigName = json_getName(ConfigItem);
const char* ConfigString = json_getValue(ConfigItem);
fextl::vector<const json_t*> ConfigBlocks;
ConfigBlocks.push_back(ConfigList);
if (!ConfigName) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get config name for an item", Config);
return;
if (AppName) {
const json_t* OverrideList = json_getProperty(json, "AppOverrides");
if (OverrideList) {
for (const json_t* Item = json_getChild(OverrideList); Item != nullptr; Item = json_getSibling(Item)) {
const char* AppPattern = json_getName(Item);
// Find the first match, then break
if (FEXCore::Utils::Wildcard::Matches(AppPattern, *AppName)) {
ConfigBlocks.push_back(Item);
break;
}
}
}
}
if (!ConfigString) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get value for config item '{}'", Config, ConfigName);
return;
for (auto ConfigBlock : ConfigBlocks) {
for (const json_t* ConfigItem = json_getChild(ConfigBlock); ConfigItem != nullptr; ConfigItem = json_getSibling(ConfigItem)) {
const char* ConfigName = json_getName(ConfigItem);
const char* ConfigString = json_getValue(ConfigItem);
if (!ConfigString) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get value for config item '{}'", Config, ConfigName);
return;
}
Func(ConfigName, ConfigString);
}
Func(ConfigName, ConfigString);
}
}
} // namespace JSON
@@ -167,22 +183,24 @@ protected:
class MainLoader final : public OptionMapper {
public:
explicit MainLoader(FEXCore::Config::LayerType Type);
explicit MainLoader(fextl::string ConfigFile);
explicit MainLoader(FEXCore::Config::LayerType Type, std::optional<fextl::string> AppName = std::nullopt);
explicit MainLoader(fextl::string ConfigFile, std::optional<fextl::string> AppName = std::nullopt);
explicit MainLoader(FEXCore::Config::LayerType Type, std::string_view ConfigFile);
void Load() override;
private:
std::optional<fextl::string> AppName;
fextl::string Config;
};
class AppLoader final : public OptionMapper {
public:
explicit AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type);
explicit AppLoader(const fextl::string& AppName, FEXCore::Config::LayerType Type);
void Load();
private:
const fextl::string AppName;
fextl::string Config;
};
@@ -221,12 +239,14 @@ void OptionMapper::MapNameToOption(const char* ConfigName, const char* ConfigStr
#include <FEXCore/Config/ConfigOptions.inl>
}
MainLoader::MainLoader(FEXCore::Config::LayerType Type)
MainLoader::MainLoader(FEXCore::Config::LayerType Type, std::optional<fextl::string> AppName)
: OptionMapper(Type)
, AppName {AppName}
, Config {FEXCore::Config::GetConfigFileLocation(Type == FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN)} {}
MainLoader::MainLoader(fextl::string ConfigFile)
MainLoader::MainLoader(fextl::string ConfigFile, std::optional<fextl::string> AppName)
: OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, AppName {AppName}
, Config {std::move(ConfigFile)} {}
@@ -236,13 +256,14 @@ MainLoader::MainLoader(FEXCore::Config::LayerType Type, std::string_view ConfigF
void MainLoader::Load() {
SetCurrentConfigFile(Config);
JSON::LoadJSonConfig(Config, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
JSON::LoadJSonConfig(Config, AppName, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
}
AppLoader::AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type)
: OptionMapper(Type) {
AppLoader::AppLoader(const fextl::string& AppName, FEXCore::Config::LayerType Type)
: OptionMapper(Type)
, AppName {AppName} {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP || Type == FEXCore::Config::LayerType::LAYER_GLOBAL_APP;
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
Config = FEXCore::Config::GetApplicationConfig(AppName, Global);
// Immediately load so we can reload the meta layer
Load();
@@ -250,7 +271,7 @@ AppLoader::AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType T
void AppLoader::Load() {
SetCurrentConfigFile(Config);
JSON::LoadJSonConfig(Config, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
JSON::LoadJSonConfig(Config, AppName, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
}
EnvLoader::EnvLoader(char* const _envp[])
@@ -320,11 +341,11 @@ fextl::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer() {
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN);
}
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File) {
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File, std::optional<fextl::string> AppName) {
if (File) {
return fextl::make_unique<MainLoader>(*File);
return fextl::make_unique<MainLoader>(*File, std::move(AppName));
} else {
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN);
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN, std::move(AppName));
}
}
@@ -461,7 +482,7 @@ void LoadConfig(fextl::string ProgramName, char** const envp, const PortableInfo
if (!IsPortable) {
FEXCore::Config::AddLayer(CreateGlobalMainLayer());
}
FEXCore::Config::AddLayer(CreateMainLayer());
FEXCore::Config::AddLayer(CreateMainLayer(nullptr, ProgramName.empty() ? std::nullopt : std::optional {ProgramName}));
if (!ProgramName.empty()) {
if (!IsPortable) {
@@ -633,17 +654,10 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
}
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && (Global || !ConfigOverride)) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
} else if (PortableInfo.IsPortable && ConfigOverride && !Global) {
if (PortableInfo.IsPortable && Global) {
return fextl::fmt::format("{}/../share/fex-emu/", PortableInfo.InterpreterPath);
} else if (ConfigOverride && !Global) {
fextl::string AppConfigStr = ConfigOverride;
if (FHU::Filesystem::IsRelative(AppConfigStr)) {
AppConfigStr = PortableInfo.InterpreterPath + AppConfigStr;
@@ -652,6 +666,13 @@ fextl::string GetConfigDirectory(bool Global, const PortableInformation& Portabl
return AppConfigStr;
}
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
fextl::string ConfigDir;
if (Global) {
return GLOBAL_DATA_DIRECTORY;
+1 -1
View File
@@ -81,7 +81,7 @@ fextl::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer();
*
* @return unique_ptr for that layer
*/
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File = nullptr);
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File = nullptr, std::optional<fextl::string> AppName = std::nullopt);
fextl::unique_ptr<FEXCore::Config::Layer> CreateUserOverrideLayer(std::string_view AppConfig);
/**
+24 -10
View File
@@ -17,9 +17,9 @@
#include <fcntl.h>
#include <linux/limits.h>
#include <unistd.h>
#include <sys/poll.h>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/types.h>
@@ -137,21 +137,31 @@ fextl::string GetServerSocketName() {
return ServerSocketPath;
}
fextl::string GetServerSocketPath() {
fextl::string GetServerSocketPath(bool ForceTmp) {
fextl::string name {};
fextl::string Folder {};
#ifndef FEX_STEAM_SUPPORT
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
name = ServerSocketPath();
if (!ForceTmp) {
name = ServerSocketPath();
if (name.starts_with("/")) {
return name;
if (name.starts_with("/")) {
return name;
}
}
auto Folder = GetTempFolder();
Folder = GetTempFolder();
#else
// Under Steam the FEXServer's socket is a game-specific directory.
auto Folder = GetServerLockFolder();
if (ForceTmp) {
// If we're forcing temporary directory usage then the server socket path has exceeded sun_path 108 byte limit.
// Let's be a bit nice and put some more metadata in the server socket path.
const auto SteamID = getenv("SteamAppId") ?: "";
return fextl::fmt::format("{}/{}.FEXServer.Socket", GetTempFolder(), SteamID);
} else {
// Under Steam the FEXServer's socket is a game-specific directory.
Folder = GetServerLockFolder();
}
#endif
if (name.empty()) {
@@ -203,7 +213,11 @@ int ConnectToServer(ConnectionOption ConnectionOption) {
// Try again with a path-based socket, since abstract sockets will fail if we have been
// placed in a new netns as part of a sandbox.
auto ServerSocketPath = GetServerSocketPath();
auto ServerSocketPath = GetServerSocketPath(false);
if (ServerSocketPath.size() > sizeof(sockaddr_un::sun_path) - 1) {
LogMan::Msg::EFmt("Socket path '{}' too large for Unix domain sockets. Moving to tmp", ServerSocketPath);
ServerSocketPath = FEXServerClient::GetServerSocketPath(true);
}
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketPath.size(), sizeof(addr.sun_path) - 1);
+1 -1
View File
@@ -64,7 +64,7 @@ fextl::string GetServerRootFSLockFile();
fextl::string GetTempFolder();
fextl::string GetServerMountFolder();
fextl::string GetServerSocketName();
fextl::string GetServerSocketPath();
fextl::string GetServerSocketPath(bool ForceTmp);
int GetServerFD();
bool SetupClient(std::string_view InterpreterPath);
+55 -39
View File
@@ -502,6 +502,7 @@ static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEW
ENABLE_DISABLE_OPTION(SupportsWFXT, WFXT, WFXT);
ENABLE_DISABLE_OPTION(Supports3DNow, 3DNOW, 3DNOW);
ENABLE_DISABLE_OPTION(SupportsSSE4a, SSE4A, SSE4A);
ENABLE_DISABLE_OPTION(SupportsMOPS, MOPS, MOPS);
GET_SINGLE_OPTION(Crypto, CRYPTO);
#undef ENABLE_DISABLE_OPTION
@@ -538,6 +539,9 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
constexpr uint32_t PartNum_Oryon3 = 0x002;
constexpr uint32_t Implementer_Ampere = 0xc0;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
@@ -551,7 +555,7 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef ARCHITECTURE_arm64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
if (MIDR_Implementer == Implementer_QCOM && (MIDR_PartNum == PartNum_Oryon1 || MIDR_PartNum == PartNum_Oryon3)) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
@@ -591,6 +595,15 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
break;
}
}
if (MIDR_Implementer == Implementer_Ampere) {
// Ampere Computing CPUs that support CLZero should prefer using `dc zva` for vzero{upper,all} as its faster there.
// For Cortex CPUs it doesn't matter one way or the other.
// For Oryon CPUs, it is dramatically faster to avoid `dc zva` as it has dramatic stalls around barriers and overlapping `dc zva`.
//
// Because the `dc zva` optimization was implemented for Ampere, only use that path on the hardware.
HostFeatures->PreferZVAForVZero = HostFeatures->SupportsCLZERO;
}
}
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
@@ -625,11 +638,50 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
// Hardcode enable SVE with 256-bit wide registers.
HostFeatures.SupportsSVE128 = ForceSVEWidth() ? ForceSVEWidth() >= 128 : true;
HostFeatures.SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
HostFeatures.SupportsMOPS = true;
// Simulator has a hardcoded ZVA size of 64-bytes.
HostFeatures.SupportsCLZERO = true;
HostFeatures.SupportsAES = true;
HostFeatures.SupportsCRC = true;
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
HostFeatures.SupportsSVE128 = Features.Supports(CPUFeatures::Feature::SVE2);
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && Features.GetSVEVectorLengthInBits() >= 256;
HostFeatures.SupportsMOPS = Features.Supports(CPUFeatures::Feature::MOPS);
// Check if we can support cacheline clears
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsAES256 = HostFeatures.SupportsAVX && HostFeatures.SupportsAES;
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HostFeatures.PreferZVAForVZero = false;
if (CTR) {
HostFeatures.DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
HostFeatures.ICacheLineSize = 4 << (CTR & 0xF);
} else {
HostFeatures.DCacheLineSize = 64;
HostFeatures.ICacheLineSize = 64;
}
if (!HostFeatures.SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _WIN32
// Disable 3DNow! by default to better match the set of extensions exposed on modern CPUs.
@@ -640,12 +692,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.Supports3DNow = true;
#endif
HostFeatures.SupportsAES256 = HostFeatures.SupportsAVX && HostFeatures.SupportsAES;
if (!HostFeatures.SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef ARCHITECTURE_arm64
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
@@ -666,36 +712,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
SetFPCR(OriginalFPCR);
#endif
#ifdef VIXL_SIMULATOR
// simulator has a hardcoded ZVA size of 64-bytes.
HostFeatures.SupportsCLZERO = true;
HostFeatures.SupportsAES = true;
HostFeatures.SupportsCRC = true;
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
// Check if we can support cacheline clears
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
if (CTR) {
HostFeatures.DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
HostFeatures.ICacheLineSize = 4 << (CTR & 0xF);
} else {
HostFeatures.DCacheLineSize = HostFeatures.ICacheLineSize = 64;
}
#if defined(ARCHITECTURE_x86_64) && !defined(VIXL_SIMULATOR)
FEX::X86::Features Feature {};
HostFeatures.SupportsAES = Feature.Feat_aes;
@@ -712,7 +728,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.SupportsAFP = true;
HostFeatures.SupportsFloatExceptions = true;
#endif
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HandleErrata(&HostFeatures, MIDR);
OverrideFeatures(&HostFeatures, ForceSVEWidth());
@@ -739,7 +754,7 @@ FEXCore::HostFeatures FetchHostFeatures() {
uint64_t CTR = 0;
uint64_t MIDR = 0;
#ifdef ARCHITECTURE_arm64
#if defined(ARCHITECTURE_arm64) && !defined(VIXL_SIMULATOR)
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
__asm volatile("mrs %[ctr], ctr_el0" : [ctr] "=r"(CTR));
@@ -751,6 +766,7 @@ FEXCore::HostFeatures FetchHostFeatures() {
FetchHostFeatures(Features, HostFeatures, true, CTR, MIDR);
HostFeatures.SupportsCPUIndexInTPIDRRO = false;
HostFeatures.HostType = FEXCore::HostFeatures::HostTypeEnum::Linux;
return HostFeatures;
}
} // namespace FEX
Loaded 100 of 520 files, more files were not shown because too many files have changed in this diff. Show more