Compare commits

...
418 Commits
Author SHA1 Message Date
Ryan Houdek 1cc4b93e7a Docs: Update for release FEX-2607 2026-07-02 17:47:31 -07:00
LC b4e2f5118a Merge pull request #5645 from Sonicadvance1/180
Windows: Fixes SHM stats reallocation
2026-07-02 19:33:37 -04:00
Ryan Houdek 812b6398e5 Windows: Fixes SHM stats reallocation
This was accidentally setting `CurrentSize` instead of just returning
the newly allocated size to the frontend. This was causing the frontend
to then fail to detect the reallocation actually occured and no longer
get stats for new threads.

Also happened to not use `NewSize` but instead `CurrentSize * 2` which
didn't matter as it matched the growth pattern, but was technically
incorrect.

Fixes SHM stats since the introduction of the unixlib, ezpz.
2026-07-02 13:38:14 -07:00
LC 1db45e2a70 Merge pull request #5639 from Sonicadvance1/179
FEXServerClient: Workaround sun_path 108 byte limit
2026-07-02 04:03:48 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
Ryan Houdek 5f1c8efe0e Merge pull request #5643 from lioncash/rec
VectorOps: Avoid temporary if able in 256-bit VFRecp
2026-07-01 17:25:33 -07:00
Ryan Houdek b9aeccf13b Merge pull request #5642 from lioncash/feature
HostFeatures: Put SVE support querying into single function
2026-07-01 15:29:02 -07:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
Ryan Houdek 5d8d052a77 Merge pull request #5640 from lioncash/move
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka 7fa4d78269 Merge pull request #5635 from Sonicadvance1/178
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek 6bf0db7df6 FEXServerClient: Workaround sun_path 108 byte limit
We really don't want to do this, but in the case that the AF_UNIX path
is longer than the 108-byte limit that sun_path provides we don't really
have a choice. The alternative choice would be to switch /entirely/ away
from AF_UNIX and instead use pipes. We need a bandage fix for now, so
throw the socket in to a temp folder if the path is too long.
2026-06-30 17:31:23 -07:00
Ryan Houdek 110313e7de Merge pull request #5638 from lioncash/swap
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek 44e24c9e6b Merge pull request #5637 from mrpippy/unicodestring
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks 6c58fef220 Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism. 2026-06-30 15:37:55 -07:00
LC 3d593ce87d Merge pull request #5626 from Sonicadvance1/177
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek 36e1b5107e InstcountCI: Update 2026-06-30 15:01:46 -07:00
Ryan Houdek f89123f489 unittests/ASM: Allow up to 3-bits of precision loss 2026-06-30 15:00:16 -07:00
Paulo Matos b7280a765d asm_tests: Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek a4f89b79a3 Merge pull request #5636 from lioncash/move
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek 5bde4d875a FEXOfflineCompiler: Fixes HostFeature detection under Win32
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
  differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
  WINE

Ensure that when built for Win32 that it uses the correct feature
fetching.

One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek a79c471c31 Merge pull request #5634 from lioncash/shuffle
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
LC d3a85e14d9 VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
Ryan Houdek 417bd8604c Merge pull request #5632 from lioncash/vpblendd_test
unittests: Add test for stress-testing VPBLENDD selectors
2026-06-29 22:13:05 -07:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek d138c854f3 Merge pull request #5631 from lioncash/vpblendd
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek e26a792b70 Merge pull request #5630 from lioncash/comiss
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek 7e2d3b07c0 Merge pull request #5629 from lioncash/comment
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek 394a6f28db Merge pull request #5628 from lioncash/movmsk
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek 72e01274c9 Merge pull request #5627 from lioncash/vpblendw
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 59919a0b0c unittests: Expand VPBLENDW selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 07:03:16 -04:00
LC bf1857ecc7 unittests: Expand VPBLENDD selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 06:56:37 -04:00
LC b40920596a unittests: Add selector stress tests for VPSHUFD
Ensures any added optimization paths result in the same output.
2026-06-29 06:56:34 -04:00
LC 3b14c322e8 unittests: Add selector stress tests for VPSHUFLW/VPSHUFHW
Ensures any added optimization paths result in the same output.
2026-06-29 06:45:18 -04:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 6ca2d27c82 unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator 2026-06-29 08:00:58 +02:00
Simon Scherer d2f26d1969 InstcountCI: Update 2026-06-29 07:55:05 +02:00
LC fc25443827 unittests: Move full_vpblendw_imm test into the VEX folder
Keeps all of the selector tests in the same location.
2026-06-28 23:17:19 -04:00
LC 986800885a unittests: Add test for stress-testing VPBLENDD selectors
Drops a test in like the one for VPERMQ to ensure that, even if different
optimization paths are introduced, the behavior remains consistent.
2026-06-28 23:13:39 -04:00
LC 52c6ab1cec instcountci/VEX_map3: Add missing third param to VPBLENDD
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek 83a989c6dc Merge pull request #5625 from lioncash/perm
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC 32b96c259b Merge pull request #5623 from Sonicadvance1/176
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek 954581c750 Windows/UnixLib: Adds remaining helpers
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.

Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC 24da43f823 Merge pull request #5613 from Sonicadvance1/175
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek 9740f488cf Merge pull request #5624 from lioncash/psad
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC f81467fd30 instcountci/VEX_map1: Remove obsolete comments
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC a215bb9709 instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek 848c4b2d68 Merge pull request #5621 from wsxarcher/fixfexfolder
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek 4f995cbc1c Merge pull request #5620 from wsxarcher/fixuafbash
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
Ryan Houdek 3470dd1e7b Merge pull request #5619 from lioncash/dpps
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00
wsxarcher 2ba35286ef Do not exec FEX if it is a folder 2026-06-28 17:28:41 +02:00
wsxarcher bbd8212ced Fix UAF of PS1 2026-06-28 17:18:05 +02:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
Simon Scherer b9052ed7f0 unittests/ASM: Test zeroing of vcvtps2ph 2026-06-28 13:14:26 +02:00
Ryan Houdek 4160a92621 Windows: Fixes duplicated hardware TSO handling
This was handled in both the Module.cpp files and also the common
TSOHandlerConfig on accident. Wouldn't have caused an issue but it was
definitely a bit weird.
2026-06-27 20:59:15 -07:00
Ryan Houdek 201bb73980 Windows/UnixLib: Adds support for Hardware TSO support
Including fallback to non-unixlib path because we need to support both.

Showcases how these are going to be implemented without throwing the
entire world at it right away. Next PR will be implementing the
remaining four necessary unixlib handlers that we will require:

- Kernel unaligned atomic control
- shm_stats thing
- madvise operation
- prctl vma naming
2026-06-27 19:59:35 -07:00
LC 78832cc0d0 Merge pull request #5612 from Sonicadvance1/174
Windows: Load unixlib if possible
2026-06-27 22:25:43 -04:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek 70fe9a4405 Merge pull request #5614 from lioncash/whoops
Vector: Fix typo in VPERMQOp
2026-06-26 21:45:33 -07:00
Ryan Houdek c09225f868 Windows: Load unixlib if possible
Currently does nothing other than load it (as the library also doesn't
do anything yet). Ensured it was working by temporarily creating a test
entrypoint and doing `Call` on to it.

Next step after this is to reimplement some of the nasty hacks FEX is
doing inside the unixlib code itself.
2026-06-26 20:50:24 -07:00
Ryan Houdek 126bcd365d winternl: Update enums
Newer WINE has a better mechanism for asking to load unix libraries.
Older WINE like what is in Proton doesn't have this yet. Add definitions
for both so we can try either one.
2026-06-26 20:50:19 -07:00
LC fe4d2bc6c5 Merge pull request #5611 from Sonicadvance1/173
Windows: Adds empty Linux side unix library
2026-06-26 23:48:51 -04:00
Ryan Houdek dbaf22372c Windows: Adds empty Linux side unix library
We are going to need a unix library. Going to take this one step at a
time without AI/ML so I fully understand all the pieces of the puzzle,
and to ensure we don't lose any functionality before we're ready.

This only ensures that we are building the Linux facing .so files for
arm64ec and wow64, but they are empty today. Next PR will be
initializing it on the PE side.
2026-06-26 19:05:08 -07:00
LC 43bd243457 Merge pull request #5609 from Sonicadvance1/170
OpcodeDispatcher: Fixes CRC32 with high 8-bit register
2026-06-26 15:30:00 -04:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
Ryan Houdek 26266c6a94 InstcountCI: Adds CRC32 with high 8-bit register 2026-06-26 11:38:08 -07:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Simon Scherer 0c1a35f297 unittests/ASM: Add unit test for crc32 with 8bit register operand 2026-06-26 11:10:06 -07:00
Tony Wasserka 9ac608ca43 Merge pull request #5606 from Sonicadvance1/169
LibraryForwarding/cuda: Convert constexpr to const
2026-06-26 11:17:49 +02:00
Ryan Houdek 7ae55d73c1 Merge pull request #5607 from lioncash/broadcast
AVX: Handle easily broadcastable permutations in VPERMQ
2026-06-25 20:23:55 -07:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek f8491794d3 thunks/cuda: Convert constexpr to const
Apparently some compilers or libstdc++ or libc++ takes offence to
constexpr std::array that gets filled by GOT. My compiler this generates
the same code regardless but I guess this'll probably fix #5582.
2026-06-25 16:53:59 -07:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
Ryan Houdek 3d66be9e5a Merge pull request #5604 from lioncash/pcmpstr
instcountci: Add 16-bit pcmpxstrx variants
2026-06-25 10:23:59 -07:00
Ryan Houdek ec1b24d05b Merge pull request #5603 from lioncash/permq
AVX: Handle transpose cases in VPERMQ
2026-06-25 01:18:01 -07:00
Ryan Houdek d93997c1cb Merge pull request #5602 from lioncash/same
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek e19aa975c8 Merge pull request #5597 from simon902/BTOpTypo
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00
Simon Scherer 3f02dd0a36 OpcodeDispatcher: Update BTOp comment to clarify AMD vs Intel flag behavior 2026-06-25 07:36:32 +02:00
Ryan Houdek ab4e0f653a Merge pull request #5600 from lioncash/palign
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek fda023e7dc Merge pull request #5601 from lioncash/pshufb
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00
Ryan Houdek ad618be979 Merge pull request #5599 from lioncash/unused
Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
2026-06-24 09:47:38 -07:00
Ryan Houdek 5881266256 Merge pull request #5598 from lioncash/ilpd
AVX: Wire up helper to VPERMILPD
2026-06-24 09:03:22 -07:00
Simon Scherer e03187852b OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved. 2026-06-24 16:11:13 +02:00
Ryan Houdek b846cb6c2c Merge pull request #5596 from lioncash/ilps
AVX: Wire up lane helper for VPERMILPS imm variant
2026-06-24 00:56:36 -07:00
Ryan Houdek 62ecdd650f Merge pull request #5595 from lioncash/pspd
AVX: Wire up lane helper for VSHUFPD/VSHUFPS
2026-06-24 00:22:12 -07:00
Ryan Houdek 9f195ff377 Merge pull request #5594 from lioncash/pshuf
AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
2026-06-23 20:31:04 -07:00
Ryan Houdek 27a5f09185 Merge pull request #5592 from ShadowCurse/fixes
JIT: Arm64: fix the loop in CacheLineClear/Clean
2026-06-23 17:15:07 -07:00
Ryan Houdek 01b0b4e653 Merge pull request #5593 from lioncash/dup
VectorOps: Avoid dup if able in VInsElement 128-bit element path
2026-06-23 17:10:35 -07:00
Egor Lazarchuk ff5dfff5bb IR: fix typo in CacheLineClean description 2026-06-24 00:38:59 +01:00
Egor Lazarchuk e3e9777ee6 JIT: Arm64: fix the loop in CacheLineClear/Clean
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
2026-06-24 00:38:48 +01:00
Ryan Houdek 7dc2dc8749 Merge pull request #5591 from lioncash/typo
OpcodeDispatcher: Fix typo in comment
2026-06-23 16:16:29 -07:00
Ryan Houdek 4eb5694872 Merge pull request #5590 from lioncash/insert
AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
2026-06-23 15:21:45 -07:00
Ryan Houdek 681c5e8097 Merge pull request #5589 from lioncash/selector
OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
2026-06-23 13:01:41 -07:00
Ryan Houdek 5ad06b0255 Merge pull request #5588 from lioncash/pinsr
AVX: Remove unnecessary moves from PINSRX ops
2026-06-23 12:34:03 -07:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC a4cd5f7584 instcountci: Add 16-bit pcmpxstrx variants
Also adds expanded mask variants. Lets us get a better whole picture on
all the main paths of these instructions.
2026-06-23 09:25:37 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek 1619374252 Merge pull request #5587 from lioncash/pd
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
Ryan Houdek c13064e201 Merge pull request #5586 from lioncash/mov
[SVE256] Remove unnecessary move in VCVTPS2PD
2026-06-22 20:44:23 -07:00
LC 20647f2287 AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
Ryan Houdek 37e32fbcb9 Merge pull request #5585 from lioncash/cmp-move
[SVE256] Remove heavy handed moves from scalar compares
2026-06-22 17:26:37 -07:00
Ryan Houdek 27acbba52e Merge pull request #5584 from lioncash/cmp_scalar
[SVE256] More comprehensively test SSE insertions for scalar comparisons
2026-06-22 15:49:57 -07:00
Ryan Houdek 6f33d2b4c4 Merge pull request #5583 from lioncash/scalar
[SVE256] Add more SSE scalar variant unit tests
2026-06-22 12:41:52 -07:00
LC a2e4209f4f AVX: Reduce moves in 256-bit VPSHUFB
Just a minor reduction by avoiding insertion overhead.
2026-06-22 10:47:49 -04:00
LC 73d6716828 AVX: Shave some moves off 256-bit VPALIGNR
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
2026-06-22 10:09:13 -04:00
LC 1fb2419be2 Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
No behavior change, just a reduction in noise.
2026-06-22 09:47:46 -04:00
LC 3394808c06 AVX: Wire up helper to VPERMILPD
We can just leverage the shuffle handler for this, since VPSHUFD
essentially functions like VPERMILPD
2026-06-22 09:03:22 -04:00
LC c252b58a15 AVX: Wire up lane helper for VPERMILPS imm variant
Makes for some more trivial savings. Will need handling for VPERMILPD
added separately, since selector behavior is different.
2026-06-22 05:21:10 -04:00
LC 39ae8c3ea0 AVX: Wire up lane helper for VSHUFPD/VSHUFPS
Also allows collapsing quite a bit of emitted code, like with
the shuffles in #5594
2026-06-22 04:46:22 -04:00
LC eb8c2d964c Vector: Factor out 128-bit path in SHUFOpImpl
We can leverage this for the 256-bit path
2026-06-22 03:47:20 -04:00
LC bd9cf9ca11 AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.

Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.

For example:

vpshufd ymm0, ymm1, 0b00000011

drops from 50 instructions to 9
2026-06-22 00:16:39 -04:00
Ryan Houdek 280568df2f Merge pull request #5581 from lioncash/fwd
Passes: Trim unnecessary forward declarations
2026-06-21 20:05:23 -07:00
Ryan Houdek b87ff1e2dc Merge pull request #5580 from lioncash/list
IntrusiveIRList: Amend signature for PostRA()
2026-06-21 20:04:51 -07:00
Ryan Houdek 55c90cfc38 Merge pull request #5579 from lioncash/typo
JIT: Amend op typos in implementations
2026-06-21 20:04:19 -07:00
Ryan Houdek 500d2374a5 Merge pull request #5578 from lioncash/tidy
Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
2026-06-21 20:03:47 -07:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC 9ab0920f01 OpcodeDispatcher: Fix typo in comment
It's the bits in general, not just the even ones (whoops).
2026-06-21 19:16:38 -04:00
LC a6e7fba433 AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.

In some cases, this can be quite drastic, like with:

vpblendw ymm0, ymm1, ymm2, 0b00000001

being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC 8905e39439 OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC df41b85827 OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC 9fa3b9345e AVX: Remove unnecessary moves from PINSRX ops
These are old paths still around from when StoreResult used to
automatically perform truncating moves.

These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
LC 42af6c8508 [SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC b7df1bc259 [SVE256] Remove unnecessary move in VCVTPS2PD
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC b40f9db735 [SVE256] Remove heavy handed moves from scalar compares
(See #3799)

I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)

This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC 938841c658 [SVE256] More comprehensively test SSE insertions for scalar comparisons
See #3799

Drops in the facilities to ensure all of the available SSE scalar comparison
paths are tested for proper insertion behavior.
2026-06-21 00:35:47 -04:00
LC 34344e5769 Merge pull request #5576 from Sonicadvance1/167
Proton: Fixes Mafia 3
2026-06-20 23:18:33 -04:00
Ryan Houdek 538a9624ec ArchHelpers/Arm64: Fixes zero register usage
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
2026-06-20 19:56:18 -07:00
Ryan Houdek c4c69ca8de ArchHelpers/Arm64: Support CAS/CASP in non-JIT SIGBUS handler
Proton was using this
2026-06-20 19:55:39 -07:00
LC 90330ab3e4 [SVE256] Add more SSE scalar variant unit tests
See #3799

These were technically already covered when the work was done to make
SSE insertion behavior conform to hardware, so this just adds tests
that ensure that behavior holds over time.
2026-06-20 19:25:19 -04:00
Ryan Houdek c5880e7618 Merge pull request #5572 from lioncash/str
[SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
2026-06-20 15:13:21 -07:00
Ryan Houdek 9d0c05d9cc Merge pull request #5575 from lioncash/movq2dq
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek f6a68cb7fd Merge pull request #5574 from lioncash/sse4a
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
Ryan Houdek 462c785418 Merge pull request #5573 from lioncash/cvtpi
[SVE256] Handle SSE insertions for CVTPI2PD
2026-06-20 12:44:21 -07:00
LC a1f90dd8d3 Passes: Trim unnecessary forward declarations
Less visual noise and lingering types left in the header.
2026-06-20 13:59:18 -04:00
LC 2a67261eac IntrusiveIRList: Amend signature for PostRA()
PostRA is a bool, not an unsigned value. We can also adjust SpillSlots()
to use uint32_t like its returned data member.
2026-06-20 13:39:35 -04:00
LC 1d3403fdc2 JIT: Amend op typos in implementations
Mostly benign, but ensures that they're correct in the event any of
their IR definitions change.
2026-06-20 13:16:12 -04:00
LC 53301b0f56 Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
Same thing, just a little less verbose.
2026-06-20 12:32:23 -04:00
LC 8989ce1766 [SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
See #3799
2026-06-19 18:53:50 -04:00
LC faa121e9ef Vector: Move MOVQ2DQ over to Bind
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC f48759e83a [SVE256] Handle SSE insertions for MOVQ2DQ
See #3799
2026-06-19 17:42:21 -04:00
LC 2eca733603 [SVE256] Handle SSE insertions for EXTRQ/INSERTQ
See #3799
2026-06-19 17:05:59 -04:00
LC 14b65cec43 [SVE256] Handle SSE insertions for CVTPI2PD
See #3799

CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
Ryan Houdek f5477039fa Merge pull request #5571 from simon902/16bitleave
Fix incorrect RSP update for 16bit leave
2026-06-19 10:27:33 -07:00
Simon Scherer 9fa3221687 FEXCore: Fix incorrect RSP update for 16bit leave 2026-06-19 14:23:16 +02:00
Simon Scherer c8c63faf15 unittests/ASM: Adds unit test for 16bit leave 2026-06-19 14:16:15 +02:00
Ryan Houdek ee4794c99e Merge pull request #5570 from lioncash/pmadd
[SVE256] Handle SSE insertions for PMADDWD
2026-06-17 22:21:16 -07:00
Ryan Houdek 3a23bb4b73 Merge pull request #5569 from lioncash/cmp
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00
Ryan Houdek 46ffb25f84 Merge pull request #5568 from lioncash/mov3
[SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
2026-06-17 21:21:50 -07:00
Ryan Houdek 069a4025c5 Merge pull request #5567 from lioncash/mov2
[SVE256] Handle SSE insertion for aligned/unaligned loads and non-temporal loads
2026-06-17 20:18:20 -07:00
LC edd044752d [SVE256] Handle SSE insertions for PMADDWD
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek 3929d25dcc Merge pull request #5566 from lioncash/mov
[SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
2026-06-17 18:55:22 -07:00
LC 99baa4f3d9 [SVE256] Handle SSE insertions for CMPSD/CMPSS
See #3799
2026-06-17 21:05:31 -04:00
LC f64d4c571b [SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
See #3799
2026-06-17 20:38:58 -04:00
LC 9372fa169a [SVE256] Handle SSE insertions for MOVSD/MOVSS
See #3799
2026-06-17 20:09:17 -04:00
Ryan Houdek f98ac7f268 Merge pull request #5565 from lioncash/xor
[SVE256] Handle SSE insertions for XOR special case
2026-06-17 16:47:50 -07:00
LC daaa6ec129 [SVE256] Handle SSE insertions for aligned and unaligned moves
See #3799
2026-06-17 19:39:06 -04:00
LC 88afc22d5b [SVE256] Handle SSE insertions for MOVNTDQA
See #3799
2026-06-17 19:38:57 -04:00
Ryan Houdek 23099100b6 Merge pull request #5564 from lioncash/misc2
[SVE256] Handle SSE insertions for INSERTPS, PSIGN, PINSR, and shuffles
2026-06-17 16:24:26 -07:00
LC b7ea9e30df [SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
See #3799

Gets a few of the moves out of the way.
2026-06-17 17:39:13 -04:00
LC 844c3bb197 [SVE256] Handle SSE insertions for XOR special case
See #3799

Ensures that our special case maintains insertion behavior
2026-06-17 16:33:05 -04:00
LC 208c6d3eac [SVE256] Handle SSE insertions for vector unary ops
See #3799
2026-06-17 14:43:56 -04:00
LC 886a2e74ac [SVE256] Handle SSE insertions for pack ops 2026-06-17 13:36:18 -04:00
Ryan Houdek adad3c27dd Merge pull request #5563 from lioncash/shift
[SVE256] Handle SSE insertions for shifts
2026-06-17 08:02:41 -07:00
Ryan Houdek 470aeab215 Merge pull request #5562 from lioncash/misc 2026-06-17 05:54:59 -07:00
LC db1d90ec9d [SVE256] Handle SSE insertions for shuffles
See #3799
2026-06-17 08:44:45 -04:00
LC 124ce8420a [SVE256] Handle SSE insertions for PINSR(B,D,Q,W)
See #3799
2026-06-17 08:22:26 -04:00
LC 4433eaf242 [SVE256] Handle SSE insertions for INSERTPS
See #3799
2026-06-17 08:11:49 -04:00
LC 5ba070f600 [SVE256] Handle SSE insertions for PSIGN(B,D,W)
See #3799
2026-06-17 07:59:33 -04:00
LC 2d1a42aa00 [SVE256] Handle SSE insertions for shifts
See #3799
2026-06-17 06:06:08 -04:00
LC 74d9f5a3e2 [SVE256] Handle SSE insertions for MOVDDUP
See #3799
2026-06-17 05:13:08 -04:00
LC 634fbb5a73 [SVE256] Handle SSE insertions for Float->Int/Int->Float conversions
See #3799
2026-06-17 04:44:39 -04:00
LC e16948bf80 [SVE256] Handle SSE insertions for CVTPD2PS/CVTPS2PD
See #3799
2026-06-17 03:51:36 -04:00
LC 76c9833ee5 [SVE256] Handle SSE insertions for CMPPD/CMPPS
See #3799
2026-06-17 03:33:26 -04:00
Ryan Houdek 99662b70ff Merge pull request #5560 from lioncash/psad
[SVE256] Handle SSE insertions for more misc ops
2026-06-17 00:14:30 -07:00
LC fcde9eabbf [SVE256] Handle SSE insertions for PALIGNR
See #3799
2026-06-17 02:13:37 -04:00
LC 4ef15951a8 [SVE256] Handle SSE insertions for PACKSS/PACKUS ops
See #3799
2026-06-17 02:08:24 -04:00
LC 1bc51c2290 [SVE256] Handle SSE insertions for PMULUDQ
See #3799
2026-06-17 02:00:11 -04:00
LC be1025901a [SVE256] Handle SSE insertions for ADDSUBPD/ADDSUBPS
See #3799
2026-06-17 01:54:55 -04:00
Ryan Houdek 1a606de29f Merge pull request #5559 from lioncash/phmin
[SVE256] Handle SSE insertions for PHMINPOSUW, DPPD, and DPPS
2026-06-16 22:50:53 -07:00
LC 454c0b31cb [SVE256] Handle SSE insertions for MPSADBW
See #3799
2026-06-17 01:47:21 -04:00
Ryan Houdek 223e0f4e53 Merge pull request #5558 from lioncash/blend
[SVE256] Handle SSE insertions for blends
2026-06-16 22:34:32 -07:00
LC 3a84091945 [SVE256] Handle SSE insertions for DPPD/DPPS
See #3799
2026-06-17 01:32:36 -04:00
LC a0e8f1097f [SVE256] Handle SSE insertions for PHMINPOSUW
See #3799
2026-06-17 01:21:19 -04:00
Ryan Houdek 08ed4fb983 Merge pull request #5546 from neobrain/feature_woa_fexofflinecompiler
CodeCache: Support targeting WOW64/ARM64EC in FEXOfflineCompiler
2026-06-16 22:15:49 -07:00
Ryan Houdek 0b1f336e03 Merge pull request #5557 from lioncash/round 2026-06-16 22:09:20 -07:00
LC 0258fcb116 [SVE256] Handle SSE insertions for blends
See #3799
2026-06-17 01:05:46 -04:00
LC 3272aa3f08 [SVE256] Handle SSE insertions for ROUNDPD/ROUNDPS
See #3799
2026-06-17 00:42:20 -04:00
Ryan Houdek d9ea6651f8 Merge pull request #5556 from lioncash/madd
[SVE256] Handle more SSE insertions for some one-off instructions
2026-06-16 21:22:29 -07:00
LC 32b11603d8 [SVE256] Handle SSE insertions for PMOVSX/PMOVZX ops 2026-06-17 00:04:03 -04:00
LC 538fd2672d [SVE256] Handle SSE insertions for PSADBW
See #3799
2026-06-17 00:04:03 -04:00
LC 3e37724e3e [SVE256] Handle SSE insertions for PHADDSW
See #379
2026-06-17 00:04:03 -04:00
LC 1b249ba76b [SVE256] Handle SSE insertion for PHSUBD/PHSUBW/PHSUBSW 2026-06-17 00:04:03 -04:00
LC df1295fbd0 [SVE256] Handle SSE insertions for HSUBPD/HSUBPS
See #3799
2026-06-17 00:04:03 -04:00
LC dcd71fe126 [SVE256] Handle SSE insertions for PMULHW/PMULHRSW
See #3799
2026-06-17 00:04:00 -04:00
LC 610ee5db76 [SVE256] Handle SSE insertions for PMADDUBSW
See #3799
2026-06-16 23:02:25 -04:00
Ryan Houdek 9d5494d9f0 Merge pull request #5555 from lioncash/alu
[SVE256] Vector: Handle SSE insertion properly for various ALU operations
2026-06-16 19:56:38 -07:00
Ryan Houdek dd44bc8d00 Merge pull request #5554 from lioncash/vmov
OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
2026-06-16 19:47:01 -07:00
Ryan Houdek 110c7cb62b Merge pull request #5553 from lioncash/bind
OpcodeDispatcher: Make use of Bind consistently
2026-06-16 19:44:44 -07:00
Ryan Houdek 09aa5abbdf Merge pull request #5552 from lioncash/literal
OpcodeDispatcher: Move a few stray literal accesses to Literal()
2026-06-16 19:37:28 -07:00
LC 1d8b6df630 [SVE256] Vector: Handle SSE insertion properly for various ALU operations
See #3799 for the bulk of the issue explanation. Ensures that emulated
SSE operation on aarch64 don't end up obliterating the upper 128-bit
lane when SVE-256 is present (Adv. SIMD operations zero-extend)

Knocks out quite a few SSE instructions right off the jump.
2026-06-16 22:26:28 -04:00
LC 195058752e OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
Will make removing the TODO in LoadSource regarding partial loads a
little easier.
2026-06-16 17:47:01 -04:00
LC c62805e86d OpcodeDispatcher: Make use of Bind consistently
We had a few places that were using Bind, and a few other places
that were using specializations as a means to composing the instruction
tables. Instead, we can just use Bind consistently, which lets us tidy
up a bunch of the implementations (and gets rid of some unnecessary
codegen).
2026-06-16 15:07:59 -04:00
LC d15b175c33 OpcodeDispatcher: Move a few stray literal accesses to Literal()
Same core behavior, but ensures that the immediates are valid literals
when assertions are enabled.
2026-06-16 11:50:48 -04:00
Tony Wasserka 6cb73adfd5 CodeCache: Integrate FEXOfflineCompiler backend for WoA 2026-06-16 17:33:16 +02:00
Billy Laws 6d1cd67900 Windows: Implement WOW64/ARM64EC offline compiler backend
Implements an offline JIT compiler backend that compiles x86 code blocks
from a given code map into an ARM64 code cache on Windows. Separate
binaries are built for WOW64 (32-bit) and ARM64EC (64-bit) targets.

NOTE: This patch originally added a separate binary; instead the new
      functionality is added to the existing FEXOfflineCompiler and will
      be properly integrated in the next patches.
2026-06-16 17:30:32 +02:00
Ryan Houdek e12bd27106 Merge pull request #5536 from Sonicadvance1/105
LibraryForwarding: Implement support for CUDA
2026-06-12 15:14:38 -07:00
Ryan Houdek cb018257cf Merge pull request #5547 from neobrain/feature_fexofflinecompiler_process_all
FEXOfflineCompiler: Add "process-all" verb
2026-06-12 02:11:33 -07:00
LC e02953dc17 Merge pull request #5548 from Sonicadvance1/166
Frontend: Fix vsyscall page tracking.
2026-06-03 21:20:20 -04:00
Ryan Houdek ba56f8e0c5 unittests/FEXLinuxTests: Adds 64-bit vsyscall test
We never actually had a unittest to ensure these keep working, so add
one now.
2026-06-03 18:03:31 -07:00
Ryan Houdek ac12dd55c3 Frontend: Fix vsyscall page tracking.
Now that NX is tracked in the frontend, we need to ensure that adjusted
RIP pages are tracked correctly. Keep around both instruction stream
pointers, validate the the original RIP is executable, and read from the
adjusted RIP as appropriate.

Fixes #5544
2026-06-03 18:03:31 -07:00
Ryan Houdek 24647820d7 Linux: Ensure 64-bit vsyscall page is tracked
It's a purely virtual page even on x86, so we need to manually add it to
tracking.
2026-06-03 18:03:31 -07:00
Ryan Houdek 7e8aa711ef Linux: Pass gettimeofday through glibc
This ensures that it hits the VDSO path if possible.
2026-06-03 17:45:41 -07:00
Tony Wasserka 329e561eff FEXOfflineCompiler: Add "process-all" verb
This operation will take care of any pending code cache operations:
* import new code maps from `CACHE_DIR/codemap/new` and process them to `CACHE_DIR/codemap/ready`
* generate caches for updated code maps with new blocks
* ensure caches already exist for all other code maps (and generate them if needed)
2026-06-03 18:42:37 +02:00
LC d848cbbc0f Merge pull request #5543 from Sonicadvance1/165
FEXCore: Ensure LOCK prefix instructions are handled correctly
2026-06-02 23:39:02 -04:00
Ryan Houdek cd46e43c20 unittests/FEXLinuxTests: Add a LOCK prefix test
Ensures we handle lock prefixing correctly.
2026-06-02 19:41:04 -07:00
Ryan Houdek 98674c1cc8 unittests/FEXLinuxTests: Support redirecting RIP entirely 2026-06-02 19:41:04 -07:00
Ryan Houdek 8bfae631b1 Frontend: Support raising unimplemented instruction on LOCK failure
When the LOCK prefix is on an instruction that doesn't support LOCK then
it raises a SIGILL. Make sure to pass that up.

Additionally if the instruction does support lock prefix, has a lock
prefix, but the destination is not memory then that is also invalid.
2026-06-02 19:41:03 -07:00
Ryan Houdek d00c7cf3a3 Frontend: Support passing the decode failure type through decoding
Only used for invalid inst currently
2026-06-02 19:41:03 -07:00
Ryan Houdek b45665fee4 OpcodeDispatcher: Support instruction type of unimplement operation 2026-06-02 19:41:03 -07:00
Ryan Houdek 1b58664541 X86Tables: Describe instructions that support LOCK prefix 2026-06-02 19:41:02 -07:00
LC ed6a178ae3 Merge pull request #5542 from Sonicadvance1/164
OpcodeDispatcher: Fixes 64-bit LODs with address size override
2026-06-02 22:40:53 -04:00
Ryan Houdek e925ca509d unittests/ASM: Adds lods tests with address size override
If the override isn't handled then it'll fall back to either using the
wrong address and getting the wrong data or crashing depending on how it
is broken.
2026-06-02 18:21:00 -07:00
Ryan Houdek ca310cf815 OpcodeDispatcher: Fixes 64-bit LODs with address size override
Fairly trivial but just need to be careful with address size wraparound
as usual.
2026-06-02 18:21:00 -07:00
Ryan Houdek 6fa27aac42 Thunks: Implement support for CUDA
This is enough to get less complex cuda applications running, and is a
good starting spot to slowly finish off the remaining implementation.

Some information:
- 429 functions in total
- 153 only compiled for 64-bit (35.6%)
- 7 functions disabled entirely (1.6%)

The main thing /not/ working with this initial implementation is .cu
files compiled in to an ELF using the static cuda runtime. This is due
to the `cuGetExportTable` function being stubbed out and the static cuda
RT requires at least two interfaces from that function before it
continues.

This function isn't publicly documented by NVIDIA but has been publicly
reverse engineered to be fairly trivial. It's just a jump table with the
first element being the size of the table in bytes.

That will be the next step of the implementation.
2026-06-01 16:32:45 -07:00
Tony Wasserka a5c3fc4751 Merge pull request #5514 from peppergrayxyz/snd_htimestamp_t
LibraryForwarding: Add annotation for snd_htimestamp_t
2026-06-01 12:53:42 +02:00
LC 65b05fa8c1 Merge pull request #5540 from Sonicadvance1/163
arm64ec: Single instruction optimization in EC map lookup
2026-06-01 05:41:34 -04:00
Pepper Gray 3ee556d858 add template for snd_htimestamp_t
building on musl fails with:
`error: Unsupported parameter type 'snd_htimestamp_t *' (aka 'timespec *')`

due to empty padding members in `alltypes.h`:
```
STRUCT timespec {
  time_t tv_sec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER==4321);
  long tv_nsec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER!=4321);
};
```
add (missing) annotation to libasound interface

fix: #5513
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-31 17:38:25 +02:00
Ryan Houdek fd3546a999 arm64ec: Single instruction optimization in EC map lookup
We can merge the lsr+and in to a single lsr by 18 and then use ldr with
LSL of 3 to accomplish the same result. Modern Cortex doesn't even
generate an additional integer pipeline uop for this ldr+lsl instruction
anymore.
2026-05-29 19:17:13 -07:00
Ryan Houdek f5fafa5b96 Merge pull request #5492 from FrontMage/codex/int29-failfast-probe
Windows: Trace interrupt translation and prototype INT 0x29 fail-fast mapping
2026-05-29 16:59:16 -07:00
Ryan Houdek 7ff2069e60 Merge pull request #5539 from Sonicadvance1/162
arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
2026-05-29 16:49:14 -07:00
Ryan Houdek 154ff43d7f arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
We were accidentally allocating some things without this flag and it was
causing us to dump memory in to the lower 32-bits on arm64ec.

This was causing the game
[Below](https://store.steampowered.com/app/250680/BELOW/) to run out of
memory to allocate for its LUA JIT and causes it to crash.
2026-05-29 16:17:04 -07:00
Ryan Houdek b754fe4810 External/rpmalloc: update 2026-05-29 16:16:46 -07:00
Tony Wasserka a1071ec01a Merge pull request #5538 from peppergrayxyz/syscall_headers
LinuxSyscalls: add missing thread header
2026-05-29 14:35:43 +02:00
Pepper Gray 92dce9a2ea add missing header <thread>
building using musl fails due to missing defintions:

```
Source/Tools/LinuxEmulation/LinuxSyscalls/Syscalls.cpp:913:23: error: no member named 'sleep_for' in namespace 'std::this_thread'
  913 |     std::this_thread::sleep_for(std::chrono::milliseconds {10});
      |                       ^~~~~~~~~
1 error generated.
```

Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-29 13:19:24 +02:00
LC 5fd917ec2f Merge pull request #5537 from Sonicadvance1/161
Format: Fix missed clang-format
2026-05-29 00:52:34 -04:00
Ryan Houdek d7cda23b25 Format: Fix missed clang-format
Minor clang-format version differences causing differing behaviour
again.
2026-05-28 12:12:24 -07:00
Ryan Houdek cae5da5777 Merge pull request #5530 from neobrain/fix_ccache_time_macros
Build: Enable ccache sloppiness for time macros
2026-05-25 10:55:27 -07:00
Ryan Houdek 97f1f47fa5 Merge pull request #5531 from neobrain/refactor_drop_cmake_settings
Drop unused CMakeSettings.json
2026-05-25 10:53:44 -07:00
Tony Wasserka 53c269ee25 Merge pull request #5516 from peppergrayxyz/thunk_rootfs
set sysroot to X86_DEV_ROOTFS for guest toolchain
2026-05-25 18:08:45 +02:00
Pepper Gray 1420d3cc10 set sysroot to X86_DEV_ROOTFS for guest toolchain
**Faulty Behaviour:**
When not building on Ubuntu `unittests/ThunkLibs` fails to find
c++ header and fails:

```
gen_input.cpp:2:10: fatal error: 'cstddef' file not found
    2 | #include <cstddef>
      |          ^~~~~~~~~
1 error generated.
```

This also includes building with nix-shell:
```
nix-shell ../Data/nix/LibraryForwarding/shell.nix --run "cmake --build . --target thunkgen_tests"
```

**Root Cause**:
`X86_DEV_ROOTFS` is not used, but  header paths are hard
coded to Ubuntu's multilib layout (`/usr/i686-linux-gnu/include/`,
`/usr/x86_64-linux-gnu/include/`). It works on Ubuntu but the
mechanism to use to another directory is broken and the include
paths always point to the host.

**Solution**:
set `--sysroot ${X86_DEV_ROOTFS}` for guest builds and remove
hard coded include paths.

Fix: #5515
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-25 17:18:47 +02:00
Tony Wasserka ce401b5ca1 Build: Enable ccache sloppiness for time macros
Ccache won't attempt to cache files that use __DATE__/__TIME__ since their
contents would always be out of date. Setting sloppiness disables this
behavior, which works for us since we don't use __DATE__/__TIME__ for anything
that needs accurate values.
2026-05-25 12:03:33 +02:00
Tony Wasserka cd412bd0f5 Drop unused CMakeSettings.json
Visual Studio (but not VS Code) used to read this file, but this convention is
discouraged nowadays.
2026-05-25 11:27:29 +02:00
FrontMage f33f88072e Windows: Fix INT exception formatting 2026-05-24 11:42:51 +08:00
LC 1240a00fa5 Merge pull request #5522 from Sonicadvance1/160
New CPL0 instructions from #5510 but with unittests
2026-05-23 22:57:53 -04:00
LC 83f325de0d Merge pull request #5519 from Sonicadvance1/157
meta: Add CONTRIBUTING.md
2026-05-23 22:55:26 -04:00
LC 0b871bf54e Merge pull request #5518 from Sonicadvance1/156
code-format-helper: More dependabot changes
2026-05-23 22:54:36 -04:00
LC 07f7aa3c8f Merge pull request #5521 from Sonicadvance1/159
HostFeatures: Don't capture CTR/MIDR under simulator
2026-05-22 22:10:38 -04:00
Ryan Houdek 7208bc6cdd unittests: Extend unittests for CPL0 instructions 2026-05-22 15:33:42 -07:00
Ryan Houdek fef5a98602 Fix build failure. 2026-05-22 15:33:41 -07:00
Daniel Lu 5cce65cdfa OpcodeDispatcher: Decode INVD and WBINVD through privileged op handling 2026-05-22 15:25:22 -07:00
Ryan Houdek 268081e5d0 Merge pull request #5520 from Sonicadvance1/158
Cherry-pick #5508 with instcountci changes
2026-05-22 14:46:20 -07:00
Ryan Houdek f5f179117e HostFeatures: Don't capture CTR/MIDR under simulator
If the simulator was selected, we would still capture CTR and MIDR on
the host AArch64 system. Potentially modifying codegen in unexpected
ways.

Ensure we return 0/0 like under x86 with simulator to simulate
"unknown".
2026-05-22 14:04:52 -07:00
Ryan Houdek 8c85096f98 meta: Add CONTRIBUTING.md 2026-05-22 14:03:14 -07:00
Ryan Houdek c5e7675c4b InstcountCI: Update 2026-05-22 13:58:10 -07:00
Daniel Lu 03009912ac JIT: Avoid clobbering guest rdx while raising generated faults 2026-05-22 13:56:40 -07:00
Ryan Houdek e4a1138291 code-format-helper: More dependabot changes 2026-05-22 13:44:22 -07:00
FrontMage a5bc54d2d5 Windows: Preserve INT 0x2D exception parameter source 2026-05-22 12:56:44 +08:00
LC df73e84725 Merge pull request #5507 from Sonicadvance1/155
CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588
2026-05-21 22:39:14 -04:00
Ryan Houdek 5bf07c2e77 CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588 2026-05-21 18:04:29 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
LC c98cef0da1 Merge pull request #5506 from Sonicadvance1/154
HostFeatures: Only enable `dc zva` optimization on Ampere CPUs
2026-05-19 22:29:11 -04:00
Ryan Houdek 1d9c52be02 InstcountCI: Update 2026-05-19 17:38:10 -07:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00
Ryan Houdek f66368b191 Merge pull request #5505 from fixedcat/main
Arm64: Fix byte-size handling in unaligned STLXR emulation
2026-05-19 12:28:23 -07:00
LC e4daea406e Merge pull request #5503 from Sonicadvance1/154
FEXCore: Allow InterruptFaultPage to be significantly further away
2026-05-19 12:22:18 -04:00
fixedcat 6216f22cb9 Arm64: Fix byte-size handling in unaligned STLXR emulation 2026-05-19 18:53:44 +08:00
LC d4c80d9094 Merge pull request #5422 from Sonicadvance1/138
FEXCore: Add support for developer single stepping, read/write watching.
2026-05-19 01:11:04 -04:00
Ryan Houdek 27324ded87 Merge pull request #5499 from bylaws/winstuff
Windows additions for code caching
2026-05-18 16:05:59 -07:00
Ryan Houdek 7d1c625e32 FEXCore: Allow InterruptFaultPage to be significantly further away
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.

Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
2026-05-18 15:57:25 -07:00
Ryan Houdek b4fe65f2c0 Merge pull request #5496 from neobrain/fix_libfwd_findpkg
Library Forwarding: Various build system improvements
2026-05-18 14:55:53 -07:00
Ryan Houdek f5efdac2e5 Merge pull request #5423 from bylaws/depenencey
WOW64: Support disabling DEP
2026-05-18 14:39:44 -07:00
Billy Laws 23de875516 Windows: Add NtUnmapViewOfSection prototype 2026-05-17 23:06:02 +01:00
Billy Laws 82030b8286 Windows: Declare winternl relocation APIs 2026-05-17 23:03:11 +01:00
Ryan Houdek af4da43bb8 Merge pull request #5494 from neobrain/fix_determine_va
Allocator: Fix and optimize VA range detection
2026-05-17 14:51:59 -07:00
Ryan Houdek 33f3b8659c Merge pull request #5495 from neobrain/fix_portable_config
Config: Use more sensible default for portable config location
2026-05-17 14:51:04 -07:00
Ryan Houdek ed724a61a7 Merge pull request #5498 from bylaws/evmd
ImageTracker: Support using image IDs as an extended volatile metadata key
2026-05-17 14:49:16 -07:00
Billy Laws 053bd74aa3 Windows/Common: Add ScopedHandle::reset() 2026-05-17 22:35:06 +01:00
Billy Laws 3c0410e59f ImageTracker: Support using image IDs as an extended volatile metadata key 2026-05-17 19:34:41 +01:00
Tony Wasserka e621f6c753 LibraryForwarding/Build: Explicitly look up LLVM headers
This could previously set up incorrect header paths when clang and LLVM were
installed in different directories (such as when using nix).
2026-05-15 16:07:13 +02:00
Tony Wasserka b05f000f42 LibraryForwarding/Build: Try harder to properly discover header locations 2026-05-15 16:07:13 +02:00
Tony Wasserka e60bfc6d23 LibraryForwarding/Build: Allow specifying system header location externally 2026-05-15 16:07:09 +02:00
Tony Wasserka 9a3d3201f9 Config: Use more sensible default for portable config location 2026-05-15 15:59:22 +02:00
Tony Wasserka abf9724424 Allocator: Fix and optimize VA range detection 2026-05-15 15:47:38 +02:00
LC ab9a8c62ab Merge pull request #5493 from Sonicadvance1/153
FEXGetConfig: Even more correctness changes for X2E
2026-05-14 09:39:31 -04:00
FrontMage e2fe936152 Windows: Handle INT 0x29 as fast-fail 2026-05-14 12:05:08 +08:00
Ryan Houdek 1d71650379 FEXGetConfig: Even more correctness changes for X2E
Some of the information was incorrect, so make sure it shows the
hardware correctly.
2026-05-13 20:00:57 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Tony Wasserka 52ad434d24 LinuxSyscalls: Defer MappedResource deletion until after code invalidation
This ensures that any code buffer memory owned by the MappedResource is
invalidated before being deallocated.
2026-05-13 21:25:47 +02:00
Tony Wasserka d69d111bb6 CodeCache: Align code section within cache files
This allows mapping the code directly into memory for execution.
2026-05-13 21:25:47 +02:00
LC 50f4494875 Merge pull request #5490 from Sonicadvance1/152
FEXGetConfig: Showcase RMW versus loadstore atomic differences
2026-05-12 16:32:04 -04:00
Ryan Houdek 8a4982383a FEXGetConfig: Showcase RMW versus loadstore atomic differences
This differs on X2E, so it's good to showcase it.
2026-05-12 12:38:44 -07:00
Ryan Houdek 0d72890482 Merge pull request #5432 from pmatos/f64-fprem
JIT-inline FPREM/FPREM1 for reduced precision x87 path
2026-05-11 14:31:31 -07:00
Ryan Houdek 9be7d6d112 Merge pull request #5489 from neobrain/feature_better_fexbash
FEXBash: Drop implicit -c and add colored PS1
2026-05-11 12:22:00 -07:00
Tony Wasserka 3c4121ba07 FEXBash: Use a shiny rainbow for PS1 2026-05-11 20:30:21 +02:00
Tony Wasserka b6e44b04d6 FEXBash: Clean up path handling 2026-05-11 20:30:21 +02:00
Tony Wasserka 02c11afbc6 FEXBash: Don't imply "-c" to behave more closely like bash
Implicitly adding "-c" breaks argument passing for scripts. For example, the
command "FEXBash ./steam.sh -silent" will process steam.sh but the script
wouldn't see the "-silent" argument previously.
2026-05-11 19:49:15 +02:00
Paulo Matos 2b8f5b57eb instcountci: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 0519c9467c asm_tests: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00
Ryan Houdek 7506cb4ea1 Merge pull request #5483 from neobrain/fix_guest_wrapper_code_cache
CodeCache: Delay cache loading for guest library wrappers until after LoadLib
2026-05-06 18:04:55 -07:00
Tony Wasserka 60bc5944db CodeCache: Slightly optimize cache file writing
ftruncate only requires one call (and one extra seek) instead up to 64 manual
zero writes.
2026-05-06 17:08:15 +02:00
Tony Wasserka 8f0572283a Windows/CRT: Implement ftruncate and _chsize 2026-05-06 17:08:04 +02:00
Tony Wasserka b13b46eefe CodeCache: Delay cache loading for guest library wrappers until after LoadLib
These libraries need to be initialized before relocating their caches,
since the guest function hashes won't be registered before.
2026-05-05 16:29:39 +02:00
LC 4db2a98d7f Merge pull request #5481 from Sonicadvance1/148
win32: Query DCZID_EL0 so clzero works
2026-05-05 08:42:31 -04:00
LC 05ebb07753 Merge pull request #5482 from Sonicadvance1/149
OpcodeDispatcher: Optimize MMX pshufw
2026-05-05 08:41:22 -04:00
Ryan Houdek 694e68b838 Merge pull request #5468 from peppergrayxyz/proc_self_stat
read /proc/self/stat using %lu
2026-05-04 20:21:16 -07:00
Pepper Gray 78320e1433 read /proc/self/stat using %lu
building on clang/musl causes warnings:

```
FEX/Source/Tools/FEXInterpreter/ELFCodeLoader.h:782:29: warning: format specifies type 'unsigned long long *' but the argument has type 'uint64_t *' (aka 'unsigned long *') [-Wformat]
  776 |                             "%llu %llu %llu %*u %*u "   // 26 to 30
      |                              ~~~~
      |                              %lu
  777 |                             "%*u %*u %*u %*u %*u "      // 31 to 35
  778 |                             "%*u %*u %*d %*d %*u "      // 36 to 40
  779 |                             "%*u %*u %*u %*d %llu "     // 40 to 45
  780 |                             "%llu %llu %llu %llu %llu " // 46 to 50
  781 |                             "%llu",                     // 51
  782 |                             &map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
      |                             ^~~~~~~~~~~~~~~
```

according to the [man page](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html)
`/proc/self/stat` uses `%lu`:

read the values as unsigned long (%lu) and then write them to
prctl_mm_map (platform specific format).

fixes: #5467
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-05 04:24:31 +02:00
Ryan Houdek 082e7b2695 InstcountCI: Update 2026-05-04 17:52:54 -07:00
Ryan Houdek cdbffb80c7 unittests: Adds a full coverage pshufw test 2026-05-04 17:52:53 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 47e173e549 Merge pull request #5472 from peppergrayxyz/format
use portable format specifiers
2026-05-04 14:18:28 -07:00
Ryan Houdek 162bd4be97 win32: Query DCZID_EL0 so clzero works
EL0 registers are readable without going through the registry, but we
were failing to populate this register, which was causing clzero to not
be supported.
2026-05-04 12:42:18 -07:00
Ryan Houdek f0764aeafe Merge pull request #5480 from neobrain/refactor_musl_sigmask
SignalDelegator: Simplify support for musl's sigset_t
2026-05-04 10:39:42 -07:00
Ryan Houdek 1efed71696 Merge pull request #5479 from neobrain/refactor_drop_compile_service
FEXCore: Drop unused CompileService
2026-05-04 10:36:44 -07:00
Tony Wasserka 06d77c1c19 SignalDelegator: Simplify support for musl's sigset_t 2026-05-04 16:28:47 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray abae5dd93b use portable format specifiers
building with clang/musl causes these warnings:

```
FEX/Source/Tools/FEXServer/ProcessPipe.cpp:100:96: warning: format specifies type 'ssize_t' (aka 'long') but the argument has type 'rlim_t' (aka 'unsigned long long') [-Wformat]
```

- cast platform specific MaxFDs members to uintmax_t and print as PRIuMAX
- use %zu for GetNumFilesOpen (size_t)

fix: #5471
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-04 11:20:55 +02:00
Ryan Houdek 93015a0266 Merge pull request #5458 from peppergrayxyz/largefile64
make 64bit symbols visible to enhance portability (musl)
2026-05-03 23:36:59 -07:00
Ryan Houdek d238db69d3 Merge pull request #5477 from peppergrayxyz/unistd
include missing header unistd.h
2026-05-03 15:18:41 -07:00
Ryan Houdek e0ead236b6 Merge pull request #5473 from peppergrayxyz/ObjectCacheRefCounter
remove dead code (ObjectCacheRefCounter)
2026-05-03 15:16:57 -07:00
Pepper Gray c548262664 include missing header unistd.h
build on clang/musl fails with:

```
FEX/unittests/APITests/Allocator.cpp:17:5: error: use of undeclared identifier 'close'
FEX/unittests/APITests/Allocator.cpp:23:5: error: use of undeclared identifier 'lseek'; did you mean 'fseek'?
FEX/unittests/APITests/Allocator.cpp:23:11: error: cannot initialize a parameter of type 'FILE *' (aka 'struct _IO_FILE *') with an lvalue of type 'int'
FEX/unittests/APITests/Allocator.cpp:24:5: error: use of undeclared identifier 'write'; did you mean '_IO_cookie_io_functions_t::write'?
FEX/unittests/APITests/Allocator.cpp:24:5: error: invalid use of non-static data member 'write'
```

include header to provide defintions

fix: #5476
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:49:40 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Ryan Houdek 85773995e1 Merge pull request #5470 from peppergrayxyz/header_redirect
fix include redirect for <poll.h> and <signal.h>
2026-05-03 04:41:18 -07:00
Pepper Gray 00b4777290 fix include redirect for <poll.h> and <signal.h>
building on musl/clang causes redirecting incorrect #includes warnings:

```
warning: redirecting incorrect #include <sys/poll.h> to <poll.h> [-W#warnings]
warning: redirecting incorrect #include <sys/signal.h> to <signal.h> [-W#warnings]
```

include headers instead of sys/headers.

fix: #5469
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 13:22:32 +02:00
Ryan Houdek 197e6de194 Merge pull request #5462 from peppergrayxyz/uc_sigmask
determine sigset_t fieldname to enhance portability (musl)
2026-05-03 03:48:53 -07:00
Ryan Houdek 9908ea4c2f Merge pull request #5466 from peppergrayxyz/libgen_h
include <libgen.h> for basename
2026-05-03 03:46:43 -07:00
Ryan Houdek c402b15bd3 Merge pull request #5464 from peppergrayxyz/tgkill
add header and classpath for tgkill
2026-05-03 03:46:07 -07:00
Ryan Houdek a3f3118ac5 Merge pull request #5460 from peppergrayxyz/sigset_t
use <signal.h> instead of glibc header to enhance portability (musl)
2026-05-03 03:18:20 -07:00
Pepper Gray b1aab0e498 determine sigset_t fieldname to enhance portability (musl)
musl build fails due to access to internal glibc member:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
  655 |       .SigMask = _context->uc_sigmask.__val[0],
      |                  ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```

add check to determine private glibc or musl member name or throw an error.

fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:17:09 +02:00
Pepper Gray 927c0ce54a include <libgen.h> for basename
building on clang/musl build fails due to missing symbol:

```
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:321:35: error: use of undeclared identifier 'basename'
  321 |   auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
      |                                   ^~~~~~~~
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:327:43: error: use of undeclared identifier 'basename'
  327 |     fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
      |                                           ^~~~~~~~
```

include missing header.

fixes: #5465
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:10:07 +02:00
Pepper Gray 57d9dc037d add header and classpath for tgkill
building on clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
 1174 |     tgkill(::getpid(), ::getpid(), SIGKILL);
      |     ^~~~~~
1 error generated.
```

include and use `FHU::Syscalls::tgkill`.

fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:50:23 +02:00
Pepper Gray e3cfe28848 use <signal.h> instead of glibc header to enhance portability (musl)
building using musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/ThreadManager.h:33:10: fatal error: 'bits/types/sigset_t.h' file not found
   33 | #include <bits/types/sigset_t.h>
      |          ^~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
```

`<bits/types/sigset_t.h>` is a glibc internal header, use <signal.h> instead.

fixes: #5459
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:03:38 +02:00
Pepper Gray ac91f583b8 make 64bit symbols visible to enhance portability (musl)
building using clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/x32/Types.h:548:5: error: member access into incomplete type 'const struct statfs64'
  548 |     COPY(f_bsize);
      |     ^
```

add `_LARGEFILE64_SOURCE` to define large-file feature macros

fix: #5457
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 10:52:09 +02:00
Ryan Houdek 215658bf29 Merge pull request #5455 from peppergrayxyz/sys_prctl
use <sys/prctl.h> to enhance portability (clang)
2026-05-02 15:11:33 -07:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek e91bda7765 Merge pull request #5447 from bylaws/claudefix5
SoftFloat: Fix FSCALE(0, +Inf) to raise IE and return a quiet NaN
2026-04-30 14:41:24 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Ryan Houdek b1381fd3b7 Merge pull request #5444 from bylaws/claudefix2
VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
2026-04-30 14:39:56 -07:00
Ryan Houdek e5f6a7d85e Merge pull request #5450 from neobrain/fix_code_cache_portable
CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
2026-04-30 14:17:16 -07:00
Ryan Houdek feae76fc4f Merge pull request #5451 from Sonicadvance1/146
arm64ec: Fixes crash in many games with SDL+Dualsense
2026-04-30 14:16:06 -07:00
Ryan Houdek 015f3cffb9 arm64ec: Fixes crash in many games with SDL+Dualsense
We were pointing to an incorrect function pointer and exploding when a
pending suspend doorbell had occured.
2026-04-30 12:59:53 -07:00
Tony Wasserka 3e5c17ae80 CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
In portable mode, FEXOfflineCompiler may not be in PATH (and if it is, it's
most likely not a compatible version). Instead, use the executable next to
the FEXServer binary.
2026-04-30 17:14:22 +02:00
Billy Laws 9a1d06c6ab WOW64: Support disabling DEP
Required for older 32-bit games that assumes the execute bit is implicit
from read.
2026-04-29 03:17:03 +00:00
Billy Laws 3e278b42f8 InstcountCI: Update 2026-04-29 02:54:34 +00:00
Billy Laws fa953445e9 InstcountCI: Update 2026-04-29 02:53:07 +00:00
Billy Laws a412b1d3b7 InstcountCI: Update 2026-04-29 02:43:19 +00:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c260b45e1 unittests/ASM: Adds tests for FXTRACT Inf/NaN 2026-04-29 02:29:23 +00:00
Ryan Houdek 098c4c57b4 Merge pull request #5443 from bylaws/claudefix1
X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel
2026-04-28 19:25:15 -07:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Billy Laws 1bd2ff3fc3 unittests/ASM: Adds test for FSCALE(0, +Inf) raising IE 2026-04-29 02:16:57 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Billy Laws cf20647b25 unittests/ASM: Adds test for 16-bit FIST with denormal input not setting IE 2026-04-29 02:09:54 +00:00
Billy Laws 8d7071e549 VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
2026-04-29 02:00:58 +00:00
Billy Laws fb006b2c6d unittests/ASM: Adds test for MAXPS/MAXPD NaN and signed-zero tie 2026-04-29 01:57:41 +00:00
Billy Laws 9039eeb3cd X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel 2026-04-29 01:57:18 +00:00
Billy Laws dd0702d30f unittests/ASM: Mark SSE4a/CLZERO as required for tests using them 2026-04-29 01:56:35 +00:00
LC 886faf0bd4 Merge pull request #5442 from neobrain/refactor_code_cache_check
CodeCache: Move bounds check to FEXOfflineCompiler
2026-04-28 19:47:52 -04:00
Tony Wasserka 86e28c6d34 CodeCache: Move bounds check to FEXOfflineCompiler
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.

It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
2026-04-28 17:31:20 +02:00
LC 821efab8aa Merge pull request #5440 from Sonicadvance1/145
OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
2026-04-28 07:55:10 -04:00
Ryan Houdek 5295365dd0 InstcountCI: Update 2026-04-27 17:55:49 -07:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
LC dd145aaa88 Merge pull request #5439 from Sonicadvance1/144
Fix push/pop fs/gs segments and unittests
2026-04-27 19:43:44 -04:00
Ryan Houdek 34b3adc23d unittests/ASM: Adds unit test to ensure push/pop segment of o16 works
Only ensures we are pushing and popping the correct size, not any of the
selector data within it, as 64-bit systems with the FSGSBase extension
don't use them selectors anyway.

Can't test the 32-bit side currently because we would corrupt FS/GS in
CI and the host testharnessrunner can't fix that right now.
2026-04-27 15:21:47 -07:00
Simon Scherer ab14882761 FEXCore: Fix 2byte stack access for 0x66 PUSH/POP FS/GS 2026-04-27 15:21:11 -07:00
LC 7dc1f54fb6 Merge pull request #5435 from Sonicadvance1/143
unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting
2026-04-25 10:34:07 -04:00
Ryan Houdek c09fb03eda unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting 2026-04-25 00:46:40 -07:00
Ryan Houdek dbf2761fb7 InstcountCI: Update 2026-04-25 00:45:05 -07:00
Ryan Houdek fd1378f778 InstcountCI: Fix incorrect instruction 2026-04-25 00:43:50 -07:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
LC 4b02c04afc Merge pull request #5429 from Sonicadvance1/142
Steam/CompatTool: Fixes Graphics Provider path handling
2026-04-23 20:13:13 -04:00
Ryan Houdek deed99e7a3 Merge pull request #5425 from pmatos/f64-atan-fyl2x
JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path
2026-04-23 15:16:04 -07:00
Ryan Houdek 7bffc4a177 Steam/CompatTool: Fixes Graphics Provider path handling
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.

Misunderstood the assignment before.
2026-04-23 15:13:42 -07:00
Ryan Houdek 701555e400 Merge pull request #5428 from Sonicadvance1/141
Steam/CompatTool: Support `STEAM_COMPAT_GRAPHICS_PROVIDER` for rootfs path
2026-04-21 12:50:22 -07:00
Ryan Houdek 49fa86d0b5 Steam/CompatTool: Support STEAM_COMPAT_GRAPHICS_PROVIDER for rootfs path
If we have been provided a graphics provider path through an environment
variable, then use that path directly rather than searching.
2026-04-21 12:37:40 -07:00
Paulo Matos adbace8810 instcountci: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos 050138bcea asm_tests: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
LC 59755ec115 Merge pull request #5426 from Sonicadvance1/139
Snapdragon X2 Elite fixes
2026-04-20 11:53:51 -04:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Ryan Houdek d41d52b889 Merge pull request #5419 from pmatos/f64-scale-f2xm1
JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path
2026-04-17 14:38:22 -07:00
Paulo Matos 18f69fb16d instcountci: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-17 17:28:13 +02:00
Ryan Houdek 739e85032b FEXCore: Add support for developer single stepping, read/write watching. 2026-04-16 14:13:19 -07:00
Paulo Matos d165711f2e asm_tests: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:24 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 441116e1e6 Merge pull request #5421 from Sonicadvance1/136
Scripts: Move arch check first in InstallFEX
2026-04-15 16:24:04 -04:00
Ryan Houdek ce97ef0ab1 Scripts: Move arch check first in InstallFEX
Don't give people false hope that the script might work on distros that
aren't Ubuntu.

Fixes #5420
2026-04-15 13:01:37 -07:00
LC 2ea0de92f4 Merge pull request #5418 from Sonicadvance1/135
ArchHelpers: Allow atomic memory operations in non-JIT handler
2026-04-15 07:22:25 -04:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
449 changed files with 22781 additions and 6286 deletions

No files matched your search

+14
View File
@@ -33,3 +33,17 @@ runs:
- name: Install
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_${{ inputs.target }} -t install
- name: Configure UnixLib
shell: bash
run: |
cmake -S Source/Windows/UnixLib -B build_unixlib_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE \
-G Ninja -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-unix -DCMAKE_INSTALL_PREFIX=/usr
- name: Build UnixLib
shell: bash
run: cmake --build build_unixlib_${{ inputs.target }}
- name: Install UnixLib
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_unixlib_${{ inputs.target }} -t install
+3 -1
View File
@@ -50,6 +50,8 @@ jobs:
with:
overwrite: true
name: wine_dll_artifacts
path: ${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
path: |
${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
${{ github.workspace }}/install/usr/lib/wine/aarch64-unix/lib*.so
retention-days: 60
compression-level: 9
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+11 -2
View File
@@ -194,6 +194,8 @@ if (ENABLE_GDB_SYMBOLS)
add_compile_definitions(GDB_SYMBOLS_ENABLED=1)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -251,8 +253,15 @@ endif()
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
execute_process(COMMAND "${CCACHE_PROGRAM}" --print-version
OUTPUT_VARIABLE CCACHE_VERSION OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "Enabling ccache ${CCACHE_VERSION}")
if (CCACHE_VERSION VERSION_GREATER_EQUAL "4.8")
# Set sloppiness to enable caching even for files that use __DATE__/__TIME__ macros
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM} sloppiness=time_macros")
else()
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
endif()
endif()
endif()
-132
View File
@@ -1,132 +0,0 @@
{
"environments": [
{
"BuildPath": "${projectDir}\\out\\build\\${name}",
"InstallPath": "${projectDir}\\out\\install\\${name}",
"clangcl": "clang-cl.exe",
"cc": "clang",
"cxx": "clang++"
}
],
"configurations": [
{
"name": "WSL-Clang-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "WSL-Clang-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "x86-Clang-Cross-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "x64-Clang-Cross-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "Linux-Clang-Remote-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"cmakeExecutable": "/usr/bin/cmake",
"remoteCopySourcesExclusionList": [ ".vs", ".vscode", ".git", ".github", "build", "out", "bin" ],
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "linux_clang_x64" ],
"remoteMachineName": "${env.fexremote}",
"remoteCMakeListsRoot": "$HOME/projects/.vs/${projectDirName}/src",
"remoteBuildRoot": "$HOME/projects/.vs/${projectDirName}/build/${name}",
"remoteInstallRoot": "$HOME/projects/.vs/${projectDirName}/install/${name}",
"remoteCopySources": true,
"rsyncCommandArgs": "-t --delete --delete-excluded",
"remoteCopyBuildOutput": false,
"remoteCopySourcesMethod": "rsync",
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": []
}
]
}
+1
View File
@@ -0,0 +1 @@
No AI/ML/LLM/etc code contributions.
+7
View File
@@ -46,6 +46,13 @@
"@PREFIX_LIB@/libwayland-client.so.0",
"@PREFIX_LIB@/libwayland-client.so.0.20.0"
]
},
"cuda": {
"Library" : "libcuda-guest.so",
"Overlay": [
"@PREFIX_LIB@/libcuda.so",
"@PREFIX_LIB@/libcuda.so.1"
]
}
}
}
+60 -60
View File
@@ -1,5 +1,5 @@
#
# This file is autogenerated by pip-compile with Python 3.13
# This file is autogenerated by pip-compile with Python 3.14
# by the following command:
#
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
@@ -210,56 +210,56 @@ click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==46.0.5 \
--hash=sha256:02f547fce831f5096c9a567fd41bc12ca8f11df260959ecc7c3202555cc47a72 \
--hash=sha256:039917b0dc418bb9f6edce8a906572d69e74bd330b0b3fea4f79dab7f8ddd235 \
--hash=sha256:1abfdb89b41c3be0365328a410baa9df3ff8a9110fb75e7b52e66803ddabc9a9 \
--hash=sha256:2ae6971afd6246710480e3f15824ed3029a60fc16991db250034efd0b9fb4356 \
--hash=sha256:2b7a67c9cd56372f3249b39699f2ad479f6991e62ea15800973b956f4b73e257 \
--hash=sha256:351695ada9ea9618b3500b490ad54c739860883df6c1f555e088eaf25b1bbaad \
--hash=sha256:38946c54b16c885c72c4f59846be9743d699eee2b69b6988e0a00a01f46a61a4 \
--hash=sha256:3b4995dc971c9fb83c25aa44cf45f02ba86f71ee600d81091c2f0cbae116b06c \
--hash=sha256:3ce58ba46e1bc2aac4f7d9290223cead56743fa6ab94a5d53292ffaac6a91614 \
--hash=sha256:3ee190460e2fbe447175cda91b88b84ae8322a104fc27766ad09428754a618ed \
--hash=sha256:4108d4c09fbbf2789d0c926eb4152ae1760d5a2d97612b92d508d96c861e4d31 \
--hash=sha256:420d0e909050490d04359e7fdb5ed7e667ca5c3c402b809ae2563d7e66a92229 \
--hash=sha256:47fb8a66058b80e509c47118ef8a75d14c455e81ac369050f20ba0d23e77fee0 \
--hash=sha256:4c3341037c136030cb46e4b1e17b7418ea4cbd9dd207e4a6f3b2b24e0d4ac731 \
--hash=sha256:4d7e3d356b8cd4ea5aff04f129d5f66ebdc7b6f8eae802b93739ed520c47c79b \
--hash=sha256:4d8ae8659ab18c65ced284993c2265910f6c9e650189d4e3f68445ef82a810e4 \
--hash=sha256:4e817a8920bfbcff8940ecfd60f23d01836408242b30f1a708d93198393a80b4 \
--hash=sha256:50bfb6925eff619c9c023b967d5b77a54e04256c4281b0e21336a130cd7fc263 \
--hash=sha256:556e106ee01aa13484ce9b0239bca667be5004efb0aabbed28d353df86445595 \
--hash=sha256:582f5fcd2afa31622f317f80426a027f30dc792e9c80ffee87b993200ea115f1 \
--hash=sha256:5be7bf2fb40769e05739dd0046e7b26f9d4670badc7b032d6ce4db64dddc0678 \
--hash=sha256:60ee7e19e95104d4c03871d7d7dfb3d22ef8a9b9c6778c94e1c8fcc8365afd48 \
--hash=sha256:61aa400dce22cb001a98014f647dc21cda08f7915ceb95df0c9eaf84b4b6af76 \
--hash=sha256:68f68d13f2e1cb95163fa3b4db4bf9a159a418f5f6e7242564fc75fcae667fd0 \
--hash=sha256:7d1f30a86d2757199cb2d56e48cce14deddf1f9c95f1ef1b64ee91ea43fe2e18 \
--hash=sha256:7d731d4b107030987fd61a7f8ab512b25b53cef8f233a97379ede116f30eb67d \
--hash=sha256:803812e111e75d1aa73690d2facc295eaefd4439be1023fefc4995eaea2af90d \
--hash=sha256:80a8d7bfdf38f87ca30a5391c0c9ce4ed2926918e017c29ddf643d0ed2778ea1 \
--hash=sha256:8293f3dea7fc929ef7240796ba231413afa7b68ce38fd21da2995549f5961981 \
--hash=sha256:8456928655f856c6e1533ff59d5be76578a7157224dbd9ce6872f25055ab9ab7 \
--hash=sha256:890bcb4abd5a2d3f852196437129eb3667d62630333aacc13dfd470fad3aaa82 \
--hash=sha256:94a76daa32eb78d61339aff7952ea819b1734b46f73646a07decb40e5b3448e2 \
--hash=sha256:9f16fbdf4da055efb21c22d81b89f155f02ba420558db21288b3d0035bafd5f4 \
--hash=sha256:a3d1fae9863299076f05cb8a778c467578262fae09f9dc0ee9b12eb4268ce663 \
--hash=sha256:a3d507bb6a513ca96ba84443226af944b0f7f47dcc9a399d110cd6146481d24c \
--hash=sha256:abace499247268e3757271b2f1e244b36b06f8515cf27c4d49468fc9eb16e93d \
--hash=sha256:ba2a27ff02f48193fc4daeadf8ad2590516fa3d0adeeb34336b96f7fa64c1e3a \
--hash=sha256:bc84e875994c3b445871ea7181d424588171efec3e185dced958dad9e001950a \
--hash=sha256:bfd56bb4b37ed4f330b82402f6f435845a5f5648edf1ad497da51a8452d5d62d \
--hash=sha256:c18ff11e86df2e28854939acde2d003f7984f721eba450b56a200ad90eeb0e6b \
--hash=sha256:c3bcce8521d785d510b2aad26ae2c966092b7daa8f45dd8f44734a104dc0bc1a \
--hash=sha256:c4143987a42a2397f2fc3b4d7e3a7d313fbe684f67ff443999e803dd75a76826 \
--hash=sha256:c69fd885df7d089548a42d5ec05be26050ebcd2283d89b3d30676eb32ff87dee \
--hash=sha256:ced80795227d70549a411a4ab66e8ce307899fad2220ce5ab2f296e687eacde9 \
--hash=sha256:d66e421495fdb797610a08f43b05269e0a5ea7f5e652a89bfd5a7d3c1dee3648 \
--hash=sha256:d861ee9e76ace6cf36a6a89b959ec08e7bc2493ee39d07ffe5acb23ef46d27da \
--hash=sha256:e9251e3be159d1020c4030bd2e5f84d6a43fe54b6c19c12f51cde9542a2817b2 \
--hash=sha256:f145bba11b878005c496e93e257c1e88f154d278d2638e6450d17e0f31e558d2 \
--hash=sha256:fe346b143ff9685e40192a4960938545c699054ba11d4f9029f94751e3f71d87
cryptography==48.0.0 \
--hash=sha256:0890f502ddf7d9c6426129c3f49f5c0a39278ed7cd6322c8755ffca6ee675a13 \
--hash=sha256:0c558d2cdffd8f4bbb30fc7134c74d2ca9a476f830bb053074498fbc86f41ed6 \
--hash=sha256:16cd65b9330583e4619939b3a3843eec1e6e789744bb01e7c7e2e62e33c239c8 \
--hash=sha256:18349bbc56f4743c8b12dc32e2bccb2cf83ee8b69a3bba74ef8ae857e26b3d25 \
--hash=sha256:1e2d54c8be6152856a36f0882ab231e70f8ec7f14e93cf87db8a2ed056bf160c \
--hash=sha256:22a5cb272895dce158b2cacdfdc3debd299019659f42947dbdac6f32d68fe832 \
--hash=sha256:27241b1dc9962e056062a8eef1991d02c3a24569c95975bd2322a8a52c6e5e12 \
--hash=sha256:2b4d59804e8408e2fea7d1fbaf218e5ec984325221db76e6a241a9abd6cdd95c \
--hash=sha256:2eb992bbd4661238c5a397594c83f5b4dc2bc5b848c365c8f991b6780efcc5c7 \
--hash=sha256:369a6348999f94bbd53435c894377b20ab95f25a9065c283570e70150d8abc3c \
--hash=sha256:3cb07a3ed6431663cd321ea8a000a1314c74211f823e4177fefa2255e057d1ec \
--hash=sha256:40ba1f85eaa6959837b1d51c9767e230e14612eea4ef110ee8854ada22da1bf5 \
--hash=sha256:4defde8685ae324a9eb9d818717e93b4638ef67070ac9bc15b8ca85f63048355 \
--hash=sha256:55b7718303bf06a5753dcdccf2f3945cf18ad7bffde41b61226e4db31ab89a9c \
--hash=sha256:561215ea3879cb1cbbf272867e2efda62476f240fb58c64de6b393ae19246741 \
--hash=sha256:58d00498e8933e4a194f3076aee1b4a97dfec1a6da444535755822fe5d8b0b86 \
--hash=sha256:59baa2cb386c4f0b9905bd6eb4c2a79a69a128408fd31d32ca4d7102d4156321 \
--hash=sha256:5a5ed8fde7a1d09376ca0b40e68cd59c69fe23b1f9768bd5824f54681626032a \
--hash=sha256:5b012212e08b8dd5edc78ef54da83dd9892fd9105323b3993eff6bea65dc21d7 \
--hash=sha256:5c3932f4436d1cccb036cb0eaef46e6e2db91035166f1ad6505c3c9d5a635920 \
--hash=sha256:614d0949f4790582d2cc25553abd09dd723025f0c0e7c67376a1d77196743d6e \
--hash=sha256:76341972e1eff8b4bea859f09c0d3e64b96ce931b084f9b9b7db8ef364c30eff \
--hash=sha256:77a2ccbbe917f6710e05ba9adaa25fb5075620bf3ea6fb751997875aff4ae4bd \
--hash=sha256:7995ef305d7165c3f11ae07f2517e5a4f1d5c18da1376a0a9ed496336b69e5f3 \
--hash=sha256:7ce4bfae76319a532a2dc68f82cc32f5676ee792a983187dac07183690e5c66f \
--hash=sha256:7e8eac43dfca5c4cccc6dad9a80504436fca53bb9bc3100a2386d730fbe6b602 \
--hash=sha256:84cf79f0dc8b36ac5da873481716e87aef31fcfa0444f9e1d8b4b2cece142855 \
--hash=sha256:8c7378637d7d88016fa6791c159f698b3d3eed28ebf844ac36b9dc04a14dae18 \
--hash=sha256:8cd666227ef7af430aa5914a9910e0ddd703e75f039cef0825cd0da71b6b711a \
--hash=sha256:906cbf0670286c6e0044156bc7d4af9cbb0ef6db9f73e52c3ec56ba6bdde5336 \
--hash=sha256:9071196d81abc88b3516ac8cdfad32e2b66dd4a5393a8e68a961e9161ddc6239 \
--hash=sha256:9249e3cd978541d665967ac2cb2787fd6a62bddf1e75b3e347a594d7dacf4f74 \
--hash=sha256:984a20b0f62a26f48a3396c72e4bc34c66e356d356bf370053066b3b6d54634a \
--hash=sha256:9be5aafa5736574f8f15f262adc81b2a9869e2cfe9014d52a44633905b40d52c \
--hash=sha256:9c459db21422be75e2809370b829a87eb37f74cd785fc4aa9ea1e5f43b47cda4 \
--hash=sha256:9ccdac7d40688ecb5a3b4a604b8a88c8002e3442d6c60aead1db2a89a041560c \
--hash=sha256:a0e692c683f4df67815a2d258b324e66f4738bd7a96a218c826dce4f4bd05d8f \
--hash=sha256:a5da777e32ffed6f85a7b2b3f7c5cbc88c146bfcd0a1d7baf5fcc6c52ee35dd4 \
--hash=sha256:a64697c641c7b1b2178e573cbc31c7c6684cd56883a478d75143dbb7118036db \
--hash=sha256:ad64688338ed4bc1a6618076ba75fd7194a5f1797ac60b47afe926285adb3166 \
--hash=sha256:bd72e68b06bb1e96913f97dd4901119bc17f39d4586a5adf2d3e47bc2b9d58b5 \
--hash=sha256:c17dfe85494deaeddc5ce251aebd1d60bbe6afc8b62071bb0b469431a000124f \
--hash=sha256:c18684a7f0cc9a3cb60328f496b8e3372def7c5d2df39ac267878b05565aaaae \
--hash=sha256:cc90c0b39b2e3c65ef52c804b72e3c58f8a04ab2a1871272798e5f9572c17d20 \
--hash=sha256:db63bf618e5dea46c07de12e900fe1cdd2541e6dc9dbae772a70b7d4d4765f6a \
--hash=sha256:ea8990436d914540a40ab24b6a77c0969695ed52f4a4874c5137ccf7045a7057 \
--hash=sha256:ecde28a596bead48b0cfd2a1b4416c3d43074c2d785e3a398d7ec1fc4d0f7fbb \
--hash=sha256:f5333311663ea94f75dd408665686aaf426563556bb5283554a3539177e03b8c \
--hash=sha256:fdfef35d751d510fcef5252703621574364fec16418c4a1e5e1055248401054b
# via
# -r requirements_formatting.txt.in
# pyjwt
@@ -281,9 +281,9 @@ graylint==1.1.1 \
--hash=sha256:0fd8e02972ca03d0ef2bf0adea76b5343efcd492d7afb5f658f3e3a724f55a36 \
--hash=sha256:b7e0eab6c159684dbf5ef84e942c3340f6a6549b02a3d11b1a1763cc4f8f0593
# via darker
idna==3.10 \
--hash=sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9 \
--hash=sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3
idna==3.16 \
--hash=sha256:cc246e3a3f89580c3a951b5ad298ca4638078b2cdd4f115654332b5c26daded5 \
--hash=sha256:d7a6da03db833450fca25d2358ac9ff06cd624577a4aea3a596d5c0f77b8e03d
# via
# -r requirements_formatting.txt.in
# requests
@@ -390,9 +390,9 @@ pytokens==0.4.1 \
--hash=sha256:ee44d0f85b803321710f9239f335aafe16553b39106384cef8e6de40cb4ef2f6 \
--hash=sha256:f66a6bbe741bd431f6d741e617e0f39ec7257ca1f89089593479347cc4d13324
# via black
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# -r requirements_formatting.txt.in
# pygithub
@@ -406,9 +406,9 @@ typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.6.3 \
--hash=sha256:1b62b6884944a57dbe321509ab94fd4d3b307075e0c2eae991ac71ee15ad38ed \
--hash=sha256:bf272323e553dfb2e87d9bfd225ca7b0f467b919d7bbd355436d3fd37cb0acd4
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via
# -r requirements_formatting.txt.in
# pygithub
+4 -4
View File
@@ -1,10 +1,10 @@
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=46.0.5
urllib3>=2.6.3
requests>=2.32.4
idna>=3.7
cryptography>=46.0.7
urllib3>=2.7.0
requests>=2.33.0
idna>=3.15
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.12.1
+19
View File
@@ -260,6 +260,10 @@ struct FEX_PACKED X80SoftFloat {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
// Inf/NaN pass through unchanged in the significand slot.
if (lhs.Top.Exponent == 0x7FFF) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Top.Exponent = 0x3FFF;
Tmp.Top.Sign = lhs.Top.Sign;
@@ -288,6 +292,14 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
// +/-Inf returns +Inf in the exponent slot; NaN propagates.
if (lhs.Top.Exponent == 0x7FFF) {
if ((lhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
X80SoftFloat Result(0, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
return lhs;
}
int32_t TrueExp = lhs.Top.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
@@ -324,6 +336,13 @@ struct FEX_PACKED X80SoftFloat {
#else
extFloat80_t Zero {0, 0};
if (extF80_eq(state, lhs, Zero)) {
// FSCALE(0, +Inf) is 0 * Inf, which is invalid. FSCALE(0, anything
// else) is still 0.
if (rhs.Top.Exponent == 0x7FFF && rhs.Top.Sign == 0 && (rhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
X80SoftFloat QNaN(0, 0x7FFFUL, 0xC000000000000000ULL);
return QNaN;
}
return lhs;
}
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
@@ -23,6 +23,13 @@
"Enable the code caching subsystem"
]
},
"EnableLazyCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable lazy loading of chunks in code caches"
]
},
"EnableCodeCacheValidation": {
"Type": "bool",
"Default": "false",
+1 -1
View File
@@ -53,6 +53,6 @@ FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionN
}
bool FEXCore::Context::ContextImpl::IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const {
return Thread->CPUBackend->IsAddressInCodeBuffer(Address);
return Thread->CPUBackend->IsAddressInCodeBuffer(Address) || CodeCache.IsAddressInMappedCodeBuffer(Address);
}
} // namespace FEXCore::Context
+89 -2
View File
@@ -64,6 +64,8 @@ struct CustomIRResult {
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
constexpr static bool BLOCK_DEBUGGING = false;
class CodeCache : public AbstractCodeCache {
public:
CodeCache(ContextImpl&);
@@ -76,11 +78,17 @@ public:
bool IsGeneratingCache = false;
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
FEX_CONFIG_OPT(EnableLazyCodeCaching, ENABLELAZYCODECACHINGWIP);
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
fextl::unique_ptr<MappedCodeCacheFile> LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo&, uint64_t FileStartVA) override;
bool EnableLoadedSection(Core::InternalThreadState*, MappedCodeCacheFile&, const ExecutableFileSectionInfo&) override;
void FinalizeCodePages(MappedCodeCacheFile&, std::span<std::byte> CodeRange) override;
/**
* Performs expensive extra validation on the loaded code cache data.
@@ -112,12 +120,14 @@ public:
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param RelocationOffset Offset to subtract from relocation target offsets
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations,
uint32_t RelocationOffset, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -233,6 +243,83 @@ public:
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
// Manual debugging tooling which is useful for developers.
struct TrackingEmpty {
// RIP stepping handling
virtual void AddSingleStepTarget(uint64_t GuestRIP) {}
virtual void AllTargetSingleStep() {}
virtual void RemoveSingleStepTarget(uint64_t GuestRIP) {}
virtual bool IsSingleStepTarget(uint64_t GuestRIP) {
return false;
}
// Watchpoints
virtual void AddWriteWatchPoint(uint64_t Ptr) {}
virtual void AddReadWatchPoint(uint64_t Ptr) {}
virtual bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
virtual bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
};
struct TrackingPossible final : public TrackingEmpty {
void AddSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.emplace(GuestRIP);
}
void RemoveSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.erase(GuestRIP);
}
void AllTargetSingleStep() override {
SingleStepEverything = true;
}
bool IsSingleStepTarget(uint64_t GuestRIP) override {
return SingleStepEverything || SingleStepTargets.contains(GuestRIP);
}
void AddWriteWatchPoint(uint64_t Ptr) override {
WatchWriteTargets.emplace(Ptr);
}
void AddReadWatchPoint(uint64_t Ptr) override {
WatchReadTargets.emplace(Ptr);
}
bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchWriteTargets, Ptr, Size);
}
bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchReadTargets, Ptr, Size);
}
private:
bool SingleStepEverything {};
fextl::set<uint64_t> SingleStepTargets {};
fextl::set<uint64_t> WatchWriteTargets {};
fextl::set<uint64_t> WatchReadTargets {};
static bool ContainsRange(const fextl::set<uint64_t>& Set, uint64_t Ptr, size_t Size) {
for (auto it = Set.lower_bound(Ptr); it != Set.end(); --it) {
auto Watch = *it;
if (Watch < Ptr) {
break;
}
if (Watch >= Ptr && Watch < (Ptr + Size)) {
return true;
}
}
return false;
}
};
using TrackingStructure = std::conditional<BLOCK_DEBUGGING, TrackingPossible, TrackingEmpty>::type;
TrackingStructure BlockDebuggerTracker {};
public:
struct {
uint64_t VirtualMemSize {1ULL << 36};
@@ -586,8 +586,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
}};
for (auto& RegPair : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, -16);
for (const auto& [rt, rt2] : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(rt, rt2, ARMEmitter::Reg::rsp, -16);
}
// Additionally we need to store the lower 64bits of v8-v15
@@ -604,9 +604,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We just saved x19 so it is safe
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r19, ARMEmitter::Reg::rsp, 0);
for (auto& RegQuad : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::r19, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::r19, 32);
}
}
@@ -616,9 +615,8 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
}};
for (auto& RegQuad : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::rsp, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::rsp, 32);
}
constexpr static std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
@@ -630,8 +628,8 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
}};
for (auto& RegPair : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, 16);
for (const auto& [rt, rt2] : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(rt, rt2, ARMEmitter::Reg::rsp, 16);
}
}
@@ -661,7 +659,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
#endif
if (SetPredRegs && (EmitterCTX->HostFeatures.SupportsSVE256 || EmitterCTX->HostFeatures.SupportsSVE128)) {
if (SetPredRegs && EmitterCTX->HostFeatures.SupportsSVE()) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
+2 -1
View File
@@ -12,7 +12,6 @@
#include <cstdint>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -44,6 +43,8 @@ namespace CPU {
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB_UPPER
{0x0706'0504'0302'0100ULL, 0x1716'1514'1312'1110ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP
{0x0F0E'0D0C'0B0A'0908ULL, 0x1F1E'1D1C'1B1A'1918ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_ONE
{0xD49A'784B'CD1B'8AFEULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_LOG2_10
{0xB8AA'3B29'5C17'F0BCULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_LOG2_E
+3 -1
View File
@@ -92,6 +92,7 @@ namespace ProductNames {
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_ORYON_3[] = "Oryon-3";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
@@ -186,8 +187,9 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 67> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 68> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x002, 1, ProductNames::ARM_ORYON_3}, // Qualcomm Oryon-3
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x039, 1, ProductNames::ARM_Avalanche_M2Max}, // Apple Avalanche (M2 Max)
+402 -185
View File
@@ -1,4 +1,9 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/Utils/LogManager.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Utils/TypeDefines.h"
#include "FEXCore/fextl/memory.h"
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SpinWaitLock.h>
#include <Interface/Context/Context.h>
@@ -16,10 +21,14 @@
#include <FEXHeaderUtils/Filesystem.h>
#include <algorithm>
#include <git_version.h>
#include <span>
#include <xxhash.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <fstream>
namespace FEXCore {
@@ -32,6 +41,38 @@ ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
MappedCodeCacheFile::~MappedCodeCacheFile() {
if (CacheManager) {
CacheManager->UnregisterMappedCodeBuffer(*this);
}
#ifndef _WIN32
if (!CodeBuffer.empty()) {
FEXCore::Allocator::munmap(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#endif
}
void AbstractCodeCache::RegisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
MappedCodeBuffers.push_back(Code.CodeBuffer);
// Unregister on destruction of Code
Code.CacheManager = this;
}
void AbstractCodeCache::UnregisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
std::erase_if(MappedCodeBuffers, [&](const auto& Elem) { return Elem.data() == Code.CodeBuffer.data(); });
}
bool AbstractCodeCache::IsAddressInMappedCodeBuffer(uintptr_t Address) const {
for (const auto& Range : MappedCodeBuffers) {
auto Start = reinterpret_cast<uintptr_t>(Range.data());
if (Address >= Start && Address < Start + Range.size_bytes()) {
return true;
}
}
return false;
}
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
@@ -233,7 +274,10 @@ uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
struct CodeCacheHeader {
std::array<char, 4> Magic = ExpectedMagic;
uint32_t FormatVersion = 1;
// Version history:
// 1: Initial version
// 2: Padding code buffer data to enable direct mapping
uint32_t FormatVersion = 2;
uint8_t FEXVersion[20] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
@@ -260,7 +304,7 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
std::ranges::copy(GIT_HASH, header.FEXVersion);
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.CodeBufferSize = FEXCore::AlignUp(CTX.LatestOffset, Utils::FEX_PAGE_SIZE);
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
@@ -300,21 +344,25 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
{
auto AlignedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, AlignedSize);
lseek(fd, AlignedSize, SEEK_SET);
}
// Dump the host code (relocated for position-independent serialization)
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, 0, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Pad to next page in file for mmap
{
auto PaddedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, PaddedSize);
lseek(fd, PaddedSize, SEEK_SET);
}
// Dump code pages
static_assert(OrderedContainer<decltype(LookupCache.CodePages)>, "Non-deterministic data source");
@@ -332,178 +380,6 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
return true;
}
bool CodeCache::LoadData(Core::InternalThreadState* Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
// Read file header
CodeCacheHeader header {};
::memcpy(&header, MappedCacheFile, sizeof(header));
MappedCacheFile += sizeof(header);
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", header.NumBlocks, BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (!ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return false;
}
if (!ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return false;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return false;
}
// Read guest<->host block mappings
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(header.NumBlocks);
{
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, MappedCacheFile, sizeof(BlockPtr.first));
MappedCacheFile += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, MappedCacheFile, sizeof(BlockPtr.second.HostCode));
MappedCacheFile += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, MappedCacheFile, sizeof(NumGuestPages));
MappedCacheFile += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), MappedCacheFile, std::span {BlockPtr.second.CodePages}.size_bytes());
MappedCacheFile += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Consistency check: VMA regions at the top and end should belong to the same file
auto [min_val, max_val] = ranges::minmax_element(BlockList, std::less {}, &decltype(BlockList)::value_type::first);
auto MinBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, min_val->first + BinarySection.FileStartVA);
auto MaxBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, max_val->first + BinarySection.FileStartVA);
if (&MinBound->FileInfo != &BinarySection.FileInfo || &MaxBound->FileInfo != &BinarySection.FileInfo) {
ERROR_AND_DIE_FMT("Cached blocks offsets {:#x}-{:#x} out of bounds for guest library {} ({:016x} @ {:#x}) while trying to load "
"section {:#x}-{:#x}!",
min_val->first, max_val->first, BinarySection.FileInfo.Filename, BinarySection.FileInfo.FileId,
BinarySection.FileStartVA, BinarySection.BeginVA, BinarySection.EndVA);
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
if (begin == end) {
// Not an error since there is just no data to load
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
// Read relocations
fextl::vector<FEXCore::CPU::Relocation> Relocations(header.NumRelocations, FEXCore::CPU::Relocation::Default());
::memcpy(Relocations.data(), MappedCacheFile, Relocations.size() * sizeof(Relocations[0]));
MappedCacheFile += Relocations.size() * sizeof(Relocations[0]);
// Pad to next page in file, which contains CodeBuffer data
MappedCacheFile = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(MappedCacheFile), Utils::FEX_PAGE_SIZE));
// Prepare CodeBuffer: Page aligned and big enough to hold all cached data
auto Lock = std::unique_lock {CTX.CodeBufferWriteMutex};
if (Thread) {
if (auto Prev = Thread->CPUBackend->CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = Thread->LookupCache->AcquireWriteLock();
Thread->LookupCache->ChangeGuestToHostMapping(*Prev, *CTX.GetLatest()->LookupCache, lk);
}
}
auto CodeBuffer = CTX.GetLatest();
LOGMAN_THROW_A_FMT(reinterpret_cast<uintptr_t>(CodeBuffer->Ptr) % 0x1000 == 0, "Expected CodeBuffer base to be page-aligned");
const auto Delta = AlignUp(CTX.LatestOffset, 0x1000) - CTX.LatestOffset;
CTX.LatestOffset += Delta;
while (CTX.LatestOffset + header.CodeBufferSize > CodeBuffer->UsableSize()) {
if (Thread) {
CTX.ClearCodeCache(Thread);
CodeBuffer = CTX.GetLatest();
LogMan::Msg::IFmt("Increased code buffer size to {} MiB for cache load", CodeBuffer->AllocatedSize / 1024 / 1024);
} else {
ERROR_AND_DIE_FMT("Cannot extend codebuffer without thread!");
}
}
// Read CodeBuffer data from file. Make sure the destination is page-aligned.
// TODO: Only load the data needed for the selected section
auto CodeBufferRange =
std::as_writable_bytes(std::span {CodeBuffer->Ptr, CodeBuffer->UsableSize()}).subspan(CTX.LatestOffset, header.CodeBufferSize);
::memcpy(CodeBufferRange.data(), MappedCacheFile, header.CodeBufferSize);
MappedCacheFile += header.CodeBufferSize;
CTX.LatestOffset += header.CodeBufferSize;
// Apply FEX relocations
auto Ret = ApplyCodeRelocations(BinarySection.FileStartVA, CodeBufferRange, Relocations, false);
LOGMAN_THROW_A_FMT(Ret == true, "Failed to apply code cache relocations");
{
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
// Register blocks to LookupCache
for (auto& [Guest, Host] : BlockList) {
for (auto& CodePage : Host.CodePages) {
CodePage += BinarySection.FileStartVA;
}
auto HostCode = reinterpret_cast<void*>(Host.HostCode + reinterpret_cast<uintptr_t>(CodeBufferRange.data()));
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Host.CodePages), HostCode, WriteLock);
}
// Register loaded code ranges
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < header.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, MappedCacheFile, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
MappedCacheFile += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, MappedCacheFile, sizeof(NumEntrypoints));
MappedCacheFile += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), MappedCacheFile, NumEntrypoints * sizeof(Entrypoints[0]));
MappedCacheFile += NumEntrypoints * sizeof(Entrypoints[0]);
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, CodeBufferRange);
}
return true;
}
void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<uint64_t> GuestBlocks, const fextl::set<uint64_t>& HostBlocks,
std::span<std::byte> CachedCode) {
LOGMAN_THROW_A_FMT(!HostBlocks.empty(), "Tried to validate without any host blocks");
@@ -558,7 +434,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
NewRelocations.erase(std::remove_if(NewRelocations.begin(), NewRelocations.end(), [](const CPU::Relocation& Reloc) {
return Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL && Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
}));
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, false);
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, 0, false);
if (ValidationCTX->LatestOffset <= CodeBufferRangeRef.size()) {
// Reference compilation produced fewer bytes than our cache, so validation is going to fail.
@@ -616,15 +492,17 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
ValidationThread->LookupCache->ClearCache(ValidationThread->LookupCache->AcquireWriteLock());
ValidationCTX->LatestOffset = 0;
LogMan::Msg::IFmt("\tSuccessfully validated cache");
LogMan::Msg::IFmt(" successfully validated cache");
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
std::span<const FEXCore::CPU::Relocation> EntryRelocations, uint32_t RelocationOffset, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
LOGMAN_THROW_A_FMT(Reloc.Header.Offset >= RelocationOffset, "Invalid relocation offset");
LOGMAN_THROW_A_FMT(Reloc.Header.Offset - RelocationOffset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Header.Offset - RelocationOffset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
@@ -663,4 +541,343 @@ bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> C
return true;
}
fextl::unique_ptr<MappedCodeCacheFile>
CodeCache::LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo& FileInfo, uint64_t FileStartVA) {
if (!EnableCodeCaching) {
return nullptr;
}
FEXCORE_PROFILE_SCOPED("LoadCache");
// Read file header
CodeCacheHeader header {};
::memcpy(&header, CacheFile.data(), sizeof(header));
if (!std::ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return nullptr;
}
if (!std::ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return nullptr;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return nullptr;
}
// Skip over BlockEntry data since it won't be used until EnableLoadedSection
// TODO: Store direct offset to relocations in the header
auto* BlockListStart = CacheFile.data() + sizeof(header);
auto* Cursor = BlockListStart;
for (uint32_t i = 0; i < header.NumBlocks; ++i) {
Cursor += sizeof(uint64_t); // guest address
Cursor += sizeof(uint64_t); // host code address
uint64_t NumGuestCodePages;
::memcpy(&NumGuestCodePages, Cursor, sizeof(NumGuestCodePages));
Cursor += sizeof(NumGuestCodePages);
Cursor += NumGuestCodePages * sizeof(uint64_t);
}
auto Relocations = std::span {reinterpret_cast<const FEXCore::CPU::Relocation*>(Cursor), header.NumRelocations};
Cursor += Relocations.size_bytes();
// Pad to next page to get the code buffer data
Cursor = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(Cursor), Utils::FEX_PAGE_SIZE));
auto CodeDataInFile = std::span {Cursor, header.CodeBufferSize};
#ifndef _WIN32
// Allocate target memory for post-relocation code. This is PROT_NONE until
// the first execution, so that contents can be lazily populated in a
// frontend-provided segfault handler.
void* CodeBufferAllocation = Allocator::mmap(nullptr, header.CodeBufferSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (CodeBufferAllocation == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to reserve target memory for code cache");
return nullptr;
}
auto CodeBuffer = std::span {static_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#else
// TODO: Implement lazy mapping on Windows
auto CodeBuffer = CodeDataInFile;
#endif
// Group relocations by page
size_t NumPages = header.CodeBufferSize / Utils::FEX_PAGE_SIZE;
fextl::vector<MappedCodeCacheFile::PageRelocationRange> PageRelocationRanges(NumPages, {0, 0});
auto RelocBaseOffset = std::as_bytes(Relocations).data() - CacheFile.data();
auto RelocIt = Relocations.begin();
for (size_t Page = 0; Page < NumPages; ++Page) {
auto EndRelocIt = std::upper_bound(RelocIt, Relocations.end(), Page,
[](auto& Page, auto& Reloc) { return Page < Reloc.Header.Offset / Utils::FEX_PAGE_SIZE; });
PageRelocationRanges.at(Page) = {static_cast<uint32_t>(RelocBaseOffset + (RelocIt - Relocations.begin()) * sizeof(CPU::Relocation)),
static_cast<uint32_t>(EndRelocIt - RelocIt)};
RelocIt = EndRelocIt;
}
auto Storage = FEXCore::Allocator::aligned_alloc(alignof(MappedCodeCacheFile), sizeof(MappedCodeCacheFile));
return fextl::unique_ptr<MappedCodeCacheFile>(
new (Storage) MappedCodeCacheFile {this, CacheFile, CodeDataInFile, CodeBuffer, BlockListStart, header.NumBlocks, header.NumCodePages,
std::move(PageRelocationRanges), fextl::vector<bool>(NumPages), FileStartVA});
}
bool CodeCache::EnableLoadedSection(Core::InternalThreadState* Thread, MappedCodeCacheFile& Code, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
FEXCORE_PROFILE_SCOPED("EnableLoadedSection");
// Read block list from cache file
// TODO: Store section-ized BlockLists in cache file
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(Code.NumBlocks);
{
auto* Cursor = Code.BlockListInFile;
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, Cursor, sizeof(BlockPtr.first));
Cursor += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, Cursor, sizeof(BlockPtr.second.HostCode));
Cursor += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, Cursor, sizeof(NumGuestPages));
Cursor += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), Cursor, std::span {BlockPtr.second.CodePages}.size_bytes());
Cursor += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
if (begin == end) {
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", BlockList.size(), BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (EnableLazyCodeCaching) {
LogMan::Msg::IFmt(" lazy mapping: base={:#14x} -> host={}; cache_source={}", BinarySection.FileStartVA,
fmt::ptr(Code.CodeBuffer.data()), fmt::ptr(Code.MappedFile.data()));
}
// Register blocks to LookupCache.
// The host addresses will point into the protected code buffer, so that FEX
// can lazily apply relocations on first execution of each page.
auto CodeBuffer = CTX.GetLatest();
{
FEXCORE_PROFILE_SCOPED("Decode");
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
for (auto& [Guest, Block] : BlockList) {
for (auto& CodePage : Block.CodePages) {
CodePage += BinarySection.FileStartVA;
}
LOGMAN_THROW_A_FMT(Block.HostCode < Code.CodeBuffer.size_bytes(), "Host offset {:#x} out of range ({:#x})", Block.HostCode,
Code.CodeBuffer.size_bytes());
auto HostCode = &Code.CodeBuffer[Block.HostCode];
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Block.CodePages), HostCode, WriteLock);
}
// Guest code pages
auto* Cursor = Code.CodeBufferInFile.data() + Code.CodeBufferInFile.size_bytes();
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < Code.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, Cursor, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
Cursor += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, Cursor, sizeof(NumEntrypoints));
Cursor += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), Cursor, std::span {Entrypoints}.size_bytes());
Cursor += std::span {Entrypoints}.size_bytes();
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
#ifndef _WIN32
if (!EnableLazyCodeCaching || EnableCodeCacheValidation) {
#else
// TODO: Implement lazy mapping on Windows
if (true) {
#endif
auto Range = SelectCodeRangeToFinalize(Code, 0, Code.CodeBuffer.size_bytes() / Utils::FEX_PAGE_SIZE);
FinalizeCodePages(Code, Range);
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, Code.CodeBuffer);
}
return true;
}
} // namespace FEXCore::Context
namespace FEXCore {
static std::span<CPU::Relocation> SpanPageRelocations(const MappedCodeCacheFile& Code, size_t PageIndex) {
auto [Offset, Count] = Code.PageRelocationRanges.at(PageIndex);
return std::span {reinterpret_cast<FEXCore::CPU::Relocation*>(Code.MappedFile.data() + Offset), Count};
}
std::span<std::byte> AbstractCodeCache::SelectCodeRangeToFinalize(MappedCodeCacheFile& Code, size_t StartPage, size_t EndPage) {
// First, check if we were racing another thread in loading this range
if (std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, false) == Code.LoadedPages.begin() + EndPage) {
return {};
}
LOGMAN_THROW_A_FMT(StartPage < EndPage, "Invalid page range [{}, {})", StartPage, EndPage);
LOGMAN_THROW_A_FMT(EndPage <= Code.NumPages(), "End page {} out of range ({})", EndPage, Code.NumPages());
// Include any pages that have relocations or block link records crossing
// into the current page range. This ensures we don't attempt to finalize
// any page twice, partially apply FEX relocations, or trigger page loads
// during block linking.
while (EndPage < Code.NumPages()) {
auto PageRelocs = SpanPageRelocations(Code, EndPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > EndPage * Utils::FEX_PAGE_SIZE) {
++EndPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, EndPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < EndPage * Utils::FEX_PAGE_SIZE + 0x18) {
++EndPage;
continue;
}
}
break;
};
while (StartPage != 0) {
auto PageRelocs = SpanPageRelocations(Code, StartPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > StartPage * Utils::FEX_PAGE_SIZE) {
--StartPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, StartPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < StartPage * Utils::FEX_PAGE_SIZE + 0x18) {
--StartPage;
continue;
}
}
break;
};
return Code.CodeBuffer.subspan(StartPage * Utils::FEX_PAGE_SIZE, (EndPage - StartPage) * Utils::FEX_PAGE_SIZE);
}
} // namespace FEXCore
namespace FEXCore::Context {
void CodeCache::FinalizeCodePages(MappedCodeCacheFile& Code, std::span<std::byte> CodeRange) {
const size_t StartOffset = CodeRange.data() - Code.CodeBuffer.data();
const auto StartPage = StartOffset / Utils::FEX_PAGE_SIZE;
const auto EndPage = StartPage + CodeRange.size_bytes() / Utils::FEX_PAGE_SIZE;
const size_t Size = CodeRange.size_bytes();
// None of the selected pages should be loaded at all; otherwise, SelectCodeRangeToFinalize returned inconsistent ranges
LOGMAN_THROW_A_FMT(std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, true) == Code.LoadedPages.begin() + EndPage,
"Inconsistent page load state");
FEXCORE_PROFILE_SCOPED("FinalizeCodePages");
#ifndef _WIN32
// Atomicity is critical when making the finalized code data visible.
// We ensure this by remapping a temporary buffer onto the PROT_NONE
// placeholder page in CodeBuffer. Some constraints to keep in mind are:
// 1. Pages can't be write-only (readability is implicitly added), so
// we can't change CodeBuffer from PROT_NONE to PROT_WRITE even for just
// a short duration
// 2. Naive mremap from CodeBufferInFile to CodeBuffer would leave a gap in
// the former, which would make cleanup overly complicated
//
// Due to (1), we can't apply relocations in place (CodeBufferInFile); at
// least a secondary buffer is needed for execution (CodeBuffer).
// Due to (2), a third buffer is temporarily allocated here and freed on
// completion. The final code data is computed here and then the memory
// is remapped onto CodeBuffer.
auto* Staging = reinterpret_cast<std::byte*>(Allocator::VirtualAlloc(nullptr, Size, true));
if (!Staging) {
ERROR_AND_DIE_FMT("Failed to allocate {} bytes of staging memory for code-cache finalization", Size);
}
// Copy code from the cache file to the staging buffer
memcpy(Staging, Code.CodeBufferInFile.data() + StartOffset, Size);
// Apply relocations
auto StagingSpan = std::span {Staging, Size};
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, StagingSpan, PageRelocations, static_cast<uint32_t>(StartOffset), false);
Code.LoadedPages[i] = true;
}
// Atomically make the finalized code data visible by remapping the staging
// buffer onto the requested CodeBuffer window. MREMAP_DONTUNMAP is used to
// leave the old VA range reserved so that we can cleanly deallocate it
// through Allocator.
void* RemapResult = ::mremap(Staging, Size, Size, MREMAP_FIXED | MREMAP_MAYMOVE | MREMAP_DONTUNMAP, CodeRange.data());
if (RemapResult == MAP_FAILED) {
ERROR_AND_DIE_FMT("{}: mremap failed: {}", __FUNCTION__, errno);
}
Allocator::VirtualFree(Staging, Size);
// Release resident file pages that will no longer be needed. The VA range is left allocated to allow cleanup with a single VirtualFree.
Allocator::VirtualDontNeed(Code.CodeBufferInFile.data() + StartOffset, Size);
#else
// TODO: Implement lazy mapping on Windows
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, Code.CodeBuffer, PageRelocations, 0, false);
Code.LoadedPages[i] = true;
}
#endif
ARMEmitter::Emitter::ClearICache(CodeRange.data(), Size);
}
} // namespace FEXCore::Context
+26 -2
View File
@@ -358,6 +358,16 @@ bool ContextImpl::InitCore() {
Config.NeedsPendingInterruptFaultCheck = true;
}
if constexpr (BLOCK_DEBUGGING) {
// If the developer wants to do any single-stepping points or watch points.
// Add them here.
//
// eg:
// BlockDebuggerTracker.AllTargetSingleStep();
// BlockDebuggerTracker.AddSingleStepTarget(0x14000'0000ULL);
// BlockDebuggerTracker.AddWriteWatchPoint(0x420BA5ED);
}
return true;
}
@@ -376,7 +386,7 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this, Thread);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
@@ -456,7 +466,6 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
if (Config.StrictInProcessSplitLocks) {
FEXCore::Utils::SpinWaitLock::unlock(&StrictSplitLockMutex);
}
return;
}
}
@@ -641,6 +650,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->StartNewBlock();
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
@@ -648,6 +658,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
Thread->OpDispatcher->StartNewBlock();
}
if (TableInfo && TableInfo->OpcodeDispatcher.OpDispatch) {
@@ -695,6 +706,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST ||
Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::BAD_RELOCATION) {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
} else if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::UNIMPLEMENTED_INST) {
Thread->OpDispatcher->UnimplementedOp(DecodedInfo);
} else {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
}
@@ -818,6 +831,17 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst) {
if constexpr (BLOCK_DEBUGGING) {
// Block debugging logic is hand-written and needs to be handled with care.
// Force MaxInst to only be one in this case.
MaxInst = 1;
// If the entrypoint is part of the single step targets then single step it.
if (BlockDebuggerTracker.IsSingleStepTarget(GuestRIP)) {
return CompileSingleStep(Frame, GuestRIP);
}
}
auto Thread = Frame->Thread;
FEXCORE_PROFILE_SCOPED("CompileBlock");
FEXCORE_PROFILE_ACCUMULATION(Thread, AccumulatedJITTime);
@@ -153,9 +153,8 @@ void Dispatcher::EmitDispatcher() {
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 18);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 3);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
(void)tbz(TMP1, 0, &l_NotECCode);
@@ -549,6 +548,15 @@ void Dispatcher::EmitDispatcher() {
EmitF64Sin();
EmitF64Cos();
EmitF64Tan();
EmitF64F2XM1();
EmitF64Scale();
EmitF64Atan();
F64Log2Constants Log2C;
EmitF64FYL2X(Log2C);
EmitF64FYL2XP1(Log2C);
EmitF64Log2Constants(Log2C);
EmitF64FPREM();
EmitF64FPREM1();
// Interpreter fallbacks
{
@@ -1246,6 +1254,935 @@ void Dispatcher::EmitF64Tan() {
dc64(0x4160'0000'0000'0000ULL); // 2^23
}
void Dispatcher::EmitF64Scale() {
// Computes result = src1 * 2^trunc(src2).
// Input: VTMP1 = base (src1), VTMP2 = exponent (src2). Output: VTMP1.
F64ScaleHandlerAddress = GetCursorAddress<uint64_t>();
ARMEmitter::ForwardLabel Fallback;
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// n = trunc(src2).
frintz(VTMP2.D(), VTMP2.D());
// NaN check: NaN != NaN sets V flag.
fcmp(VTMP2.D(), VTMP2.D());
(void)b(ARMEmitter::Condition::CC_VS, &Fallback);
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
// Range check: int_n in [-1022, 1023].
cmn(ARMEmitter::Size::i64Bit, TMP1, 1022);
(void)b(ARMEmitter::Condition::CC_LT, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP1, 1023);
(void)b(ARMEmitter::Condition::CC_GT, &Fallback);
// 2^n, then result = src1 * 2^n.
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1023);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 52);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ret();
// Fallback path.
(void)Bind(&Fallback);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SCALE].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SCALE].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
void Dispatcher::EmitF64F2XM1() {
// JIT-inlined double-precision 2^x - 1 for x in [-1, 1].
// Uses argument reduction: split x into n = round(x) and r = x - n,
// then compute 2^x - 1 = 2^n * (2^r - 1) + (2^n - 1).
// 2^r - 1 is approximated via a 13-term Horner polynomial in r.
// Input in VTMP1.D(), output in VTMP1.D().
F64F2XM1HandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel OneLabel;
ARMEmitter::ForwardLabel C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
ARMEmitter::ForwardLabel C7Label, C8Label, C9Label, C10Label, C11Label, C12Label, C13Label;
// Save q2.
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Range check: |x| > 1.0 -> fallback.
fabs(VTMP2.D(), VTMP1.D());
ldr(Accum.D(), &OneLabel);
fcmp(VTMP2.D(), Accum.D());
(void)b(ARMEmitter::Condition::CC_HI, &Fallback);
// Argument reduction: n = round(x), r = x - n.
frinta(VTMP2.D(), VTMP1.D());
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
fsub(Accum.D(), VTMP1.D(), VTMP2.D());
// scale = 2^n, scale_m1 = 2^n - 1.
// TMP1 = scale bits, TMP3 = scale_m1 bits, Accum = r.
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1023);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 52);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
ldr(VTMP1.D(), &OneLabel);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP1.D());
// Horner polynomial: p = c1 + r * (c2 + r * (... + r * c13)).
ldr(VTMP1.D(), &C13Label);
ldr(VTMP2.D(), &C12Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C11Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C10Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C9Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C8Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C7Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C6Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C5Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C4Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C3Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C2Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C1Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// q = r * p, then result = scale * q + scale_m1.
fmul(VTMP1.D(), Accum.D(), VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, Accum.D(), TMP1);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
// Restore q2 and return.
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path.
(void)Bind(&Fallback);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64F2XM1].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64F2XM1].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool
Align(16);
(void)Bind(&OneLabel);
dc64(0x3FF0'0000'0000'0000ULL); // 1.0
(void)Bind(&C1Label);
dc64(0x3FE6'2E42'FEFA'39EFULL); // ln(2)
(void)Bind(&C2Label);
dc64(0x3FCE'BFBD'FF82'C58EULL);
(void)Bind(&C3Label);
dc64(0x3FAC'6B08'D704'A0BEULL);
(void)Bind(&C4Label);
dc64(0x3F83'B2AB'6FBA'4E76ULL);
(void)Bind(&C5Label);
dc64(0x3F55'D87F'E78A'672FULL);
(void)Bind(&C6Label);
dc64(0x3F24'3091'2F86'C785ULL);
(void)Bind(&C7Label);
dc64(0x3EEF'FCBF'C588'B0C2ULL);
(void)Bind(&C8Label);
dc64(0x3EB6'2C02'23A5'C821ULL);
(void)Bind(&C9Label);
dc64(0x3E7B'5253'D395'E7C0ULL);
(void)Bind(&C10Label);
dc64(0x3E3E'4CF5'158B'8EC5ULL);
(void)Bind(&C11Label);
dc64(0x3DFE'8CAC'7351'BB20ULL);
(void)Bind(&C12Label);
dc64(0x3DBC'3BD6'50FC'2981ULL);
(void)Bind(&C13Label);
dc64(0x3D78'1619'3166'D0F5ULL);
}
// JIT-inlined double-precision atan2 for the F64 reduced precision x87 path.
// Input: VTMP1 = y, VTMP2 = x. Output: VTMP1 = atan2(y, x).
// Algorithm: 20-term Horner polynomial from ARM optimized-routines atan_data.c.
// atan(z) = z + z^3 * P(z^2), with range reduction to [0,1] and quadrant adjustment.
void Dispatcher::EmitF64Atan() {
F64AtanHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel NoSwap, PosX, NegY;
ARMEmitter::ForwardLabel PiOver2Label, PiLabel;
ARMEmitter::ForwardLabel C0Label, C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
ARMEmitter::ForwardLabel C7Label, C8Label, C9Label, C10Label, C11Label, C12Label, C13Label;
ARMEmitter::ForwardLabel C14Label, C15Label, C16Label, C17Label, C18Label, C19Label;
// Stack layout: [sp] = q2 (16B), [sp+16] = y bits (8B), [sp+24] = x bits (8B).
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -32);
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP1.D());
str(TMP2, ARMEmitter::Reg::rsp, 16);
str(TMP1, ARMEmitter::Reg::rsp, 24);
// Compute |x|, |y|, pack sign/swap flags into TMP1 = (sign_x<<2)|(sign_y<<1)|swap.
fabs(Accum.D(), VTMP2.D());
fabs(VTMP2.D(), VTMP1.D());
fcmp(VTMP2.D(), Accum.D());
lsr(ARMEmitter::Size::i64Bit, TMP1, TMP1, 63);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP2, 63);
cset(ARMEmitter::Size::i64Bit, TMP3, ARMEmitter::Condition::CC_HI);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 2);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 1);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP3);
// z = min(|y|,|x|) / max(|y|,|x|); NaN guard catches 0/0, inf/inf, NaN inputs.
fcsel(ARMEmitter::ScalarRegSize::i64Bit, VTMP1, Accum, VTMP2, ARMEmitter::Condition::CC_HI);
fcsel(ARMEmitter::ScalarRegSize::i64Bit, Accum, VTMP2, Accum, ARMEmitter::Condition::CC_HI);
fdiv(VTMP2.D(), VTMP1.D(), Accum.D());
fcmp(VTMP2.D(), VTMP2.D());
(void)b(ARMEmitter::Condition::CC_VS, &Fallback);
// P(z^2) via 20-term Horner.
fmul(Accum.D(), VTMP2.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP2.D());
ldr(VTMP1.D(), &C19Label);
ldr(VTMP2.D(), &C18Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C17Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C16Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C15Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C14Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C13Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C12Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C11Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C10Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C9Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C8Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C7Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C6Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C5Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C4Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C3Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C2Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C1Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C0Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// atan_abs = z + z^3 * P(z^2).
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmul(VTMP1.D(), Accum.D(), VTMP1.D());
fmadd(VTMP1.D(), VTMP2.D(), VTMP1.D(), VTMP2.D());
// Quadrant adjustment driven by packed flags in TMP1.
(void)tbz(TMP1, 0, &NoSwap);
ldr(VTMP2.D(), &PiOver2Label);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
(void)Bind(&NoSwap);
(void)tbz(TMP1, 2, &PosX);
ldr(VTMP2.D(), &PiLabel);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
(void)Bind(&PosX);
(void)tbz(TMP1, 1, &NegY);
fneg(VTMP1.D(), VTMP1.D());
(void)Bind(&NegY);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 32);
ret();
// Fallback path: restore original inputs from stack stash and dispatch the ABI handler.
(void)Bind(&Fallback);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr(TMP2, ARMEmitter::Reg::rsp, 16);
ldr(TMP1, ARMEmitter::Reg::rsp, 24);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 32);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64ATAN].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64ATAN].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
Align(16);
(void)Bind(&C19Label);
dc64(0x3EF3'5885'1160'A528ULL);
(void)Bind(&C18Label);
dc64(0xBF2A'B24D'A7BE'7402ULL);
(void)Bind(&C17Label);
dc64(0x3F51'7739'E210'171AULL);
(void)Bind(&C16Label);
dc64(0xBF6D'0062'B42F'E3BFULL);
(void)Bind(&C15Label);
dc64(0x3F81'4E9D'C19A'4A4EULL);
(void)Bind(&C14Label);
dc64(0xBF90'0513'8172'2A59ULL);
(void)Bind(&C13Label);
dc64(0x3F98'6089'7B29'E5EFULL);
(void)Bind(&C12Label);
dc64(0xBFA0'0E6E'ECE7'DE80ULL);
(void)Bind(&C11Label);
dc64(0x3FA3'38E3'1EB2'FBBCULL);
(void)Bind(&C10Label);
dc64(0xBFA5'D301'40AE'5E99ULL);
(void)Bind(&C9Label);
dc64(0x3FA8'42DB'E9B0'D916ULL);
(void)Bind(&C8Label);
dc64(0xBFAA'EBFE'7B41'8581ULL);
(void)Bind(&C7Label);
dc64(0x3FAE'1D0F'9696'F63BULL);
(void)Bind(&C6Label);
dc64(0xBFB1'1100'EE08'4227ULL);
(void)Bind(&C5Label);
dc64(0x3FB3'B139'B6A8'8BA1ULL);
(void)Bind(&C4Label);
dc64(0xBFB7'45D1'60A7'E368ULL);
(void)Bind(&C3Label);
dc64(0x3FBC'71C7'1BC3'951CULL);
(void)Bind(&C2Label);
dc64(0xBFC2'4924'9247'8F88ULL);
(void)Bind(&C1Label);
dc64(0x3FC9'9999'9999'96C1ULL);
(void)Bind(&C0Label);
dc64(0xBFD5'5555'5555'5555ULL);
(void)Bind(&PiOver2Label);
dc64(0x3FF9'21FB'5444'2D18ULL);
(void)Bind(&PiLabel);
dc64(0x4009'21FB'5444'2D18ULL);
}
// JIT-inlined double-precision y * log2(x) for the F64 reduced precision x87 path.
// Input: VTMP1 = x, VTMP2 = y. Output: VTMP1 = y * log2(x).
// Constants in `C` are shared with EmitF64FYL2XP1 and emitted by EmitF64Log2Constants.
void Dispatcher::EmitF64FYL2X(F64Log2Constants& C) {
F64FYL2XHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Reject x <= 0, subnormal, inf and NaN before any FPR is clobbered,
// so VTMP1/VTMP2 still hold the original inputs at the fallback.
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
(void)tbnz(TMP1, 63, &Fallback);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP1, 52);
(void)cbz(ARMEmitter::Size::i64Bit, TMP2, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP2, 0x7FF);
(void)b(ARMEmitter::Condition::CC_EQ, &Fallback);
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP2.D());
// k = unbiased exponent; m bits = mantissa | (0x3FF << 52).
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1023);
ubfx(ARMEmitter::Size::i64Bit, TMP1, TMP1, 0, 52);
movz(ARMEmitter::Size::i64Bit, TMP4, 0x3FF0, 48);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4);
// Index = top 6 mantissa bits (bits 51..46 of m).
ubfx(ARMEmitter::Size::i64Bit, TMP4, TMP1, 46, 6);
// Set m as F64 in VTMP1.
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
// Load (recip, logc) from LUT[index].
(void)adr(TMP1, &C.Table);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
ldp<ARMEmitter::IndexType::OFFSET>(VTMP2.D(), Accum.D(), TMP1, 0);
// Stash logc bits in TMP4 so Accum can be reused for the polynomial.
fmov(ARMEmitter::Size::i64Bit, TMP4, Accum.D());
// r = recip * m - 1.
fmul(VTMP1.D(), VTMP2.D(), VTMP1.D());
ldr(VTMP2.D(), &C.One);
fsub(VTMP1.D(), VTMP1.D(), VTMP2.D());
// Horner: poly = a0 + r*(a1 + r*(a2 + ... + r*a7)).
ldr(Accum.D(), &C.A7);
ldr(VTMP2.D(), &C.A6);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A5);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A4);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A3);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A2);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A1);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A0);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
// log2(1+r) = r * Accum.
fmul(VTMP1.D(), VTMP1.D(), Accum.D());
// log2(x) = log2(1+r) + log2(center[i]) + k.
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP4);
fadd(VTMP1.D(), VTMP1.D(), VTMP2.D());
scvtf(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP2);
fadd(VTMP1.D(), VTMP1.D(), VTMP2.D());
// result = y * log2(x).
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path: VTMP1/VTMP2 still hold the original x/y.
(void)Bind(&Fallback);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2X].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2X].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
// JIT-inlined double-precision y * log2(1 + x) for the F64 reduced precision x87 path.
// Input: VTMP1 = x, VTMP2 = y. Output: VTMP1 = y * log2(1 + x).
// Computes v = 1 + x in F64 and runs the same LUT-based log2 as F64FYL2X.
// Loses 1-2 ulps of precision near x=0 (FYL2XP1's original purpose) but
// matches main's lowering and avoids the range-check cliff into a fallback.
void Dispatcher::EmitF64FYL2XP1(F64Log2Constants& C) {
F64FYL2XP1HandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// v = 1 + x in Accum so VTMP1/VTMP2 still hold the original x/y at the fallback.
ldr(Accum.D(), &C.One);
fadd(Accum.D(), VTMP1.D(), Accum.D());
// Reject v <= 0, subnormal, NaN, Inf via v's bits in TMP1.
fmov(ARMEmitter::Size::i64Bit, TMP1, Accum.D());
(void)tbnz(TMP1, 63, &Fallback);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP1, 52);
(void)cbz(ARMEmitter::Size::i64Bit, TMP2, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP2, 0x7FF);
(void)b(ARMEmitter::Condition::CC_EQ, &Fallback);
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP2.D());
// k = unbiased exponent; m bits = mantissa | (0x3FF << 52).
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1023);
ubfx(ARMEmitter::Size::i64Bit, TMP1, TMP1, 0, 52);
movz(ARMEmitter::Size::i64Bit, TMP4, 0x3FF0, 48);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4);
// Index = top 6 mantissa bits (bits 51..46 of m).
ubfx(ARMEmitter::Size::i64Bit, TMP4, TMP1, 46, 6);
// Set m as F64 in VTMP1.
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
// Load (recip, logc) from LUT[index].
(void)adr(TMP1, &C.Table);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
ldp<ARMEmitter::IndexType::OFFSET>(VTMP2.D(), Accum.D(), TMP1, 0);
// Stash logc bits in TMP4 so Accum can be reused for the polynomial.
fmov(ARMEmitter::Size::i64Bit, TMP4, Accum.D());
// r = recip * m - 1.
fmul(VTMP1.D(), VTMP2.D(), VTMP1.D());
ldr(VTMP2.D(), &C.One);
fsub(VTMP1.D(), VTMP1.D(), VTMP2.D());
// Horner: poly = a0 + r*(a1 + r*(a2 + ... + r*a7)).
ldr(Accum.D(), &C.A7);
ldr(VTMP2.D(), &C.A6);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A5);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A4);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A3);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A2);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A1);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &C.A0);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
// log2(1+r) = r * Accum.
fmul(VTMP1.D(), VTMP1.D(), Accum.D());
// log2(v) = log2(1+r) + log2(center[i]) + k.
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP4);
fadd(VTMP1.D(), VTMP1.D(), VTMP2.D());
scvtf(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP2);
fadd(VTMP1.D(), VTMP1.D(), VTMP2.D());
// result = y * log2(v).
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path: VTMP1/VTMP2 still hold the original x/y.
(void)Bind(&Fallback);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2XP1].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2XP1].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
// Shared constants pool for the LUT-based F64 log2 path used by FYL2X and FYL2XP1.
// Emitted once after both handlers; their forward `ldr`/`adr` references are
// patched here. Layout: One (1.0), 8 Horner coefficients A0..A7, then a 64-entry
// LUT of (1/center[i], log2(center[i])) pairs where center[i] = 1 + (i+0.5)/64.
void Dispatcher::EmitF64Log2Constants(F64Log2Constants& C) {
Align(16);
(void)Bind(&C.One);
dc64(0x3FF0000000000000ULL); // 1.0
(void)Bind(&C.A0);
dc64(0x3FF71547652B82FEULL); // log2(e) * 1/1
(void)Bind(&C.A1);
dc64(0xBFE71547652B82FEULL); // log2(e) * -1/2
(void)Bind(&C.A2);
dc64(0x3FDEC709DC3A03FDULL); // log2(e) * 1/3
(void)Bind(&C.A3);
dc64(0xBFD71547652B82FEULL); // log2(e) * -1/4
(void)Bind(&C.A4);
dc64(0x3FD2776C50EF9BFEULL); // log2(e) * 1/5
(void)Bind(&C.A5);
dc64(0xBFCEC709DC3A03FDULL); // log2(e) * -1/6
(void)Bind(&C.A6);
dc64(0x3FCA61762A7ADED9ULL); // log2(e) * 1/7
(void)Bind(&C.A7);
dc64(0xBFC71547652B82FEULL); // log2(e) * -1/8
Align(16);
(void)Bind(&C.Table);
dc64(0x3FEFC07F01FC07F0ULL);
dc64(0x3F86FE50B6EF0851ULL); // i= 0
dc64(0x3FEF44659E4A4271ULL);
dc64(0x3FA11CD1D5133413ULL); // i= 1
dc64(0x3FEECC07B301ECC0ULL);
dc64(0x3FAC4DFAB90AAB5FULL); // i= 2
dc64(0x3FEE573AC901E574ULL);
dc64(0x3FB3AA2FDD27F1C3ULL); // i= 3
dc64(0x3FEDE5D6E3F8868AULL);
dc64(0x3FB918A16E46335BULL); // i= 4
dc64(0x3FED77B654B82C34ULL);
dc64(0x3FBE72EC117FA5B2ULL); // i= 5
dc64(0x3FED0CB58F6EC074ULL);
dc64(0x3FC1DCD197552B7BULL); // i= 6
dc64(0x3FECA4B3055EE191ULL);
dc64(0x3FC476A9F983F74DULL); // i= 7
dc64(0x3FEC3F8F01C3F8F0ULL);
dc64(0x3FC70742D4EF027FULL); // i= 8
dc64(0x3FEBDD2B899406F7ULL);
dc64(0x3FC98EDD077E70DFULL); // i= 9
dc64(0x3FEB7D6C3DDA338BULL);
dc64(0x3FCC0DB6CDD94DEEULL); // i=10
dc64(0x3FEB2036406C80D9ULL);
dc64(0x3FCE840BE74E6A4DULL); // i=11
dc64(0x3FEAC5701AC5701BULL);
dc64(0x3FD0790ADBB03009ULL); // i=12
dc64(0x3FEA6D01A6D01A6DULL);
dc64(0x3FD1AC05B291F070ULL); // i=13
dc64(0x3FEA16D3F97A4B02ULL);
dc64(0x3FD2DB10FC4D9AAFULL); // i=14
dc64(0x3FE9C2D14EE4A102ULL);
dc64(0x3FD406463B1B0449ULL); // i=15
dc64(0x3FE970E4F80CB872ULL);
dc64(0x3FD52DBDFC4C96B3ULL); // i=16
dc64(0x3FE920FB49D0E229ULL);
dc64(0x3FD6518FE4677BA7ULL); // i=17
dc64(0x3FE8D3018D3018D3ULL);
dc64(0x3FD771D2BA7EFB3CULL); // i=18
dc64(0x3FE886E5F0ABB04AULL);
dc64(0x3FD88E9C72E0B226ULL); // i=19
dc64(0x3FE83C977AB2BEDDULL);
dc64(0x3FD9A802391E232FULL); // i=20
dc64(0x3FE7F405FD017F40ULL);
dc64(0x3FDABE18797F1F49ULL); // i=21
dc64(0x3FE7AD2208E0ECC3ULL);
dc64(0x3FDBD0F2E9E79031ULL); // i=22
dc64(0x3FE767DCE434A9B1ULL);
dc64(0x3FDCE0A4923A587DULL); // i=23
dc64(0x3FE724287F46DEBCULL);
dc64(0x3FDDED3FD442364CULL); // i=24
dc64(0x3FE6E1F76B4337C7ULL);
dc64(0x3FDEF6D67328E220ULL); // i=25
dc64(0x3FE6A13CD1537290ULL);
dc64(0x3FDFFD799A83FF9BULL); // i=26
dc64(0x3FE661EC6A5122F9ULL);
dc64(0x3FE0809CF27F703DULL); // i=27
dc64(0x3FE623FA77016240ULL);
dc64(0x3FE10113B153C8EAULL); // i=28
dc64(0x3FE5E75BB8D015E7ULL);
dc64(0x3FE18028CF72976AULL); // i=29
dc64(0x3FE5AC056B015AC0ULL);
dc64(0x3FE1FDE3D30E8126ULL); // i=30
dc64(0x3FE571ED3C506B3AULL);
dc64(0x3FE27A4C0585CBF8ULL); // i=31
dc64(0x3FE5390948F40FEBULL);
dc64(0x3FE2F56875EB3F26ULL); // i=32
dc64(0x3FE5015015015015ULL);
dc64(0x3FE36F3FFB6D9162ULL); // i=33
dc64(0x3FE4CAB88725AF6EULL);
dc64(0x3FE3E7D9379F7016ULL); // i=34
dc64(0x3FE49539E3B2D067ULL);
dc64(0x3FE45F3A98A20739ULL); // i=35
dc64(0x3FE460CBC7F5CF9AULL);
dc64(0x3FE4D56A5B33CEC4ULL); // i=36
dc64(0x3FE42D6625D51F87ULL);
dc64(0x3FE54A6E8CA5438EULL); // i=37
dc64(0x3FE3FB013FB013FBULL);
dc64(0x3FE5BE4D0CB51435ULL); // i=38
dc64(0x3FE3C995A47BABE7ULL);
dc64(0x3FE6310B8F553048ULL); // i=39
dc64(0x3FE3991C2C187F63ULL);
dc64(0x3FE6A2AF9E5A0F0AULL); // i=40
dc64(0x3FE3698DF3DE0748ULL);
dc64(0x3FE7133E9B156C7CULL); // i=41
dc64(0x3FE33AE45B57BCB2ULL);
dc64(0x3FE782BDBFDDA657ULL); // i=42
dc64(0x3FE30D190130D190ULL);
dc64(0x3FE7F1322182CF16ULL); // i=43
dc64(0x3FE2E025C04B8097ULL);
dc64(0x3FE85EA0B0B27B26ULL); // i=44
dc64(0x3FE2B404AD012B40ULL);
dc64(0x3FE8CB0E3B4B3BBEULL); // i=45
dc64(0x3FE288B01288B013ULL);
dc64(0x3FE9367F6DA0AB2FULL); // i=46
dc64(0x3FE25E22708092F1ULL);
dc64(0x3FE9A0F8D3B0E050ULL); // i=47
dc64(0x3FE23456789ABCDFULL);
dc64(0x3FEA0A7EDA4C112DULL); // i=48
dc64(0x3FE20B470C67C0D9ULL);
dc64(0x3FEA7315D02F20C8ULL); // i=49
dc64(0x3FE1E2EF3B3FB874ULL);
dc64(0x3FEADAC1E711C833ULL); // i=50
dc64(0x3FE1BB4A4046ED29ULL);
dc64(0x3FEB418734A9008CULL); // i=51
dc64(0x3FE19453808CA29CULL);
dc64(0x3FEBA769B39E4964ULL); // i=52
dc64(0x3FE16E0689427379ULL);
dc64(0x3FEC0C6D447C5DD3ULL); // i=53
dc64(0x3FE1485F0E0ACD3BULL);
dc64(0x3FEC7095AE91E1C7ULL); // i=54
dc64(0x3FE12358E75D3033ULL);
dc64(0x3FECD3E6A0CA8907ULL); // i=55
dc64(0x3FE0FEF010FEF011ULL);
dc64(0x3FED3663B27F31D5ULL); // i=56
dc64(0x3FE0DB20A88F4696ULL);
dc64(0x3FED9810643D6615ULL); // i=57
dc64(0x3FE0B7E6EC259DC8ULL);
dc64(0x3FEDF8F02086AF2CULL); // i=58
dc64(0x3FE0953F39010954ULL);
dc64(0x3FEE59063C8822CEULL); // i=59
dc64(0x3FE073260A47F7C6ULL);
dc64(0x3FEEB855F8CA88FBULL); // i=60
dc64(0x3FE05197F7D73404ULL);
dc64(0x3FEF16E281DB7630ULL); // i=61
dc64(0x3FE03091B51F5E1AULL);
dc64(0x3FEF74AEF0EFAFAEULL); // i=62
dc64(0x3FE0101010101010ULL);
dc64(0x3FEFD1BE4C7F2AF9ULL); // i=63
}
void Dispatcher::EmitF64FPREM() {
// JIT-inlined double-precision FPREM (C-library style truncated remainder).
// Input: VTMP1 = dividend (src1), VTMP2 = divisor (src2). Output: VTMP1.
F64FPREMHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Scratch = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel ReturnX;
ARMEmitter::ForwardLabel MaybeExact;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel NonZeroResult;
// Save q2.
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, TMP4, VTMP2.D());
fdiv(Scratch.D(), VTMP1.D(), VTMP2.D());
frintz(VTMP1.D(), Scratch.D());
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP1, 1);
(void)cbz(ARMEmitter::Size::i64Bit, TMP2, &ReturnX);
fmov(ARMEmitter::Size::i64Bit, TMP2, Scratch.D());
ubfx(ARMEmitter::Size::i64Bit, TMP2, TMP2, 52, 11);
cmp(ARMEmitter::Size::i64Bit, TMP2, 0x7FF);
(void)b(ARMEmitter::Condition::CC_EQ, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP2, 1076);
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
fsub(VTMP2.D(), Scratch.D(), VTMP1.D());
fabs(VTMP2.D(), VTMP2.D());
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 53);
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 52);
fmov(ARMEmitter::Size::i64Bit, Scratch.D(), TMP2);
fcmp(VTMP2.D(), Scratch.D());
// err <= 0.5 ULP could mean fdiv rounded across an integer boundary; defer to MaybeExact
// which distinguishes the (safe) exact-quotient case from the (unsafe) rounded case.
(void)b(ARMEmitter::Condition::CC_LS, &MaybeExact);
fmov(ARMEmitter::ScalarRegSize::i64Bit, Scratch, 1.0f);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
fsub(Scratch.D(), Scratch.D(), VTMP1.D());
fcmp(VTMP2.D(), Scratch.D());
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
fmov(ARMEmitter::Size::i64Bit, Scratch.D(), TMP4);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmsub(VTMP2.D(), VTMP1.D(), Scratch.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP2.D());
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &NonZeroResult);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP3, 63);
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 63);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP2);
(void)Bind(&NonZeroResult);
fmov(VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
(void)Bind(&ReturnX);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP3);
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// err <= 0.5 ULP. Two cases:
// err == 0: x/y rounded to an exact integer. Either truly exact (q is right) or fdiv
// rounded a near-integer onto an integer (q is off by one). Verify with
// fmsub(q, y, x): if exactly zero, x == q*y exactly => result is sign(x)*0.
// err > 0: q_fp is genuinely between integers but within rounding slack; fall back.
(void)Bind(&MaybeExact);
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP2.D());
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
fmov(ARMEmitter::Size::i64Bit, Scratch.D(), TMP4);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmsub(VTMP2.D(), VTMP1.D(), Scratch.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP2.D());
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &Fallback);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP3, 63);
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 63);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
(void)Bind(&Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP3);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FPREM].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FPREM].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
void Dispatcher::EmitF64FPREM1() {
// JIT-inlined double-precision FPREM1 (IEEE round-to-nearest remainder).
// Input: VTMP1 = dividend (src1), VTMP2 = divisor (src2). Output: VTMP1.
F64FPREM1HandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Scratch = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel ReturnX;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel NonZeroResult;
// Save q2.
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, TMP4, VTMP2.D());
fdiv(Scratch.D(), VTMP1.D(), VTMP2.D());
frintn(VTMP1.D(), Scratch.D());
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP1, 1);
(void)cbz(ARMEmitter::Size::i64Bit, TMP2, &ReturnX);
fmov(ARMEmitter::Size::i64Bit, TMP2, Scratch.D());
ubfx(ARMEmitter::Size::i64Bit, TMP2, TMP2, 52, 11);
cmp(ARMEmitter::Size::i64Bit, TMP2, 0x7FF);
(void)b(ARMEmitter::Condition::CC_EQ, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP2, 1076);
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
fsub(VTMP2.D(), Scratch.D(), VTMP1.D());
fabs(VTMP2.D(), VTMP2.D());
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 53);
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 52);
fmov(ARMEmitter::ScalarRegSize::i64Bit, Scratch, 0.5f);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
fsub(Scratch.D(), Scratch.D(), VTMP1.D());
fcmp(VTMP2.D(), Scratch.D());
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
fmov(ARMEmitter::Size::i64Bit, Scratch.D(), TMP4);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmsub(VTMP2.D(), VTMP1.D(), Scratch.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP2.D());
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP2, &NonZeroResult);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP3, 63);
lsl(ARMEmitter::Size::i64Bit, TMP2, TMP2, 63);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP2);
(void)Bind(&NonZeroResult);
fmov(VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
(void)Bind(&ReturnX);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP3);
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
(void)Bind(&Fallback);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP3);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FPREM1].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FPREM1].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
auto Address = GetCursorAddress<uint64_t>();
constexpr static auto FallbackPointerReg = TMP4;
@@ -1693,6 +2630,13 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
Ptrs.F64SinHandler = F64SinHandlerAddress;
Ptrs.F64CosHandler = F64CosHandlerAddress;
Ptrs.F64TanHandler = F64TanHandlerAddress;
Ptrs.F64F2XM1Handler = F64F2XM1HandlerAddress;
Ptrs.F64ScaleHandler = F64ScaleHandlerAddress;
Ptrs.F64AtanHandler = F64AtanHandlerAddress;
Ptrs.F64FYL2XHandler = F64FYL2XHandlerAddress;
Ptrs.F64FYL2XP1Handler = F64FYL2XP1HandlerAddress;
Ptrs.F64FPREMHandler = F64FPREMHandlerAddress;
Ptrs.F64FPREM1Handler = F64FPREM1HandlerAddress;
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Ptrs.FallbackHandlerPointers, &ABIPointers[0]);
@@ -99,10 +99,17 @@ private:
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
// F64 trig shared handlers
// F64 reduced-precision shared handlers
uint64_t F64SinHandlerAddress {};
uint64_t F64CosHandlerAddress {};
uint64_t F64TanHandlerAddress {};
uint64_t F64F2XM1HandlerAddress {};
uint64_t F64ScaleHandlerAddress {};
uint64_t F64AtanHandlerAddress {};
uint64_t F64FYL2XHandlerAddress {};
uint64_t F64FYL2XP1HandlerAddress {};
uint64_t F64FPREMHandlerAddress {};
uint64_t F64FPREM1HandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
@@ -114,9 +121,25 @@ private:
void EmitF32ToExtF80();
void EmitF64ToExtF80();
// Shared label set for the LUT-based F64 log2 path used by both FYL2X and
// FYL2XP1. The pool is emitted once via EmitF64Log2Constants.
struct F64Log2Constants {
ARMEmitter::ForwardLabel One;
ARMEmitter::ForwardLabel A0, A1, A2, A3, A4, A5, A6, A7;
ARMEmitter::ForwardLabel Table;
};
void EmitF64Sin();
void EmitF64Cos();
void EmitF64Tan();
void EmitF64F2XM1();
void EmitF64Scale();
void EmitF64Atan();
void EmitF64FYL2X(F64Log2Constants& C);
void EmitF64FYL2XP1(F64Log2Constants& C);
void EmitF64Log2Constants(F64Log2Constants& C);
void EmitF64FPREM();
void EmitF64FPREM1();
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
+61 -45
View File
@@ -124,9 +124,9 @@ uint8_t Decoder::ReadByte() {
}
std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream + InstructionSize + Offset);
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize + Offset);
if (CheckRangeExecutable(ByteAddress, 1)) {
return InstStream[InstructionSize + Offset];
return InstStream.AdjustedInstStream[InstructionSize + Offset];
} else {
return std::nullopt;
}
@@ -136,9 +136,9 @@ std::pair<uint64_t, bool> Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
uint64_t Address = reinterpret_cast<uint64_t>(InstStream + InstructionSize);
uint64_t Address = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize);
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream[InstructionSize], Size);
std::memcpy(&Res, &InstStream.AdjustedInstStream[InstructionSize], Size);
} else {
HitNonExecutableRange = true;
// See PeekByte, this specific case may cause some executable memory to read as 0 but it doesn't matter as the entire instruction will be rolled back anyway.
@@ -342,7 +342,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
}
bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
Decoder::DecodedBlockStatus Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
// TODO: Move this in to `NormalOpHeader`, Dispatch tables have a bug currently where some subtables don't inherit flags correctly.
@@ -354,11 +354,16 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SUPPORTS_LOCK) && (DecodeInst->Flags & DecodeFlags::FLAG_LOCK)) {
// Instruction has lock prefix but doesn't support lock.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P), "Group Ops "
@@ -390,15 +395,15 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
if (Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
const bool UseVEXL = Options.L && !(Info->Flags & InstFlags::FLAGS_VEX_L_IGNORE);
@@ -507,7 +512,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -576,7 +581,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const auto VEXOperand = Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_SRC_MASK;
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_NO_OPERAND && Options.vvvv) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) {
@@ -594,11 +599,11 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM) {
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SF_MOD_DST) {
if (!ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
} else {
if (!ModRMOperand(DecodeInst->Dest, DecodeInst->Src[CurrentSrc], HasXMMDst, HasXMMSrc, HasMMDst, HasMMSrc, Is8BitDest, Is8BitSrc)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
++CurrentSrc;
@@ -660,22 +665,27 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
Bytes = 0;
}
if ((DecodeInst->Flags & DecodeFlags::FLAG_LOCK) && DecodeInst->Dest.IsGPR()) {
// Instruction has lock prefix, but the destination isn't memory, this is invalid.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
return DecodedBlockStatus::SUCCESS;
}
bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
Decoder::DecodedBlockStatus Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
DecodeInst->OPRaw = DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
@@ -732,7 +742,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
};
uint8_t Field = RegToField[ModRM.reg];
if (Field == 255) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LocalOp = (Field << 3) | ModRM.rm;
@@ -751,7 +761,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
if (!VEXTable) {
// AVX not enabled.
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint16_t map_select = 1;
@@ -761,7 +771,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
if ((Byte1 & 0b10000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
@@ -772,7 +782,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte1 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.L = (Byte1 & 0b100) != 0;
@@ -783,14 +793,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte2 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
@@ -801,7 +811,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -831,14 +841,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
FEXCORE_TELEMETRY_SET(TYPE_USES_EVEX_OPS, 1);
// EVEX unsupported
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_MSG_A_FMT("Invalid instruction decoding type");
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
Decoder::DecodedBlockStatus Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
LastEscapePrefix = 0;
Instruction.fill(0);
@@ -849,7 +859,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
for (;;) {
if (InstructionSize >= MAX_INST_SIZE) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint8_t Op = ReadByte();
switch (Op) {
@@ -1035,10 +1045,10 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
}
if (DecodeInst->Dest.IsGPR()) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
return true;
return DecodedBlockStatus::SUCCESS;
}
void Decoder::DecodeREXIfValid(int8_t ExpectedOffset) {
@@ -1076,16 +1086,16 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
HitBadRelocation = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
auto ErrorDuringDecoding = DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
if (ErrorDuringDecoding != DecodedBlockStatus::SUCCESS || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
auto Result = ErrorDuringDecoding != DecodedBlockStatus::SUCCESS ? ErrorDuringDecoding :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
DecodeInst->InstSize = 0;
return Result;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
@@ -1331,7 +1341,7 @@ void Decoder::AddBranchTarget(uint64_t Target) {
}
}
const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
const Decoder::DecodeStream Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
@@ -1342,10 +1352,16 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
// Offset 0x400: vtime
// Offset 0x800: vgetcpu
uint64_t Offset = RIP - VSyscall_Base;
return VSyscallData + Offset;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = VSyscallData + Offset,
};
}
return _InstStream - EntryPoint + RIP;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = _InstStream - EntryPoint + RIP,
};
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
@@ -1373,7 +1389,6 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
@@ -1507,10 +1522,11 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
"PartialDecode",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
BlockIt->BlockStatus == DecodedBlockStatus::UNIMPLEMENTED_INST ? "Unimplemented" :
"PartialDecode",
OpAddress);
}
break;
+27 -5
View File
@@ -32,6 +32,7 @@ public:
NOEXEC_INST,
PARTIAL_DECODE_INST,
BAD_RELOCATION,
UNIMPLEMENTED_INST,
};
// New Frontend decoding
@@ -55,6 +56,7 @@ public:
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
@@ -91,7 +93,7 @@ private:
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
@@ -110,8 +112,8 @@ private:
InstructionSize += Size;
}
bool NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
DecodedBlockStatus NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
DecodedBlockStatus NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
void DecodeREXIfValid(int8_t ExpectedOffset = -1);
@@ -126,7 +128,27 @@ private:
bool HitNonExecutableRange {};
bool HitBadRelocation {};
const uint8_t* InstStream {};
struct DecodeStream {
// Original instruction stream RIP location.
const uint8_t* InstStream;
// Adjusted location for FEX actually decodes from.
const uint8_t* AdjustedInstStream;
DecodeStream& operator-=(size_t offset) noexcept {
InstStream -= offset;
AdjustedInstStream -= offset;
return *this;
}
DecodeStream& operator+=(size_t offset) noexcept {
InstStream += offset;
AdjustedInstStream += offset;
return *this;
}
};
DecodeStream InstStream;
IR::OpSize GetGPROpSize() const {
return BlockInfo.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
@@ -168,6 +190,6 @@ private:
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_TABLE_SIZE>* VEXTable {};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_GROUP_TABLE_SIZE>* VEXTableGroup {};
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
const DecodeStream AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
};
} // namespace FEXCore::Frontend
@@ -302,6 +302,16 @@ struct OpHandlers<IR::OP_F80FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame, true};
const X80SoftFloat One {&State.State, 1.0};
return X80SoftFloat::FYL2X(&State.State, X80SoftFloat::FADD(&State.State, Src1, One), Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -417,6 +427,14 @@ struct OpHandlers<IR::OP_F64FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(1.0 + src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -72,6 +72,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80DIV>::handle)};
Info[Core::OPINDEX_F80FYL2X] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2X>::handle)};
Info[Core::OPINDEX_F80FYL2XP1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2XP1>::handle)};
Info[Core::OPINDEX_F80ATAN] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80ATAN>::handle)};
Info[Core::OPINDEX_F80FPREM1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
@@ -97,6 +99,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle)};
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle)};
Info[Core::OPINDEX_F64FYL2XP1] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2XP1>::handle)};
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle)};
@@ -254,6 +258,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
COMMON_BINARY_X87_OP(MUL)
COMMON_BINARY_X87_OP(DIV)
COMMON_BINARY_X87_OP(FYL2X)
COMMON_BINARY_X87_OP(FYL2XP1)
COMMON_BINARY_X87_OP(ATAN)
COMMON_BINARY_X87_OP(FPREM1)
COMMON_BINARY_X87_OP(FPREM)
@@ -268,6 +273,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
// Double Precision Binary
COMMON_BINARY_F64_OP(FYL2X)
COMMON_BINARY_F64_OP(FYL2XP1)
COMMON_BINARY_F64_OP(ATAN)
COMMON_BINARY_F64_OP(FPREM1)
COMMON_BINARY_F64_OP(FPREM)
+3 -3
View File
@@ -274,7 +274,7 @@ DEF_OP(CmpPairZ) {
// Restore NzCV
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 0, 0xb /* NzCV */);
rmif(TMP1, 28, 0xb /* NzCV */);
} else {
cset(ARMEmitter::Size::i32Bit, TMP2, ARMEmitter::Condition::CC_EQ);
bfi(ARMEmitter::Size::i32Bit, TMP1, TMP2, 30 /* lsb: Z */, 1);
@@ -523,7 +523,7 @@ DEF_OP(AndWithFlags) {
}
DEF_OP(AndShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
auto Op = IROp->C<IR::IROp_AndShift>();
and_(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1), GetReg(Op->Src2), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
@@ -721,7 +721,7 @@ DEF_OP(Extr) {
}
DEF_OP(PDep) {
auto Op = IROp->C<IR::IROp_PExt>();
auto Op = IROp->C<IR::IROp_PDep>();
const auto EmitSize = ConvertSize48(IROp);
const auto Dest = GetReg(Node);
@@ -55,6 +55,29 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if constexpr (Context::BLOCK_DEBUGGING) {
// Skip block linking when BLOCK_DEBUGGING as it adds overhead and is unncessary.
// This is a debug only feature and doesn't need caching help.
bool IsInlineRIP = IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP);
ARMEmitter::ForwardLabel l_ExitLink;
if (IsInlineRIP) {
ldr(TMP1, &l_ExitLink);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
} else {
auto RipReg = GetReg(Op->NewRIP);
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
}
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
br(TMP2);
if (IsInlineRIP) {
BindOrRestart(&l_ExitLink);
dc64(NewRIP);
}
return;
}
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef ARCHITECTURE_arm64ec
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
+9 -2
View File
@@ -776,8 +776,15 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
constexpr size_t InterruptPageOffset =
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState);
if constexpr (InterruptPageOffset <= 32760) {
str(ARMEmitter::XReg::zr, STATE, InterruptPageOffset);
} else {
// Need to use vector 128-bit store for this range.
// Doesn't matter which register we use to store.
str(ARMEmitter::QReg::q0, STATE, InterruptPageOffset);
}
}
#ifdef ARCHITECTURE_arm64ec
@@ -568,7 +568,7 @@ DEF_OP(LoadDF) {
DEF_OP(ContextClear) {
auto Op = IROp->C<IR::IROp_ContextClear>();
if (CTX->HostFeatures.SupportsCLZERO) {
if (CTX->HostFeatures.PreferZVAForVZero) {
// We can use CLZero directly when hardware supports it.
// Provides a fairly generous speed-up on Ampere1A hardware.
// TODO: When FEAT_MOPS hardware ships, test memset using MOPS.
@@ -2400,11 +2400,12 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CIVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
@@ -2428,11 +2429,12 @@ DEF_OP(CacheLineClean) {
auto MemReg = GetReg(Op->Addr);
// Clean dcache only
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
@@ -73,8 +73,8 @@ DEF_OP(Break) {
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Constant);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case Core::FAULT_SIGILL:
+201 -52
View File
@@ -1036,24 +1036,28 @@ DEF_OP(VMov) {
const auto Dst = GetVReg(Node);
const auto Source = GetVReg(Op->Source);
const auto Sub64BitHandler = [&](ARMEmitter::SubRegSize InsertSize) {
if (Dst != Source) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
ins(InsertSize, Dst, 0, Source, 0);
} else {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(InsertSize, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
}
};
switch (OpSize) {
case IR::OpSize::i8Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i8Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i8Bit);
break;
}
case IR::OpSize::i16Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i16Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i16Bit);
break;
}
case IR::OpSize::i32Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i32Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i32Bit);
break;
}
case IR::OpSize::i64Bit: {
@@ -1095,16 +1099,21 @@ DEF_OP(VAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE ADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE ADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of ADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
addp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
addp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1298,16 +1307,21 @@ DEF_OP(VFAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE FADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE FADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of FADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
faddp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
faddp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1434,8 +1448,8 @@ DEF_OP(VFMin) {
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on false.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector1.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
@@ -1466,7 +1480,8 @@ DEF_OP(VFMax) {
const auto Mask = PRED_TMP_32B;
const auto ComparePred = ARMEmitter::PReg::p0;
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector2.Z(), Vector1.Z());
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector1.Z(), Vector2.Z());
not_(ComparePred, Mask.Zeroing(), ComparePred);
if (Dst == Vector1) {
// Trivial case where Vector1 is also the destination.
@@ -1488,17 +1503,17 @@ DEF_OP(VFMax) {
if (Dst == Vector1) {
// Destination is already Vector1, need to insert Vector2 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
}
}
}
@@ -1525,9 +1540,14 @@ DEF_OP(VFRecp) {
return;
}
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
if (Dst != Vector) {
fmov(SubRegSize.Vector, Dst.Z(), 1.0);
fdiv(SubRegSize.Vector, Dst.Z(), Pred, Dst.Z(), Vector.Z());
} else {
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
}
} else {
if (IsScalar) {
if (ElementSize == IR::OpSize::i32Bit && HostSupportsRPRES) {
@@ -1779,10 +1799,14 @@ DEF_OP(VUMin) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1828,10 +1852,14 @@ DEF_OP(VSMin) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1877,10 +1905,14 @@ DEF_OP(VUMax) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1926,10 +1958,14 @@ DEF_OP(VSMax) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -2979,9 +3015,14 @@ DEF_OP(VInsElement) {
auto Reg = GetVReg(Op->DestVector);
if (HostSupportsSVE256 && Is256Bit) {
// Broadcast our source value across a temporary,
// then combine with the destination.
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
// Broadcast our source value across a temporary, then combine
// with the destination.
//
// We don't need to perform the dup if we're just merging a 128-bit vector into
// into an equivalent position since we have a predicate set up already.
if (!(ElementSize == IR::OpSize::i128Bit && SrcIdx == DestIdx)) {
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
}
// We don't need to move the data unnecessarily if
// DestVector just so happens to also be the IR op
@@ -2994,10 +3035,12 @@ DEF_OP(VInsElement) {
if (ElementSize == IR::OpSize::i128Bit) {
if (DestIdx == 0) {
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), VTMP2.Z());
const auto Source = SrcIdx == 0 ? SrcVector : VTMP2;
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), Source.Z());
} else {
const auto Source = SrcIdx == 1 ? SrcVector : VTMP2;
not_(Predicate, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), VTMP2.Z());
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), Source.Z());
}
} else {
const auto UpperBound = 16 >> FEXCore::ilog2(IR::OpSizeToSize(ElementSize));
@@ -4434,7 +4477,7 @@ DEF_OP(VFNMLA) {
// - SVE - FMLS
// - ASIMD - FMLS
// - Scalar - FMSUB
const auto Op = IROp->C<IR::IROp_VFMLA>();
const auto Op = IROp->C<IR::IROp_VFNMLA>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4502,7 +4545,7 @@ DEF_OP(VFNMLS) {
// - ASIMD - FMLS (With Negated addend)
// - Scalar - FNMADD
const auto Op = IROp->C<IR::IROp_VFMLS>();
const auto Op = IROp->C<IR::IROp_VFNMLS>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4611,6 +4654,36 @@ DEF_OP(VFCopySign) {
}
}
DEF_OP(F64FPREM) {
const auto Op = IROp->C<IR::IROp_F64FPREM>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREMHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64FPREM1) {
const auto Op = IROp->C<IR::IROp_F64FPREM1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SIN) {
const auto Op = IROp->C<IR::IROp_F64SIN>();
const auto Src = GetVReg(Op->Src);
@@ -4650,5 +4723,81 @@ DEF_OP(F64TAN) {
fmov(Dst.D(), VTMP1.D());
}
// Src1=y(ST1), Src2=x(ST0). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64ATAN) {
const auto Op = IROp->C<IR::IROp_F64ATAN>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64AtanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2X) {
const auto Op = IROp->C<IR::IROp_F64FYL2X>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2XP1) {
const auto Op = IROp->C<IR::IROp_F64FYL2XP1>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XP1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SCALE) {
const auto Op = IROp->C<IR::IROp_F64SCALE>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64ScaleHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64F2XM1) {
const auto Op = IROp->C<IR::IROp_F64F2XM1>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64F2XM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
} // namespace FEXCore::CPU
@@ -87,8 +87,11 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
}
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// TODO: Preserve code cache entries?
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
// TODO: Rename this member to avoid confusion with code caching
CachedCodePages.clear();
}
@@ -502,7 +502,7 @@ void OpDispatchBuilder::LEAVEOp(OpcodeArgs) {
auto NewGPR = Pop(OperandSize, SP);
// Store the new stack pointer
StoreGPRRegister(X86State::REG_RSP, SP, OperandSize);
StoreGPRRegister(X86State::REG_RSP, SP, GPRSize);
// Store what we loaded to RBP
StoreGPRRegister(X86State::REG_RBP, NewGPR, OperandSize);
@@ -2427,9 +2427,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
auto BitSelect = (Size == (LshrSize * 8)) ? Src : Src.And(Mask);
auto LshrOpSize = IR::SizeToOpSize(LshrSize);
// OF/SF/AF/PF undefined. ZF must be preserved. We choose to preserve OF/SF
// too since we just use an rmif to insert into CF directly. We could
// optimize perhaps.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
//
// Set CF before the action to save a move, except for complements where we
// can reuse the invert.
@@ -2547,7 +2548,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
Value = _Lshr(std::max(OpSize::i32Bit, GetOpSize(Value)), Value, BitSelect.Ref());
}
// OF/SF/ZF/AF/PF undefined.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
SetCFDirect(Value, 0, true);
}
}
@@ -3449,30 +3453,38 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
}
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("LODSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
StoreResultGPR(Op, Src);
// Offset the pointer
Ref TailDest_RSI = LoadGPRRegister(X86State::REG_RSI);
StoreGPRRegister(X86State::REG_RSI, OffsetByDir(TailDest_RSI, IR::OpSizeToSize(Size)));
Ref TailDest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t PtrDir) {
// XXX: Theoretically LODS could be optimized to
// RSI += {-}(RCX * Size)
// RAX = [RSI - Size]
@@ -3500,7 +3512,8 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
@@ -3516,8 +3529,13 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RSI = Add(OpSize::i64Bit, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
TailDest_RSI = Add(AddrSize, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
// Jump back to the start, we have more work to do
Jump(LoopStart);
@@ -4216,6 +4234,94 @@ void OpDispatchBuilder::UpdatePrefixFromSegment(Ref Segment, uint32_t SegmentReg
}
}
uint64_t OpDispatchBuilder::CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad) {
if constexpr (!Context::BLOCK_DEBUGGING) {
LOGMAN_MSG_A_FMT("Tried to calculate address without block debugging enabled!");
FEX_UNREACHABLE;
}
const auto GPRSize = GetGPROpSize();
const auto GPRMask = GPRSize == OpSize::i64Bit ? ~0ULL : ~0U;
// This makes the assumption that InternalThreadState is synchronized at the point of call!
uint64_t Ptr {};
if (Operand.IsLiteral()) {
Ptr = Operand.Literal();
if (Operand.Data.Literal.Size != 8 && IsLoad) {
// zero extend
uint64_t width = Operand.Data.Literal.Size * 8;
Ptr &= ((1ULL << width) - 1);
}
} else if (Operand.IsGPR()) {
// Not a memory source.
return ~0ULL;
} else if (Operand.IsGPRDirect()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
} else if (Operand.IsGPRIndirect() || Operand.IsGPRIndirectRelocation()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
Ptr += static_cast<int32_t>(Operand.Data.GPRIndirect.Displacement);
} else if (Operand.IsRIPRelative() || Operand.IsRIPRelativeRelocation()) {
// 64-bit is RIP relative, while 32-bit is absolute.
if (Is64BitMode) {
Ptr = Op->PC + Op->InstSize + static_cast<int32_t>(Operand.Data.RIPLiteral.Value) - Entry;
} else {
Ptr = Operand.Data.RIPLiteral.Value;
}
} else if (Operand.IsSIB() || Operand.IsSIBRelocation()) {
const bool IsVSIB = IsLoad && ((Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0);
if (IsVSIB) {
// TODO: Unhandled.
return ~0ULL;
}
if (Operand.Data.SIB.Base != FEXCore::X86State::REG_INVALID) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Base] & GPRMask;
}
if (Operand.Data.SIB.Index != FEXCore::X86State::REG_INVALID) {
Ptr += (Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Index] * Operand.Data.SIB.Scale) & GPRMask;
}
Ptr += static_cast<int32_t>(Operand.Data.SIB.Offset);
}
auto AppendSegment = [&](uint64_t Ptr, uint32_t Flags, uint32_t DefaultPrefix = FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX,
bool Override = false) -> uint64_t {
uint32_t Prefix = Flags & FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS;
if (Is64BitMode) {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.fs_cached;
} else if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.gs_cached;
}
// If there was any other segment in 64bit then it is ignored
} else {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX || Override) {
// If there was no prefix then use the default one if available
// Or the argument only uses a specific prefix (with override set)
Prefix = DefaultPrefix;
}
// With the segment register optimization we store the GDT bases directly in the segment register to remove indexed loads
switch (Prefix) {
[[likely]] case FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX:
return Ptr;
case FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX: return Ptr + Thread->CurrentFrame->State.es_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX: return Ptr + Thread->CurrentFrame->State.cs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX: return Ptr + Thread->CurrentFrame->State.ss_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX: return Ptr + Thread->CurrentFrame->State.ds_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX: return Ptr + Thread->CurrentFrame->State.fs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX: return Ptr + Thread->CurrentFrame->State.gs_cached;
default: FEX_UNREACHABLE;
}
}
return Ptr;
};
return AppendSegment(Ptr, Op->Flags);
};
AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand,
MemoryAccessType AccessType, bool IsLoad) {
const auto GPRSize = GetGPROpSize();
@@ -4306,6 +4412,7 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
auto [Align, LoadData, ForceLoad, AccessType, AllowUpperGarbage] = Options;
AddressMode A = DecodeAddress(Op, Operand, AccessType, true /* IsLoad */);
Ref Result {};
if (Operand.IsGPR()) {
const auto gpr = Operand.Data.GPR.GPR;
const auto highIndex = Operand.Data.GPR.HighBits ? 1 : 0;
@@ -4335,22 +4442,35 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
}
}
if ((IsOperandMem(Operand, true) && LoadData) || ForceLoad) {
const bool ShouldLoad = (IsOperandMem(Operand, true) && LoadData) || ForceLoad;
if (ShouldLoad) {
if (OpSize == OpSize::f80Bit) {
Ref MemSrc = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
return _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
if (CTX->HostFeatures.SupportsSVE()) {
Result = _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
} else {
// For X87 extended doubles, Split the load.
auto Res = _LoadMem(Class, OpSize::i64Bit, MemSrc, Align == OpSize::iInvalid ? OpSize : Align);
return _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
Result = _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
}
} else {
Result = _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
}
} else {
Result = LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (ShouldLoad && CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, true);
if (CTX->BlockDebuggerTracker.ContainsReadWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit read watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
return _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
} else {
return LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
return Result;
}
Ref OpDispatchBuilder::LoadGPRRegister(uint32_t GPR, IR::OpSize Size, uint8_t Offset, bool AllowUpperGarbage) {
@@ -4465,7 +4585,7 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
if (OpSize == OpSize::f80Bit) {
Ref MemStoreDst = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
if (CTX->HostFeatures.SupportsSVE()) {
_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, Src, MemStoreDst);
} else {
// For X87 extended doubles, split before storing
@@ -4476,6 +4596,16 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
} else {
_StoreMemAutoTSO(Class, OpSize, A, Src, Align == OpSize::iInvalid ? OpSize : Align);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, false);
if (CTX->BlockDebuggerTracker.ContainsWriteWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit write watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
}
}
void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src,
@@ -4487,9 +4617,10 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
StoreResult(Class, Op, Op->Dest, Src, Align, AccessType);
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
, CTX {ctx} {
, CTX {ctx}
, Thread {Thread} {
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
SaveAVXStateFunc = &OpDispatchBuilder::SaveAVXState;
RestoreAVXStateFunc = &OpDispatchBuilder::RestoreAVXState;
@@ -4647,35 +4778,36 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
case 0xCD: { // INT imm8
uint8_t Literal = Op->Src[0].Literal();
#ifndef _WIN32
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Linux) {
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
} else if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Wow64 ||
CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
#else
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
#endif
#ifdef ARCHITECTURE_arm64ec
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
}
}
}
#endif
Reason.ErrorRegister = Literal << 3 | (0b010);
Reason.Signal = Core::FAULT_SIGSEGV;
@@ -4925,6 +5057,7 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
return;
}
const auto GPRSize = GetGPROpSize();
const auto SrcSize = OpSizeFromSrc(Op);
// Destination GPR size is always 4 or 8 bytes depending on widening
const auto DstSize = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REX_WIDENING ? OpSize::i64Bit : OpSize::i32Bit;
@@ -4933,16 +5066,15 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
// Incoming memory is 8, 16, 32, or 64
Ref Src {};
if (Op->Src[0].IsGPR()) {
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], GPRSize, Op->Flags);
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
} else {
Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
}
auto Result = _CRC32(Dest, Src, OpSizeFromSrc(Op));
auto Result = _CRC32(Dest, Src, SrcSize);
StoreResultGPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
template<bool Reseed>
void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
void OpDispatchBuilder::RDRANDOp(OpcodeArgs, bool Reseed) {
if (!CTX->HostFeatures.SupportsRAND) {
UnimplementedOp(Op);
return;
@@ -4967,9 +5099,6 @@ void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
}
}
template void OpDispatchBuilder::RDRANDOp<true>(OpcodeArgs);
template void OpDispatchBuilder::RDRANDOp<false>(OpcodeArgs);
void OpDispatchBuilder::BreakOp(OpcodeArgs, FEXCore::IR::BreakDefinition BreakDefinition) {
const auto GPRSize = GetGPROpSize();
+68 -130
View File
@@ -303,7 +303,7 @@ public:
StartNewBlock();
}
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
// Should only be called at the start of IR Emission.
void ResetWorkingList();
@@ -360,7 +360,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorAlignedOp(OpcodeArgs);
void MOVVectorUnalignedOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs, bool IsAVX);
void ALUOp(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, unsigned SrcIdx);
void LSLOp(OpcodeArgs);
void INTOp(OpcodeArgs);
@@ -470,8 +470,7 @@ public:
void AAMOp(OpcodeArgs);
void AADOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
void RDRANDOp(OpcodeArgs, bool Reseed);
enum class Segment {
FS,
@@ -501,8 +500,7 @@ public:
void VectorALUROp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VectorUnaryOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void RSqrt3DNowOp(OpcodeArgs, bool Duplicate);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorUnaryDuplicateOp(OpcodeArgs);
void VectorUnaryDuplicateOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void MOVQOp(OpcodeArgs, VectorOpType VectorType);
void MOVQMMXOp(OpcodeArgs);
@@ -524,36 +522,24 @@ public:
void PSLLDQ(OpcodeArgs);
void PSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void MOVDDUPOp(OpcodeArgs);
template<IR::OpSize DstElementSize>
void CVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void CVTFPR_To_GPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool Widen>
void Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void Scalar_CVT_Float_To_Float(OpcodeArgs);
void CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void Vector_CVT_Int_To_Float(OpcodeArgs, IR::OpSize SrcElementSize, bool Widen, bool IsAVX);
void Vector_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize, bool IsAVX);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void Vector_CVT_Float_To_Int(OpcodeArgs);
void Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode, bool IsAVX);
void MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs);
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void MASKMOVOp(OpcodeArgs);
void MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType);
void TZCNT(OpcodeArgs);
void LZCNT(OpcodeArgs);
template<IR::OpSize ElementSize>
void VFCMPOp(OpcodeArgs);
void VFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void SHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PINSROp(OpcodeArgs);
void PINSROp(OpcodeArgs, IR::OpSize ElementSize);
void InsertPSOp(OpcodeArgs);
void PExtrOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PSIGN(OpcodeArgs);
template<IR::OpSize ElementSize>
void VPSIGN(OpcodeArgs);
void PSIGN(OpcodeArgs, IR::OpSize ElementSize);
void VPSIGN(OpcodeArgs, IR::OpSize ElementSize);
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
@@ -576,53 +562,32 @@ public:
// AVX Ops
void AVXVectorXOROp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXVectorRound(OpcodeArgs);
void AVXVectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXScalar_CVT_Float_To_Float(OpcodeArgs);
void VectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarUnaryInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs);
void VectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void InsertMMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize>
void InsertCVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize DstElementSize>
void AVXInsertCVTGPR_To_FPR(OpcodeArgs);
void InsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
void AVXInsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void InsertScalar_CVT_Float_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs);
void InsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
RoundMode TranslateRoundType(uint8_t Mode);
template<IR::OpSize ElementSize>
void InsertScalarRound(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarRound(OpcodeArgs);
void InsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void InsertScalarFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarFCMPOp(OpcodeArgs);
void InsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize>
void AVXCVTGPR_To_FPR(OpcodeArgs);
void AVXVFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVXVFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void VADDSUBPOp(OpcodeArgs);
void VADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VAESDecOp(OpcodeArgs);
void VAESDecLastOp(OpcodeArgs);
@@ -631,34 +596,31 @@ public:
void VANDNOp(OpcodeArgs);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, Ref ZeroRegister, uint64_t Selector);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint64_t Selector);
void VBLENDPDOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPBLENDWOp(OpcodeArgs);
void VBROADCASTOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VDPPOp(OpcodeArgs);
void VDPPOp(OpcodeArgs, IR::OpSize ElementSize);
void VEXTRACT128Op(OpcodeArgs);
template<IROps IROp, IR::OpSize ElementSize>
void VHADDPOp(OpcodeArgs);
void VHADDPOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VHSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VINSERTOp(OpcodeArgs);
void VINSERTPSOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool IsStore>
void VMASKMOVOp(OpcodeArgs);
void VMASKMOVOp(OpcodeArgs, IR::OpSize ElementSize, bool IsStore);
void VMOVHPOp(OpcodeArgs);
void VMOVLPOp(OpcodeArgs);
void VMOVDDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSLDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSDOp(OpcodeArgs);
void VMOVSSOp(OpcodeArgs);
@@ -669,15 +631,14 @@ public:
void VMPSADBWOp(OpcodeArgs);
void VPACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPALIGNROp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs);
void VPCMPESTRMOp(OpcodeArgs);
void VPCMPISTRIOp(OpcodeArgs);
void VPCMPISTRMOp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPESTRMOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRMOp(OpcodeArgs, bool IsAVX);
void VCVTPH2PSOp(OpcodeArgs);
void VCVTPS2PHOp(OpcodeArgs);
@@ -690,36 +651,28 @@ public:
void VPERMILImmOp(OpcodeArgs, IR::OpSize ElementSize);
Ref VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize, Ref Src, Ref Indices);
template<IR::OpSize ElementSize>
void VPERMILRegOp(OpcodeArgs);
void VPERMILRegOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHADDSWOp(OpcodeArgs);
void VPHSUBOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHSUBSWOp(OpcodeArgs);
void VPINSRBOp(OpcodeArgs);
void VPINSRBWOp(OpcodeArgs, IR::OpSize ElementSize);
void VPINSRDQOp(OpcodeArgs);
void VPINSRWOp(OpcodeArgs);
void VPMADDUBSWOp(OpcodeArgs);
void VPMADDWDOp(OpcodeArgs);
template<bool IsStore>
void VPMASKMOVOp(OpcodeArgs);
void VPMASKMOVOp(OpcodeArgs, bool IsStore);
void VPMULHRSWOp(OpcodeArgs);
template<bool Signed>
void VPMULHWOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool Signed>
void VPMULLOp(OpcodeArgs);
void VPMULHWOp(OpcodeArgs, bool Signed);
void VPMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
void VPSADBWOp(OpcodeArgs);
void VPSHUFBOp(OpcodeArgs);
void VPSHUFWOp(OpcodeArgs, IR::OpSize ElementSize, bool Low);
void VPSLLOp(OpcodeArgs, IR::OpSize ElementSize);
@@ -728,7 +681,6 @@ public:
void VPSLLVOp(OpcodeArgs);
void VPSRAOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAVDOp(OpcodeArgs);
@@ -736,17 +688,14 @@ public:
void VPSRLDOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLDQOp(OpcodeArgs);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKHOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKLOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VSHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VTESTPOp(OpcodeArgs);
void VTESTPOp(OpcodeArgs, IR::OpSize ElementSize);
void VZEROOp(OpcodeArgs);
@@ -830,32 +779,24 @@ public:
void XSaveOp(OpcodeArgs);
void PAlignrOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void UCOMISxOp(OpcodeArgs);
void UCOMISxOp(OpcodeArgs, IR::OpSize ElementSize);
void LDMXCSR(OpcodeArgs);
void STMXCSR(OpcodeArgs);
template<IR::OpSize ElementSize>
void PACKUSOp(OpcodeArgs);
void PACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void PACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PACKSSOp(OpcodeArgs);
void PMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
template<IR::OpSize ElementSize, bool Signed>
void PMULLOp(OpcodeArgs);
void MOVQ2DQ(OpcodeArgs, bool ToXMM);
template<bool ToXMM>
void MOVQ2DQ(OpcodeArgs);
template<IR::OpSize ElementSize>
void ADDSUBPOp(OpcodeArgs);
void ADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void PFNACCOp(OpcodeArgs);
void PFPNACCOp(OpcodeArgs);
void PSWAPDOp(OpcodeArgs);
template<uint8_t CompType>
void VPFCMPOp(OpcodeArgs);
void VPFCMPOp(OpcodeArgs, uint8_t CompType);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
@@ -864,16 +805,12 @@ public:
void PMADDWD(OpcodeArgs);
void PMADDUBSW(OpcodeArgs);
template<bool Signed>
void PMULHW(OpcodeArgs);
void PMULHW(OpcodeArgs, bool Signed);
void PMULHRSW(OpcodeArgs);
void MOVBEOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void HSUBP(OpcodeArgs);
template<IR::OpSize ElementSize>
void PHSUB(OpcodeArgs);
void HSUBP(OpcodeArgs, IR::OpSize ElementSize);
void PHSUB(OpcodeArgs, IR::OpSize ElementSize);
void PHADDS(OpcodeArgs);
void PHSUBS(OpcodeArgs);
@@ -920,25 +857,24 @@ public:
};
RefVSIB LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags);
template<OpSize AddrElementSize>
void VPGATHER(OpcodeArgs);
void VPGATHER(OpcodeArgs, OpSize AddrElementSize);
template<IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed>
void ExtendVectorElements(OpcodeArgs);
template<IR::OpSize ElementSize>
void VectorRound(OpcodeArgs);
void AVXExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
void ExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
Ref VectorBlend(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VectorBlend(OpcodeArgs);
Ref VectorBlendImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorBlend(OpcodeArgs, IR::OpSize ElementSize);
void VectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
void PTestOpImpl(OpSize Size, Ref Dest, Ref Src);
void PTestOp(OpcodeArgs);
void AVXPHMINPOSUWOp(OpcodeArgs);
void PHMINPOSUWOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void DPPOp(OpcodeArgs);
void DPPOp(OpcodeArgs, IR::OpSize ElementSize);
void MPSADBWOp(OpcodeArgs);
void PCLMULQDQOp(OpcodeArgs);
@@ -1376,6 +1312,7 @@ private:
};
FEXCore::Context::ContextImpl* CTX {};
FEXCore::Core::InternalThreadState* Thread;
constexpr static unsigned FullNZCVMask = (1U << FEXCore::X86State::RFLAG_CF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_ZF_RAW_LOC) |
(1U << FEXCore::X86State::RFLAG_SF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_OF_RAW_LOC);
@@ -1443,7 +1380,7 @@ private:
Ref PALIGNROpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1, const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm, bool IsAVX);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask, bool IsAVX);
Ref PHADDSOpImpl(OpSize Size, Ref Src1, Ref Src2);
@@ -1481,7 +1418,7 @@ private:
Ref PSRLDOpImpl(OpcodeArgs, IR::OpSize ElementSize, Ref Src, Ref ShiftVec);
Ref SHUFOpImpl(OpcodeArgs, IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
Ref SHUFOpImpl(IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
void VMASKMOVOpImpl(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DataSize, bool IsStore, const X86Tables::DecodedOperand& MaskOp,
const X86Tables::DecodedOperand& DataOp);
@@ -1591,6 +1528,7 @@ private:
}
AddressMode DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, MemoryAccessType AccessType, bool IsLoad);
uint64_t CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad);
Ref LoadSource(RegClass Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {});
@@ -630,7 +630,7 @@ void OpDispatchBuilder::AVX128_VPSIGN(OpcodeArgs, IR::OpSize ElementSize) {
}
void OpDispatchBuilder::AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : ElementSize;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false);
@@ -1260,26 +1260,26 @@ void OpDispatchBuilder::AVX128_VAESKeyGenAssist(OpcodeArgs) {
}
void OpDispatchBuilder::AVX128_VPCMPESTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, false);
PCMPXSTRXOpImpl(Op, true, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPESTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, true);
PCMPXSTRXOpImpl(Op, true, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
}
void OpDispatchBuilder::AVX128_VPCMPISTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, false);
PCMPXSTRXOpImpl(Op, false, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPISTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, true);
PCMPXSTRXOpImpl(Op, false, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
@@ -1399,13 +1399,13 @@ void OpDispatchBuilder::AVX128_VSHUF(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
Result.Low = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
const uint8_t ShiftAmount = ElementSize == OpSize::i32Bit ? 0 : 2;
Result.High = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
Result.High = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -1484,12 +1484,12 @@ void OpDispatchBuilder::AVX128_VBLEND(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = VectorBlend(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
Result.Low = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
if (Is128Bit) {
Result = AVX128_Zext(Result.Low);
} else {
Result.High = VectorBlend(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
Result.High = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -2293,9 +2293,8 @@ void OpDispatchBuilder::AVX128_VCVTPS2PH(OpcodeArgs) {
_PopRoundingMode(OldFPCR);
}
// We need to eliminate upper junk if we're storing into a register with
// a 256-bit source (VCVTPS2PH's destination for registers is an XMM).
if (Op->Src[0].IsGPR() && SrcSize == OpSize::i256Bit) {
// We need to zero the upper 128 bits if we're storing into a register
if (Op->Dest.IsGPR()) {
Result = AVX128_Zext(Result.Low);
}
@@ -5,9 +5,9 @@
namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x0D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
{0x1D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x1D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, false>},
@@ -15,15 +15,15 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x8A, 1, &OpDispatchBuilder::PFNACCOp},
{0x8E, 1, &OpDispatchBuilder::PFPNACCOp},
{0x90, 1, &OpDispatchBuilder::VPFCMPOp<1>},
{0x90, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 1>},
{0x94, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryDuplicateOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x97, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, true>},
{0x9A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x9E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0xA0, 1, &OpDispatchBuilder::VPFCMPOp<2>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 2>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
// Can be treated as a move
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -32,7 +32,7 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, OpSize::i32Bit>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0xB0, 1, &OpDispatchBuilder::VPFCMPOp<0>},
{0xB0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 0>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
// Can be treated as a move
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -20,18 +20,18 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x03), 1, &OpDispatchBuilder::PHADDS},
{OPD(PF_38_NONE, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_66, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_66, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x10), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i8Bit>},
@@ -44,22 +44,22 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i32Bit>},
@@ -9,13 +9,13 @@ namespace FEXCore::IR {
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr DispatchTableEntry Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::VectorBlend<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::VectorBlend<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::VectorBlend<OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i16Bit>},
{OPD(REX, PF_3A_NONE, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(REX, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
@@ -24,17 +24,17 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x17), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::PINSROp<OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x21), 1, &OpDispatchBuilder::InsertPSOp},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, false>},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, false>},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, false>},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, false>},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
@@ -65,7 +65,7 @@ constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i32Bit>},
};
#undef PF_3A_NONE
@@ -69,12 +69,12 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
@@ -6,8 +6,7 @@ namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x07, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x06, 4, &OpDispatchBuilder::PermissionRestrictedOp},
{0x0B, 1, &OpDispatchBuilder::INTOp},
{0x0E, 1, &OpDispatchBuilder::X87EMMS},
@@ -44,7 +43,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xBE, 2, &OpDispatchBuilder::MOVSXOp},
{0xC0, 2, &OpDispatchBuilder::XADDOp},
{0xC3, 1, &OpDispatchBuilder::MOVGPRNTOp},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC8, 8, &OpDispatchBuilder::BSWAPOp},
@@ -56,10 +55,10 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::InsertMMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
@@ -71,7 +70,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
@@ -79,15 +78,15 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x70, 1, &OpDispatchBuilder::PSHUFW8ByteOp},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i8Bit>},
@@ -95,7 +94,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x77, 1, &OpDispatchBuilder::X87EMMS},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i32Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i32Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i32Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
@@ -116,9 +115,9 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -131,7 +130,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
@@ -152,23 +151,23 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x12, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, false>},
{0x16, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, false>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x6F, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, false>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQOp, OpDispatchBuilder::VectorOpType::SSE>},
@@ -176,36 +175,36 @@ constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0xB8, 1, &OpDispatchBuilder::PopcountOp},
{0xBC, 1, &OpDispatchBuilder::TZCNT},
{0xBD, 1, &OpDispatchBuilder::LZCNT},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, false>},
};
constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
// x52 = Invalid
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x78, 1, &OpDispatchBuilder::Insertq_imm},
{0x79, 1, &OpDispatchBuilder::Insertq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<false>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, false>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, false>},
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
@@ -217,10 +216,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
@@ -231,7 +230,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
@@ -239,15 +238,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x6C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
{0x6D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i64Bit>},
{0x6E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
@@ -260,15 +259,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x78, 1, nullptr}, // GROUP 17
{0x79, 1, &OpDispatchBuilder::Extrq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
{0x7F, 1, &OpDispatchBuilder::MOVVectorAlignedOp},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i64Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i32Bit>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i64Bit>},
@@ -288,10 +287,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, false>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -304,7 +303,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
File diff suppressed because it is too large. Load diff
@@ -163,7 +163,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
SubWithFlags(OpSize::i64Bit, Exponent, 0x7fff);
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
@@ -178,7 +178,8 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -623,13 +624,10 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
_F80FYL2XP1Stack();
} else {
_F80FYL2XStack();
}
_F80FYL2XStack();
}
void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
@@ -106,12 +106,24 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
bool CanUseFloatReg = Size == OpSize::i64Bit;
if (CanUseFloatReg) {
// If possible, it's faster to keep the data in an FPR than doing a GPR transfer.
if (Truncate) {
data = _Vector_FToZS(OpSize::i128Bit, OpSize::i64Bit, data);
} else {
data = _Vector_FToS(OpSize::i128Bit, OpSize::i64Bit, data);
}
StoreResultFPR_WithOpSize(Op, Op->Dest, data, OpSize::i64Bit, OpSize::i8Bit);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -370,6 +382,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// if val is +/-Inf, then sig = val, exp = +inf
// if val is NaN, then sig = val, exp = val
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
@@ -379,6 +393,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// Inf/NaN case
Ref ExpInfOnlyV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x7ff0'0000'0000'0000UL));
Ref ExpNanV = Node;
Ref SigInfV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
@@ -388,12 +407,24 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
// Mantissa non-zero => NaN (exp result = input); else Inf (exp result = +Inf)
Ref Mantissa = _And(OpSize::i64Bit, Gpr, Constant(0x000f'ffff'ffff'ffffULL));
_TestNZ(OpSize::i64Bit, Mantissa, Constant(~0ULL));
Ref ExpInfV = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfOnlyV, ExpNanV);
// Biased exponent == 0x7ff => Inf/NaN path, else non-zero-case.
Ref BiasedExp = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
SubWithFlags(OpSize::i64Bit, BiasedExp, 0x7ff);
Ref ExpNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfV, ExpNZV);
Ref SigNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigInfV, SigNZV);
// Zero folds on top.
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZOrInf);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZOrInf);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::iInvalid);
@@ -200,8 +200,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0}},
// Instructions
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -210,16 +210,16 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x06, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_06] }}},
{0x07, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_07] }}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x0E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_0E] }}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -227,8 +227,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x16, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_16] }}},
{0x17, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_17] }}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -236,24 +236,24 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x1E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1E] }}},
{0x1F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1F] }}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x27, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_27] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x2F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_2F] }}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -310,8 +310,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -34,7 +34,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> H0F3A_ArchSelect_LUT = {{
// ENTRY_1_3A_66_22
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::PINSROp<IR::OpSize::i64Bit> }},
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PINSROp, IR::OpSize::i64Bit> }},
},
}};
@@ -28,31 +28,31 @@ enum PrimaryGroup_LUT {
constexpr std::array<X86InstInfo[2], ENTRY_MAX> PrimaryGroup_ArchSelect_LUT = {{
{
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
@@ -66,23 +66,23 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr U16U8InfoStruct PrimaryGroupOpTable[] = {
// GROUP_1 | 0x80 | reg
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
// Duplicates the 0x80 opcode group
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_0] }}},
@@ -94,14 +94,14 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 6), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_6] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 7), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_7] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 2
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
@@ -161,8 +161,8 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
// GROUP 3
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 0), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 1), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 4), 1, X86InstInfo{"MUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 5), 1, X86InstInfo{"IMUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 1, X86InstInfo{"DIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -170,21 +170,21 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 0), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 1), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 4), 1, X86InstInfo{"MUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 5), 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 1, X86InstInfo{"DIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 7), 1, X86InstInfo{"IDIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// GROUP 4
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 2), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 5
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END | FLAGS_CALL , 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0}},
@@ -139,37 +139,37 @@ constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGr
{OPD(TYPE_GROUP_8, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 9
@@ -179,7 +179,7 @@ constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGr
// CMPXCHG8B/16B works with all prefixes
// Tooling fails to decode CMPXCHG with prefix
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -188,7 +188,7 @@ constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGr
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -197,7 +197,7 @@ constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGr
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -206,7 +206,7 @@ constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGr
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -31,19 +31,19 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
{
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
}};
@@ -61,8 +61,8 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0x05, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_05] }}},
{0x06, 1, X86InstInfo{"CLTS", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x07, 1, X86InstInfo{"SYSRET", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x0A, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x0B, 1, X86InstInfo{"UD2", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY, 0}},
{0x0C, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
@@ -205,23 +205,23 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xA0, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A0] }}},
{0xA1, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A1] }}},
{0xA2, 1, X86InstInfo{"CPUID", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xA4, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xA5, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xA6, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xA8, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A8] }}},
{0xA9, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A9] }}},
{0xAA, 1, X86InstInfo{"RSM", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xAC, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xAD, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xAE, 1, X86InstInfo{"", TYPE_GROUP_15, FLAGS_NO_OVERLAY, 0}},
{0xAF, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB2, 1, X86InstInfo{"LSS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB4, 1, X86InstInfo{"LFS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB5, 1, X86InstInfo{"LGS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB6, 1, X86InstInfo{"MOVZX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
@@ -229,14 +229,14 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xB8, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{0xB9, 1, X86InstInfo{"", TYPE_GROUP_10, FLAGS_NO_OVERLAY, 0}},
{0xBA, 1, X86InstInfo{"", TYPE_GROUP_8, FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xBC, 1, X86InstInfo{"BSF", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBD, 1, X86InstInfo{"BSR", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBE, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xBF, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xC2, 1, X86InstInfo{"CMPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{0xC3, 1, X86InstInfo{"MOVNTI", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST, 0}},
{0xC4, 1, X86InstInfo{"PINSRW", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX | FLAGS_SF_SRC_GPR, 1}},
@@ -474,7 +474,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::VMOVSLDUPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, true>},
{OPD(1, 0b11, 0x12), 1, &OpDispatchBuilder::VMOVDDUPOp},
{OPD(1, 0b00, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
@@ -487,7 +487,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::VMOVSHDUPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, true>},
{OPD(1, 0b00, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
@@ -496,36 +496,36 @@ namespace AVX256 {
{OPD(1, 0b00, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b01, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i64Bit>},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{OPD(1, 0b01, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRECP, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{OPD(1, 0b01, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
@@ -541,42 +541,42 @@ namespace AVX256 {
{OPD(1, 0b00, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{OPD(1, 0b01, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{OPD(1, 0b01, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, true>},
{OPD(1, 0b01, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, true>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, true>},
{OPD(1, 0b00, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i8Bit>},
{OPD(1, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i16Bit>},
@@ -607,8 +607,8 @@ namespace AVX256 {
{OPD(1, 0b00, 0x77), 1, &OpDispatchBuilder::VZEROOp},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i32Bit>},
@@ -618,19 +618,19 @@ namespace AVX256 {
{OPD(1, 0b01, 0x7F), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x7F), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::VPINSRWOp},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xC5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(1, 0b00, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xD1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xD2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i32Bit>},
@@ -653,14 +653,14 @@ namespace AVX256 {
{OPD(1, 0b01, 0xE1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xE3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::VPMULHWOp<false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::VPMULHWOp<true>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, true>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, true>},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0xE8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{OPD(1, 0b01, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
@@ -671,11 +671,11 @@ namespace AVX256 {
{OPD(1, 0b01, 0xEE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSMAX, OpSize::i16Bit>},
{OPD(1, 0b01, 0xEF), 1, &OpDispatchBuilder::AVXVectorXOROp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b01, 0xF1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xF2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xF3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF5), 1, &OpDispatchBuilder::VPMADDWDOp},
{OPD(1, 0b01, 0xF6), 1, &OpDispatchBuilder::VPSADBWOp},
{OPD(1, 0b01, 0xF7), 1, &OpDispatchBuilder::MASKMOVOp},
@@ -689,8 +689,8 @@ namespace AVX256 {
{OPD(1, 0b01, 0xFE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
{OPD(2, 0b01, 0x00), 1, &OpDispatchBuilder::VPSHUFBOp},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x03), 1, &OpDispatchBuilder::VPHADDSWOp},
{OPD(2, 0b01, 0x04), 1, &OpDispatchBuilder::VPMADDUBSWOp},
@@ -698,14 +698,14 @@ namespace AVX256 {
{OPD(2, 0b01, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPHSUBOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x07), 1, &OpDispatchBuilder::VPHSUBSWOp},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::VPSIGN<OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::VPSIGN<OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::VPSIGN<OpSize::i32Bit>},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0B), 1, &OpDispatchBuilder::VPMULHRSWOp},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x13), 1, &OpDispatchBuilder::VCVTPH2PSOp},
{OPD(2, 0b01, 0x16), 1, &OpDispatchBuilder::VPERMDOp},
@@ -717,28 +717,28 @@ namespace AVX256 {
{OPD(2, 0b01, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(2, 0b01, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(2, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPACKUSOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x36), 1, &OpDispatchBuilder::VPERMDOp},
{OPD(2, 0b01, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
@@ -752,7 +752,7 @@ namespace AVX256 {
{OPD(2, 0b01, 0x3F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VUMAX, OpSize::i32Bit>},
{OPD(2, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::AVXPHMINPOSUWOp},
{OPD(2, 0b01, 0x45), 1, &OpDispatchBuilder::VPSRLVOp},
{OPD(2, 0b01, 0x46), 1, &OpDispatchBuilder::VPSRAVDOp},
{OPD(2, 0b01, 0x47), 1, &OpDispatchBuilder::VPSLLVOp},
@@ -764,13 +764,13 @@ namespace AVX256 {
{OPD(2, 0b01, 0x78), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i8Bit>},
{OPD(2, 0b01, 0x79), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i16Bit>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::VPMASKMOVOp<false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::VPMASKMOVOp<true>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, true>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x96), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, true, 1, 3, 2>}, // VFMADDSUB
{OPD(2, 0b01, 0x97), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, false, 1, 3, 2>}, // VFMSUBADD
@@ -820,10 +820,10 @@ namespace AVX256 {
{OPD(3, 0b01, 0x04), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x06), 1, &OpDispatchBuilder::VPERM2Op},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0C), 1, &OpDispatchBuilder::VPBLENDDOp},
{OPD(3, 0b01, 0x0D), 1, &OpDispatchBuilder::VBLENDPDOp},
{OPD(3, 0b01, 0x0E), 1, &OpDispatchBuilder::VPBLENDWOp},
@@ -837,15 +837,15 @@ namespace AVX256 {
{OPD(3, 0b01, 0x18), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x19), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x1D), 1, &OpDispatchBuilder::VCVTPS2PHOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::VPINSRBOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i8Bit>},
{OPD(3, 0b01, 0x21), 1, &OpDispatchBuilder::VINSERTPSOp},
{OPD(3, 0b01, 0x22), 1, &OpDispatchBuilder::VPINSRDQOp},
{OPD(3, 0b01, 0x38), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x39), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::VDPPOp<OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::VDPPOp<OpSize::i64Bit>},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x42), 1, &OpDispatchBuilder::VMPSADBWOp},
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::VPCLMULQDQOp},
@@ -855,10 +855,10 @@ namespace AVX256 {
{OPD(3, 0b01, 0x4B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i64Bit>},
{OPD(3, 0b01, 0x4C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i8Bit>},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, true>},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, true>},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, true>},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, true>},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
@@ -392,8 +392,11 @@ namespace InstFlags {
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SUPPORTS_LOCK = (1ULL << 31);
// Flags [57..32]: Undefined
// Flags [60..58]: Dst size
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
// Flags [63..61]: Src size
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
constexpr InstFlagType SIZE_MASK = 0b111;
+25 -7
View File
@@ -717,7 +717,7 @@
"HasSideEffects": true
},
"CacheLineClean GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline cleanat the address specified",
"Desc": ["Does a 64 byte cacheline clean at the address specified",
"Only cleans the data cachelines. Doesn't do any zeroing",
"Skips the invalidation step of the CacheLineClear operation"
],
@@ -2758,27 +2758,31 @@
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM1 FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2XP1 FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": true
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
@@ -3208,6 +3212,20 @@
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"FPR = F80FYL2XP1Stack": {
"Desc": [
"Computes ST1 * log2(1 + ST0)",
"Stores the result in ST1, and pops the top of the stack.",
"Returns the new value at the top of the stack, i.e. the result of the operation."
],
"HasSideEffects": true,
"DestSize": "OpSize::i128Bit",
"X87": true
},
"FPR = F80FYL2XP1 FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"F80VBSLStack OpSize:#RegisterSize, FPR:$VectorMask, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Does a vector bitwise select.",
+4
View File
@@ -208,6 +208,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorCon
return "movmaskb";
case NamedVectorConstant::NAMED_VECTOR_MOVMASKB_UPPER:
return "movmaskb_upper";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP:
return "v256_mid_element_swap";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER:
return "v256_mid_element_swap_upper";
case NamedVectorConstant::NAMED_VECTOR_ZERO:
return "vectorzero";
case NamedVectorConstant::NAMED_VECTOR_X87_ONE:
@@ -186,12 +186,12 @@ public:
}
[[nodiscard]]
unsigned PostRA() const {
bool PostRA() const {
return GetHeader()->PostRA;
}
[[nodiscard]]
unsigned SpillSlots() const {
uint32_t SpillSlots() const {
return GetHeader()->SpillSlots;
}
+5 -10
View File
@@ -8,23 +8,18 @@ class CPUIDEmu;
struct HostFeatures;
} // namespace FEXCore
namespace FEXCore::Utils {
class IntrusivePooledAllocator;
}
namespace FEXCore::IR {
class Pass;
class RegisterAllocationPass;
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID);
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures&, OpSize GPROpSize);
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID);
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures&, OpSize GPROpSize);
namespace Validation {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
fextl::unique_ptr<Pass> CreateIRValidation();
} // namespace Validation
namespace Debug {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper();
fextl::unique_ptr<Pass> CreateIRDumper();
}
} // namespace FEXCore::IR
@@ -67,7 +67,7 @@ void IRDumper::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper() {
fextl::unique_ptr<Pass> CreateIRDumper() {
return fextl::make_unique<IRDumper>();
}
} // namespace FEXCore::IR::Debug
@@ -271,7 +271,7 @@ void IRValidation::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation() {
fextl::unique_ptr<Pass> CreateIRValidation() {
return fextl::make_unique<IRValidation>();
}
} // namespace FEXCore::IR::Validation
@@ -747,7 +747,7 @@ void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination() {
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination() {
return fextl::make_unique<DeadFlagCalculationEliminination>();
}
@@ -781,7 +781,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
IR->GetHeader()->PostRA = true;
}
fextl::unique_ptr<IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID) {
fextl::unique_ptr<IR::Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID) {
return fextl::make_unique<ConstrainedRAPass>(CPUID);
}
} // namespace FEXCore::IR
@@ -188,7 +188,7 @@ private:
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
if (Features.SupportsSVE()) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
@@ -785,6 +785,12 @@ void X87StackOptimization::Run(IREmitter* Emit) {
break;
}
case OP_F80FYL2XP1STACK: {
HandleBinopStack(OP_F64FYL2XP1, false, OP_F80FYL2XP1, 1, 0, 1);
StackPop();
break;
}
case OP_F80ATANSTACK: {
HandleBinopStack(OP_F64ATAN, false, OP_F80ATAN, 1, 1, 0);
StackPop();
@@ -1225,7 +1231,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
return;
}
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures& Features, OpSize GPROpSize) {
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures& Features, OpSize GPROpSize) {
return fextl::make_unique<X87StackOptimization>(Features, GPROpSize);
}
} // namespace FEXCore::IR
+20 -20
View File
@@ -114,27 +114,17 @@ FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
};
for (auto Bits : TLBSizes) {
uintptr_t Size = 1ULL << Bits;
// Just try allocating
// We can't actually determine VA size on ARM safely
auto Find = [](uintptr_t Size) -> bool {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void* Ptr = ::mmap(reinterpret_cast<void*>(Size - FEXCore::Utils::FEX_PAGE_SIZE * i), FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE,
MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
if (Ptr == (void*)(Size - FEXCore::Utils::FEX_PAGE_SIZE * i)) {
return true;
}
}
}
return false;
};
if (Find(Size)) {
HostVASize = Bits;
// We can't actually determine VA size on ARM safely.
// Instead, try allocating the page at the top of the range.
// If this succeeds OR the page is reported as already existing,
// we know we're in valid VA space. Otherwise, we must go lower.
void* Addr = reinterpret_cast<void*>((1ULL << Bits) - FEXCore::Utils::FEX_PAGE_SIZE);
void* Ptr = ::mmap(Addr, FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
}
if (Ptr != (void*)~0ULL || errno == EEXIST) {
return Bits;
}
}
@@ -262,9 +252,19 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
// Block remaining memory gaps
bool SupportsDontDump = true;
for (auto RegionIt = Regions.begin(); RegionIt != Regions.end(); ++RegionIt) {
auto Alloc = ::mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
if (SupportsDontDump) {
// Mark these regions as don't dump so that coredump doesn't try dumping large unmapped regions.
// Ideally coredump would be smart enough to only dump resident pages, but here we are.
auto Result = madvise(RegionIt->Ptr, RegionIt->Size, MADV_DONTDUMP);
if (Result == -1) {
SupportsDontDump = false;
}
}
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "StealMemoryRegion: mmap({}, {:x}) failed: {}", fmt::ptr(RegionIt->Ptr), RegionIt->Size, errno);
LogMan::Throw::AFmt(Alloc == RegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(RegionIt->Ptr));
}
-1
View File
@@ -2,7 +2,6 @@
#ifdef ENABLE_FEX_ALLOCATOR
#include <rpmalloc/rpmalloc.h>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#include <sys/mman.h>
#else
+54 -18
View File
@@ -343,11 +343,12 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
// 32bit
uint64_t Addr = GPRs[AddressReg];
// Lower register must be even, so only upper register can be 31.
uint32_t DesiredLower = GPRs[DesiredReg1];
uint32_t DesiredUpper = GPRs[DesiredReg2];
uint32_t DesiredUpper = DesiredReg2 == 31 ? 0 : GPRs[DesiredReg2];
uint32_t ExpectedLower = GPRs[ExpectedReg1];
uint32_t ExpectedUpper = GPRs[ExpectedReg2];
uint32_t ExpectedUpper = ExpectedReg2 == 31 ? 0 : GPRs[ExpectedReg2];
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
@@ -1352,7 +1353,9 @@ static std::optional<uint64_t> DoCAS(uint32_t Size, uint64_t Desired, uint64_t E
}
static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg, uint32_t* StrictSplitLockMutex) {
std::optional<uint64_t> Res = DoCAS(Size, GPRs[DesiredReg], GPRs[ExpectedReg], GPRs[AddressReg], StrictSplitLockMutex);
uint64_t Desired = DesiredReg == 31 ? 0 : GPRs[DesiredReg];
uint64_t Expected = ExpectedReg == 31 ? 0 : GPRs[ExpectedReg];
std::optional<uint64_t> Res = DoCAS(Size, Desired, Expected, GPRs[AddressReg], StrictSplitLockMutex);
if (!Res.has_value()) {
return false;
}
@@ -1384,6 +1387,8 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
uint8_t Op = (Instr >> 12) & 0xF;
uint64_t Source = SourceReg == 31 ? 0 : GPRs[SourceReg];
if (Size == 2) {
auto NOPExpected = [](uint16_t SrcVal, uint16_t) -> uint16_t {
return SrcVal;
@@ -1420,7 +1425,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS16<true>(GPRs[SourceReg],
auto Res = DoCAS16<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1465,7 +1470,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS32<true>(GPRs[SourceReg],
auto Res = DoCAS32<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1510,7 +1515,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS64<true>(GPRs[SourceReg],
auto Res = DoCAS64<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1572,9 +1577,11 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
uint64_t Addr = GPRs[AddressReg] + Offset;
constexpr bool DoRetry = false;
uint64_t Data = DataReg == 31 ? 0 : GPRs[DataReg];
if (Size == 2) {
DoCAS16<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint16_t SrcVal, uint16_t) -> uint16_t {
@@ -1589,7 +1596,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 4) {
DoCAS32<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint32_t SrcVal, uint32_t) -> uint32_t {
@@ -1604,7 +1611,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 8) {
DoCAS64<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint64_t SrcVal, uint64_t) -> uint64_t {
@@ -1834,6 +1841,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return Desired;
};
uint64_t Source = DataSourceReg == 31 ? 0 : GPRs[DataSourceReg];
if (Size == 2) {
using AtomicType = uint16_t;
CASDesiredFn<AtomicType> DesiredFunction {};
@@ -1852,7 +1860,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS16<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS16<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1879,7 +1887,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS32<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS32<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1906,7 +1914,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS64<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS64<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
if (AtomicFetch && ResultReg != 31) {
@@ -1947,8 +1955,24 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
uint32_t* StrictSplitLockMutex {CTX->Config.StrictInProcessSplitLocks ? &CTX->StrictSplitLockMutex : nullptr};
if (!IsJIT) [[unlikely]] {
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if ((Instr & ArchHelpers::Arm64::CASPAL_MASK) == ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (ArchHelpers::Arm64::HandleCASPAL(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & ArchHelpers::Arm64::CASAL_MASK) == ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (ArchHelpers::Arm64::HandleCASAL(GPRs, Instr, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0)) {
// Skip this instruction now
return 4;
@@ -1991,22 +2015,34 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
} else if ((Instr & ArchHelpers::Arm64::STLXR_MASK) == ArchHelpers::Arm64::STLXR_INST) { // STLXR*
uint32_t StatusReg = Instr << 11 >> 27;
// // Emulate exclusive store by validating the address and value against the last unaligned LDAXR*.
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || Size > Thread->ExclusiveStore.Size) {
uint32_t SizeBytes = 1u << Size;
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || SizeBytes > Thread->ExclusiveStore.Size) {
if (StatusReg != 31) {
GPRs[StatusReg] = 1;
}
return 4;
}
if (std::optional<uint64_t> Prev =
DoCAS(Size, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
DoCAS(SizeBytes, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
if (StatusReg != 31) {
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, Size);
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, SizeBytes);
}
Thread->ExclusiveStore.Size = 0;
return 4;
}
} else if ((Instr & ArchHelpers::Arm64::ATOMIC_MEM_MASK) == ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (ArchHelpers::Arm64::HandleAtomicMemOp(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: 0x{:x} Instruction: 0x{:08x}\n", Op, ProgramCounter, PC[0]);
return std::nullopt;
}
}
return 0;
LogMan::Msg::EFmt("Unhandled non-JIT atomic");
return std::nullopt;
}
const auto Frame = Thread->CurrentFrame;
+85 -7
View File
@@ -8,6 +8,7 @@
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <atomic>
#include <cstdint>
@@ -178,7 +179,55 @@ private:
CodeMapOpener& FileOpener;
};
class AbstractCodeCache;
/**
* Manages runtime state associated with a mapped code cache file.
*
* The mapped file pointer is managed by the frontend and must be valid
* throughout the lifetime of this object.
*/
struct MappedCodeCacheFile {
// Calls UnregisterMappedCodeBuffer internally, see its docstring about synchronization requirements
~MappedCodeCacheFile();
// If not nullptr, the MappedCodeCacheFile will be unregistered from this on destruction
AbstractCodeCache* CacheManager;
std::span<std::byte> MappedFile; // Mapped data of the whole cache file
std::span<std::byte> CodeBufferInFile; // Subspan of cached ARM64 data within MappedFile (pre-relocation)
std::span<std::byte> CodeBuffer; // Cached ARM64 data used for execution (post-relocation; owned by MappedCodeCacheFile)
std::byte* BlockListInFile; // Pointer to BlockListEntry data within MappedFile
uint32_t NumBlocks; // Number of BlockListEntry objects
uint32_t NumCodePages; // Number of code page entrypoint mappings
struct PageRelocationRange {
uint32_t Offset; // In bytes from start of file
uint32_t Length; // Number of relocations
};
// List of relocation ranges in the mapped cache file, grouped by the code page they apply to.
// This vector is indexed by the relative page offset from the start of the ARM64 code data.
//
// For example PageRelocationRanges[1] == { 0x100, 0x20 } means:
// - there are 0x20 bytes of relocation data at offset 0x100 in the cache file
// - these 0x20 bytes of relocation data will patch data at CodeBuffer[0x1000..0x2000]
fextl::vector<PageRelocationRange> PageRelocationRanges;
fextl::vector<bool> LoadedPages;
uint64_t GuestBase {}; // Guest base address for relocation application
// Helper member to prevent moving/copying without disallowing aggregate-construction
std::atomic<int> disallow_copy_or_move;
size_t NumPages() const {
return CodeBuffer.size_bytes() / FEXCore::Utils::FEX_PAGE_SIZE;
}
};
class AbstractCodeCache {
fextl::vector<std::span<std::byte>> MappedCodeBuffers;
public:
virtual ~AbstractCodeCache() = default;
@@ -190,13 +239,6 @@ public:
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
* Returns false if the provided cache file is invalid, and true otherwise.
*/
virtual bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) = 0;
/**
* Bundles the current Core state (CodeBuffer, GuestToHostMapping, ...) to a code cache and writes it to the given file descriptor.
* Returns true on success.
@@ -207,6 +249,42 @@ public:
* Function to be called before compiling any code for caching purposes
*/
virtual void InitiateCacheGeneration() = 0;
/**
* Loads a code cache from mapped memory.
*
* Code sections must be enabled in a second step (see EnableLoadedSection).
* Afterwards, individual code pages must be finalized using FinalizeCodePages.
*
* On success, this returns a MappedCodeCacheFile that must be kept alive
* as long the cache is in use.
*/
virtual fextl::unique_ptr<MappedCodeCacheFile> LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo&, uint64_t FileStartVA) = 0;
/**
* Registers cached blocks for the given file section to the LookupCache.
*
* Also runs extended cache validation if enabled.
*/
virtual bool EnableLoadedSection(Core::InternalThreadState*, MappedCodeCacheFile&, const ExecutableFileSectionInfo&) = 0;
/**
* Extend the given code range so that it can be safely finalized.
*
* This is required for example to avoid dangling page-crossing FEX relocations on the edges
*
* StartPage and EndPage a 0-based relative page offsets into the cached code.
*/
static std::span<std::byte> SelectCodeRangeToFinalize(MappedCodeCacheFile&, size_t StartPage, size_t EndPage);
/**
* Finalize code pages in the given range (see SelectCodePagesToFinalize) for execution.
*/
virtual void FinalizeCodePages(MappedCodeCacheFile&, std::span<std::byte> CodeRange) = 0;
void RegisterMappedCodeBuffer(MappedCodeCacheFile&);
void UnregisterMappedCodeBuffer(MappedCodeCacheFile&);
bool IsAddressInMappedCodeBuffer(uintptr_t Address) const;
};
} // namespace FEXCore
+9
View File
@@ -302,6 +302,7 @@ enum FallbackHandlerIndex {
OPINDEX_F80MUL,
OPINDEX_F80DIV,
OPINDEX_F80FYL2X,
OPINDEX_F80FYL2XP1,
OPINDEX_F80ATAN,
OPINDEX_F80FPREM1,
OPINDEX_F80FPREM,
@@ -315,6 +316,7 @@ enum FallbackHandlerIndex {
OPINDEX_F64ATAN,
OPINDEX_F64F2XM1,
OPINDEX_F64FYL2X,
OPINDEX_F64FYL2XP1,
OPINDEX_F64FPREM,
OPINDEX_F64FPREM1,
OPINDEX_F64SCALE,
@@ -379,6 +381,13 @@ struct JITPointers {
uint64_t F64SinHandler {};
uint64_t F64CosHandler {};
uint64_t F64TanHandler {};
uint64_t F64F2XM1Handler {};
uint64_t F64ScaleHandler {};
uint64_t F64AtanHandler {};
uint64_t F64FYL2XHandler {};
uint64_t F64FYL2XP1Handler {};
uint64_t F64FPREMHandler {};
uint64_t F64FPREM1Handler {};
/** @} */
// Copy of process-wide named vector constants data.
+23 -6
View File
@@ -5,13 +5,20 @@
#include <cstdint>
namespace FEXCore {
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
struct HostFeatures {
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
// Whether or not the host supports any kind of SVE implementation.
[[nodiscard]]
bool SupportsSVE() const {
return SupportsSVE128 || SupportsSVE256;
}
uint32_t DCacheLineSize {};
uint32_t ICacheLineSize {};
bool SupportsCacheMaintenanceOps {};
@@ -42,11 +49,21 @@ struct HostFeatures {
bool Supports3DNow {};
bool SupportsSSE4a {};
bool SupportsMOPS {};
bool PreferZVAForVZero {};
// Float exception behaviour
bool SupportsAFP {};
bool SupportsFloatExceptions {};
// Changes code generation slightly.
enum class HostTypeEnum {
Unknown,
Linux,
Wow64,
Arm64ec,
};
HostTypeEnum HostType {};
// Flag if this is InstCountCI
bool IsInstCountCI {};
@@ -15,7 +15,6 @@
namespace FEXCore {
class LookupCache;
class CompileService;
struct JITSymbolBuffer;
} // namespace FEXCore
@@ -102,10 +101,6 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
NonMovableUniquePtr<FEXCore::IR::PassManager> PassManager;
NonMovableUniquePtr<JITSymbolBuffer> SymbolBuffer;
std::shared_ptr<FEXCore::CompileService> CompileService;
std::shared_mutex ObjectCacheRefCounter {};
// This pointer is owned by the frontend.
FEXCore::SHMStats::ThreadStats* ThreadStats {};
@@ -130,9 +125,9 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
alignas(FEXCore::Utils::FEX_PAGE_SIZE) uint8_t InterruptFaultPage[FEXCore::Utils::FEX_PAGE_SIZE];
};
static_assert(std::is_standard_layout_v<FEXCore::Core::InternalThreadState>);
static_assert((offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) <
FEXCore::Utils::FEX_PAGE_SIZE,
"Fault page is outside of immediate range from CPU state");
static_assert(sizeof(FEXCore::Core::InternalThreadState) == (FEXCore::Utils::FEX_PAGE_SIZE * 2));
// Maximum unsigned-offset store range for fault page.
static_assert(
(offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState)) <= 65520,
"Fault page is outside of immediate range from CPU state");
} // namespace FEXCore::Core
+6
View File
@@ -34,6 +34,12 @@ enum NamedVectorConstant : uint8_t {
NAMED_VECTOR_MOVMASKB,
NAMED_VECTOR_MOVMASKB_UPPER,
// Used to swap [0, 1, 2, 3] into [0, 2, 1, 3] in lieu
// of Q operations introduced in SVE2.1. Can be removed when
// such operations become available.
NAMED_VECTOR_256_MID_ELEMENT_SWAP,
NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER,
NAMED_VECTOR_X87_ONE,
NAMED_VECTOR_X87_LOG2_10,
NAMED_VECTOR_X87_LOG2_E,
@@ -2,7 +2,6 @@
#pragma once
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/mman.h>
#include <sys/user.h>
#include <sys/prctl.h>
@@ -35,6 +35,9 @@ public:
const auto Result = pthread_mutex_lock(&Mutex);
LOGMAN_THROW_A_FMT(Result == 0, "{} failed to lock with {}", __func__, Result);
}
bool try_lock() {
return pthread_mutex_trylock(&Mutex) == 0;
}
void unlock() {
const auto Result = pthread_mutex_unlock(&Mutex);
LOGMAN_THROW_A_FMT(Result == 0, "{} failed to unlock with {}", __func__, Result);
+5 -5
View File
@@ -361,6 +361,11 @@ def IsSupportedKernel():
return version_check(GetKernelVersion()) >= version_check("5.15")
def main():
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
# Only run on supported arch
if not IsSupportedArch():
print ( "{} is not a supported architecture".format(GetArch()))
@@ -371,11 +376,6 @@ def main():
print ( "Kernel {} is too old. FEX needs 5.15 minimum".format(GetKernelVersion()))
ExitWithStatus(-1)
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
if GetDistro()[0] == "ubuntu":
print ("Getting PPA status: {}".format(("NotInstalled", "Installed")[GetPPAStatus()]))
+1 -1
View File
@@ -656,7 +656,7 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && Global) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
return fextl::fmt::format("{}/../share/fex-emu/", PortableInfo.InterpreterPath);
} else if (ConfigOverride && !Global) {
fextl::string AppConfigStr = ConfigOverride;
if (FHU::Filesystem::IsRelative(AppConfigStr)) {
+24 -10
View File
@@ -17,9 +17,9 @@
#include <fcntl.h>
#include <linux/limits.h>
#include <unistd.h>
#include <sys/poll.h>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/types.h>
@@ -137,21 +137,31 @@ fextl::string GetServerSocketName() {
return ServerSocketPath;
}
fextl::string GetServerSocketPath() {
fextl::string GetServerSocketPath(bool ForceTmp) {
fextl::string name {};
fextl::string Folder {};
#ifndef FEX_STEAM_SUPPORT
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
name = ServerSocketPath();
if (!ForceTmp) {
name = ServerSocketPath();
if (name.starts_with("/")) {
return name;
if (name.starts_with("/")) {
return name;
}
}
auto Folder = GetTempFolder();
Folder = GetTempFolder();
#else
// Under Steam the FEXServer's socket is a game-specific directory.
auto Folder = GetServerLockFolder();
if (ForceTmp) {
// If we're forcing temporary directory usage then the server socket path has exceeded sun_path 108 byte limit.
// Let's be a bit nice and put some more metadata in the server socket path.
const auto SteamID = getenv("SteamAppId") ?: "";
return fextl::fmt::format("{}/{}.FEXServer.Socket", GetTempFolder(), SteamID);
} else {
// Under Steam the FEXServer's socket is a game-specific directory.
Folder = GetServerLockFolder();
}
#endif
if (name.empty()) {
@@ -203,7 +213,11 @@ int ConnectToServer(ConnectionOption ConnectionOption) {
// Try again with a path-based socket, since abstract sockets will fail if we have been
// placed in a new netns as part of a sandbox.
auto ServerSocketPath = GetServerSocketPath();
auto ServerSocketPath = GetServerSocketPath(false);
if (ServerSocketPath.size() > sizeof(sockaddr_un::sun_path) - 1) {
LogMan::Msg::EFmt("Socket path '{}' too large for Unix domain sockets. Moving to tmp", ServerSocketPath);
ServerSocketPath = FEXServerClient::GetServerSocketPath(true);
}
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketPath.size(), sizeof(addr.sun_path) - 1);
+1 -1
View File
@@ -64,7 +64,7 @@ fextl::string GetServerRootFSLockFile();
fextl::string GetTempFolder();
fextl::string GetServerMountFolder();
fextl::string GetServerSocketName();
fextl::string GetServerSocketPath();
fextl::string GetServerSocketPath(bool ForceTmp);
int GetServerFD();
bool SetupClient(std::string_view InterpreterPath);
+16 -2
View File
@@ -539,6 +539,9 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
constexpr uint32_t PartNum_Oryon3 = 0x002;
constexpr uint32_t Implementer_Ampere = 0xc0;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
@@ -552,7 +555,7 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef ARCHITECTURE_arm64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
if (MIDR_Implementer == Implementer_QCOM && (MIDR_PartNum == PartNum_Oryon1 || MIDR_PartNum == PartNum_Oryon3)) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
@@ -592,6 +595,15 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
break;
}
}
if (MIDR_Implementer == Implementer_Ampere) {
// Ampere Computing CPUs that support CLZero should prefer using `dc zva` for vzero{upper,all} as its faster there.
// For Cortex CPUs it doesn't matter one way or the other.
// For Oryon CPUs, it is dramatically faster to avoid `dc zva` as it has dramatic stalls around barriers and overlapping `dc zva`.
//
// Because the `dc zva` optimization was implemented for Ampere, only use that path on the hardware.
HostFeatures->PreferZVAForVZero = HostFeatures->SupportsCLZERO;
}
}
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
@@ -657,6 +669,7 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsAES256 = HostFeatures.SupportsAVX && HostFeatures.SupportsAES;
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HostFeatures.PreferZVAForVZero = false;
if (CTR) {
HostFeatures.DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
@@ -741,7 +754,7 @@ FEXCore::HostFeatures FetchHostFeatures() {
uint64_t CTR = 0;
uint64_t MIDR = 0;
#ifdef ARCHITECTURE_arm64
#if defined(ARCHITECTURE_arm64) && !defined(VIXL_SIMULATOR)
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
__asm volatile("mrs %[ctr], ctr_el0" : [ctr] "=r"(CTR));
@@ -753,6 +766,7 @@ FEXCore::HostFeatures FetchHostFeatures() {
FetchHostFeatures(Features, HostFeatures, true, CTR, MIDR);
HostFeatures.SupportsCPUIndexInTPIDRRO = false;
HostFeatures.HostType = FEXCore::HostFeatures::HostTypeEnum::Linux;
return HostFeatures;
}
} // namespace FEX
+3 -3
View File
@@ -30,15 +30,15 @@ public:
auto data_7 = cpuid(0x7);
Feat_fsgsbase = data_7.ebx & (1U << 0);
Feat_bmi1 = data_7.ebx & (1U << 3);
Feat_avx &= data_7.ebx & (1U << 5);
Feat_avx = Feat_avx && (data_7.ebx & (1U << 5));
Feat_bmi2 = data_7.ebx & (1U << 8);
Feat_clwb = data_7.ebx & (1U << 24);
Feat_rand &= data_7.ebx & (1U << 18);
Feat_rand = Feat_rand && (data_7.ebx & (1U << 18));
Feat_adx = data_7.ebx & (1U << 19);
Feat_clflopt = data_7.ebx & (1U << 23);
Feat_sha = data_7.ebx & (1U << 29);
Feat_vaes = data_7.ecx & (1U << 9);
Feat_pclmulqdq &= data_7.ecx & (1U << 10);
Feat_pclmulqdq = Feat_pclmulqdq && (data_7.ecx & (1U << 10));
Feat_rdpid = data_7.ecx & (1U << 22);
}
+15 -14
View File
@@ -26,24 +26,25 @@ fextl::string GenerateSteamConfigTemplate(const FEX::Config::PortableInformation
return {};
}
// Try and find a mount point.
fextl::string MountPoint {};
// If the graphics provider was provided through an environment variable, then use that.
// Otherwise, try to find a mount point.
const char* GraphicsProvider = getenv("STEAM_COMPAT_GRAPHICS_PROVIDER");
const char* RuntimeDir = getenv("XDG_RUNTIME_DIR");
if (RuntimeDir) {
const char* CacheDir = getenv("XDG_CACHE_HOME");
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (GraphicsProvider) {
MountPoint = FHU::Filesystem::ParentPath(GraphicsProvider);
} else if (RuntimeDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", RuntimeDir);
} else if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else {
const char* CacheDir = getenv("XDG_CACHE_HOME");
if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
}
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
// Update the @FEX_COMPAT_TOOL@ config to point to the root of the depot.
+47 -39
View File
@@ -2,7 +2,7 @@
/*
$info$
tags: Bin|FEXBash
desc: Launches bash under FEX and passes arguments via -c to it
desc: Wrapper for invoking x86 bash from the rootfs using FEX
$end_info$
*/
@@ -13,58 +13,69 @@ $end_info$
#include <unistd.h>
#include <vector>
static std::string EnchantedPS1(const char* PS1Env) {
using namespace std::string_view_literals;
const char* ColorTerm = getenv("COLORTERM");
const bool SupportsTrueColor = isatty(STDIN_FILENO) && ColorTerm && (ColorTerm == "truecolor"sv || ColorTerm == "24bit"sv);
std::string PS1 = "PS1=FEXBash ";
if (SupportsTrueColor) {
// Rainbow FEXBash text matching the FEX logo
PS1 = "PS1="
R"(\[\e[38;2;251;0;145m\]F)"
R"(\[\e[38;2;199;0;198m\]E)"
R"(\[\e[38;2;141;68;253m\]X)"
R"(\[\e[38;2;31;159;248m\]B)"
R"(\[\e[38;2;0;218;181m\]a)"
R"(\[\e[38;2;129;229;93m\]s)"
R"(\[\e[38;2;240;220;10m\]h)"
R"(\[\e[0m\] )";
}
if (PS1Env) {
PS1 += &PS1Env[strlen("PS1=")];
} else {
PS1 += R"(\u@\h:\w> )";
}
return PS1;
}
int main(int argc, char** argv, char** const envp) {
// Skip argv[0].
const int ArgCount = argc - 1;
const bool EmptyArgs = ArgCount == 0;
// Skip argv[0]
const bool EmptyArgs = argc == 1;
std::vector<const char*> Argv;
// FEX will handle finding bash in the rootfs
// Use /bin/sh for -c commands and /bin/bash for interactive mode
// Use /bin/bash for interactive mode and /bin/sh when running a script or when using -c
const char* BashPath = EmptyArgs ? "/bin/bash" : "/bin/sh";
std::string FEXPath = std::filesystem::path(argv[0]).parent_path().string() + "/FEX";
auto FEXPath = std::filesystem::path(argv[0]).replace_filename("FEX");
// Check if a local FEX to FEXBash exists
// If it does then it takes priority over the installed one
if (!std::filesystem::exists(FEXPath)) {
char FEXBashPath[PATH_MAX];
auto Result = readlink("/proc/self/exe", FEXBashPath, PATH_MAX);
if (Result != -1) {
FEXPath = std::filesystem::path(&FEXBashPath[0], &FEXBashPath[Result]).parent_path().string() + "/FEX";
if (!std::filesystem::is_regular_file(FEXPath)) {
std::error_code ec;
auto FEXBashPath = std::filesystem::read_symlink("/proc/self/exe", ec);
if (!ec) {
FEXPath = FEXBashPath.replace_filename("FEX");
}
if (!std::filesystem::exists(FEXPath)) {
if (!std::filesystem::is_regular_file(FEXPath)) {
fmt::print(stderr, "Could not locate FEX executable\n");
std::abort();
}
}
const char* FEXArgs[] = {
FEXPath.c_str(),
BashPath,
"-c",
};
// Remove -c argument if arguments are empty
// Lets us start an emulated bash instance
const size_t FEXArgsCount = std::size(FEXArgs) - (EmptyArgs ? 1 : 0);
Argv.resize(ArgCount + FEXArgsCount);
// Pass in the FEX arguments
for (size_t i = 0; i < FEXArgsCount; ++i) {
Argv[i] = FEXArgs[i];
}
// Bring in passed in arguments
for (size_t i = 0; i < ArgCount; ++i) {
Argv[i + FEXArgsCount] = argv[i + 1];
std::vector<const char*> Argv;
Argv.emplace_back(FEXPath.c_str());
Argv.emplace_back(BashPath);
for (int i = 1; i < argc; ++i) {
Argv.emplace_back(argv[i]);
}
// Set --norc when no arguments are passed so PS1 doesn't get overwritten
const char* NoRC = "--norc";
if (EmptyArgs) {
Argv.emplace_back(NoRC);
Argv.emplace_back("--norc");
}
Argv.emplace_back(nullptr);
@@ -86,11 +97,8 @@ int main(int argc, char** argv, char** const envp) {
Envp.emplace_back(envp[i]);
}
}
std::string PS1 = "PS1=FEXBash-\\u@\\h:\\w> ";
if (PS1Env) {
PS1 += &PS1Env[strlen("PS1=")];
}
// Keep the string alive until after execve.
const std::string PS1 = EnchantedPS1(PS1Env);
Envp.emplace_back(PS1.c_str());
Envp.emplace_back(nullptr);
+1
View File
@@ -11,3 +11,4 @@ install(TARGETS FEXGetConfig RUNTIME
target_link_libraries(FEXGetConfig PRIVATE ${LIBS})
target_include_directories(FEXGetConfig PRIVATE ${CMAKE_BINARY_DIR}/generated)
target_compile_options(FEXGetConfig PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
+176 -45
View File
@@ -14,9 +14,10 @@
#include <filesystem>
#include <string>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <ucontext.h>
#ifdef ARCHITECTURE_arm64
namespace {
struct TSOEmulationFacts {
bool LSE {}, LSE2 {};
@@ -24,7 +25,6 @@ struct TSOEmulationFacts {
bool LRCPC1 {}, LRCPC2 {}, LRCPC3 {};
};
#ifdef ARCHITECTURE_arm64
bool CheckForHardwareTSO() {
// Check to see if this is supported.
auto Result = prctl(PR_GET_MEM_MODEL, 0, 0, 0, 0);
@@ -89,48 +89,77 @@ TSOEmulationFacts GetTSOEmulationFacts() {
.LRCPC3 = ((ISAR1 >> ISAR1_FIELDS::LRCPC) & IDFIELDMASK) >= 0b0011,
};
}
#else
TSOEmulationFacts GetTSOEmulationFacts() {
return {};
}
#endif
} // namespace
#ifdef ARCHITECTURE_arm64
namespace SIGBUSTest {
static bool* FaultArray {};
__attribute__((naked)) void atomic_store_u16(std::byte* Data, uint16_t Value) {
__attribute__((naked)) void atomic_load_u16(std::byte* Data) {
asm volatile(R"(
stlrh w1, [x0];
ldarh w1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u32(std::byte* Data) {
asm volatile(R"(
ldar w1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u64(std::byte* Data) {
asm volatile(R"(
ldar x1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u128(std::byte* Data) {
asm volatile(R"(
ldaxp x1, x2, [x0];
ret;
)" ::
: "x1", "x2", "x3", "memory");
}
__attribute__((naked)) void atomic_set_u16(std::byte* Data, uint16_t value) {
asm volatile(R"(
.word 0x78e13002; // ldsetalh w1, w2, [x0];
ret;
)" ::
: "memory");
}
__attribute__((naked)) void atomic_store_u32(std::byte* Data, uint32_t Value) {
__attribute__((naked)) void atomic_set_u32(std::byte* Data, uint32_t value) {
asm volatile(R"(
stlr w1, [x0];
.word 0xb8e13002; // ldsetal w1, w2, [x0];
ret;
)" ::
: "memory");
}
__attribute__((naked)) void atomic_store_u64(std::byte* Data, uint64_t Value) {
__attribute__((naked)) void atomic_set_u64(std::byte* Data, uint64_t value) {
asm volatile(R"(
stlr x1, [x0];
.word 0xf8e13002; // ldsetal x1, x2, [x0];
ret;
)" ::
: "memory");
}
__attribute__((naked)) void atomic_store_u128(std::byte* Data, uint64_t Value) {
__attribute__((naked)) void atomic_set_u128_impl(__uint128_t expected, __uint128_t desired, std::byte* Data) {
asm volatile(R"(
stlxp w3, x1, x1, [x0];
.word 0x4860fc82; // caspal x0, x1, x2, x3, [x4];
ret;
)" ::
: "memory");
}
static inline void atomic_set_u128(std::byte* Data, __uint128_t value) {
atomic_set_u128_impl(*reinterpret_cast<__uint128_t*>(Data), value, Data);
}
static void HandleSIGBUS(int, siginfo_t* info, void* context) {
FaultArray[reinterpret_cast<uintptr_t>(info->si_addr) & 63] = true;
@@ -140,7 +169,22 @@ static void HandleSIGBUS(int, siginfo_t* info, void* context) {
mcontext->pc += 4;
}
void TestSIGBUS() {
static bool CalculatedFaultOffsets {};
static bool FaultOffset_16bit[64] {};
static bool FaultOffset_32bit[64] {};
static bool FaultOffset_64bit[64] {};
static bool FaultOffset_128bit[64] {};
static bool FaultOffset_RMW_16bit[64] {};
static bool FaultOffset_RMW_32bit[64] {};
static bool FaultOffset_RMW_64bit[64] {};
static bool FaultOffset_RMW_128bit[64] {};
void RunFaultTests() {
if (CalculatedFaultOffsets) {
return;
}
struct sigaction act {};
act.sa_sigaction = HandleSIGBUS;
act.sa_flags = SA_SIGINFO;
@@ -148,12 +192,42 @@ void TestSIGBUS() {
auto ptr = reinterpret_cast<std::byte*>(mmap(nullptr, 4096, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
auto test_fault = [](bool* FaultOffsets, auto AccessFunction, std::byte* AccessArray) {
FaultArray = FaultOffsets;
for (size_t i = 0; i < 64; ++i) {
AccessFunction(AccessArray + i);
}
};
auto test_rmw_fault = [](bool* FaultOffsets, auto AccessFunction, std::byte* AccessArray) {
FaultArray = FaultOffsets;
for (size_t i = 0; i < 64; ++i) {
AccessFunction(AccessArray + i, 1);
}
};
test_fault(FaultOffset_16bit, atomic_load_u16, ptr);
test_fault(FaultOffset_32bit, atomic_load_u32, ptr);
test_fault(FaultOffset_64bit, atomic_load_u64, ptr);
test_fault(FaultOffset_128bit, atomic_load_u128, ptr);
auto TSOFacts = GetTSOEmulationFacts();
if (TSOFacts.LSE) {
test_rmw_fault(FaultOffset_RMW_16bit, atomic_set_u16, ptr);
test_rmw_fault(FaultOffset_RMW_32bit, atomic_set_u32, ptr);
test_rmw_fault(FaultOffset_RMW_64bit, atomic_set_u64, ptr);
test_rmw_fault(FaultOffset_RMW_128bit, atomic_set_u128, ptr);
}
munmap(ptr, 4096);
sigaction(SIGBUS, &act, nullptr);
CalculatedFaultOffsets = true;
}
void PrintSIGBUSInfo() {
RunFaultTests();
auto print_granule = [](const char* size, bool* FaultArray) {
std::string output {};
for (size_t i = 0; i < 64; ++i) {
@@ -171,25 +245,54 @@ void TestSIGBUS() {
fprintf(stdout, "%s: %s\n", size, output.c_str());
};
bool FaultOffset_16bit[64] {};
bool FaultOffset_32bit[64] {};
bool FaultOffset_64bit[64] {};
bool FaultOffset_128bit[64] {};
auto TSOFacts = GetTSOEmulationFacts();
const bool RMWIsDifferent = TSOFacts.LSE && (memcmp(FaultOffset_16bit, FaultOffset_RMW_16bit, sizeof(FaultOffset_16bit)) != 0 ||
memcmp(FaultOffset_32bit, FaultOffset_RMW_32bit, sizeof(FaultOffset_32bit)) != 0 ||
memcmp(FaultOffset_64bit, FaultOffset_RMW_64bit, sizeof(FaultOffset_64bit)) != 0 ||
memcmp(FaultOffset_128bit, FaultOffset_RMW_128bit, sizeof(FaultOffset_128bit)) != 0);
test_fault(FaultOffset_16bit, atomic_store_u16, ptr);
test_fault(FaultOffset_32bit, atomic_store_u32, ptr);
test_fault(FaultOffset_64bit, atomic_store_u64, ptr);
test_fault(FaultOffset_128bit, atomic_store_u128, ptr);
munmap(ptr, 4096);
sigaction(SIGBUS, &act, nullptr);
fprintf(stdout, "Fault Granularity: Split every 16 bytes\n");
if (!RMWIsDifferent) {
fprintf(stdout, "Fault Granularity: Split every 16 bytes\n");
} else {
fprintf(stdout, "Load/Store Fault Granularity: Split every 16 bytes\n");
}
print_granule(" 16-bit", FaultOffset_16bit);
print_granule(" 32-bit", FaultOffset_32bit);
print_granule(" 64-bit", FaultOffset_64bit);
print_granule("128-bit", FaultOffset_128bit);
if (RMWIsDifferent) {
fprintf(stdout, "RMW Atomic Fault Granularity: Split every 16 bytes\n");
print_granule(" 16-bit", FaultOffset_RMW_16bit);
print_granule(" 32-bit", FaultOffset_RMW_32bit);
print_granule(" 64-bit", FaultOffset_RMW_64bit);
print_granule("128-bit", FaultOffset_RMW_128bit);
}
}
struct FirstFaultInformation {
int32_t LoadStoreFaultAlignment {};
int32_t RMWFaultAlignment {};
};
FirstFaultInformation CalculateFirstFaultInformation() {
RunFaultTests();
FirstFaultInformation Info {};
auto FindFirstFaultOffset = [](bool FaultOffsets[64]) -> int32_t {
for (int32_t i = 0; i < 64; ++i) {
if (FaultOffsets[i]) {
return i;
}
}
return -1;
};
Info.LoadStoreFaultAlignment = FindFirstFaultOffset(FaultOffset_16bit) + 1;
Info.RMWFaultAlignment = FindFirstFaultOffset(FaultOffset_RMW_16bit) + 1;
return Info;
}
} // namespace SIGBUSTest
#endif
@@ -210,9 +313,8 @@ int main(int argc, char** argv, char** envp) {
Parser.add_option("--current-rootfs").action("store_true").help("Print the directory that contains the FEX rootfs. Mounted in the case of squashfs");
Parser.add_option("--tso-emulation-info").action("store_true").help("Print how FEX is emulating the x86-TSO memory model.");
#ifdef ARCHITECTURE_arm64
Parser.add_option("--tso-emulation-info").action("store_true").help("Print how FEX is emulating the x86-TSO memory model.");
Parser.add_option("--test-fault-granularity").action("store_true").help("Show SIGBUS fault granularity");
Parser.add_option("--identification-reg-info").action("store_true").help("Print identification registers");
#endif
@@ -248,7 +350,7 @@ int main(int argc, char** argv, char** envp) {
#ifdef ARCHITECTURE_arm64
if (Options.is_set_by_user("test_fault_granularity")) {
SIGBUSTest::TestSIGBUS();
SIGBUSTest::PrintSIGBUSInfo();
}
#endif
@@ -272,12 +374,23 @@ int main(int argc, char** argv, char** envp) {
}
}
#ifdef ARCHITECTURE_arm64
if (Options.is_set_by_user("tso_emulation_info")) {
auto TSOFacts = GetTSOEmulationFacts();
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(HalfBarrierTSOEnabled, HALFBARRIERTSOENABLED);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
const char* GPRMemoryTSOEmulation {};
const char* MemcpyMemoryTSOEmulation {};
const char* VectorMemoryTSOEmulation {};
const char* UnalignedMemoryLoadStoreTSOEmulation {};
const char* SplitLock16BEmulationType {};
const char* SplitLock16BConfigurationType {};
std::string UnalignedMemoryLoadStoreAlignmentGranularity {};
std::string UnalignedRMWAlignmentGranularity {};
if (TSOFacts.HardwareTSO) {
GPRMemoryTSOEmulation = "\e[32mHardware TSO\e[0m";
@@ -314,34 +427,52 @@ int main(int argc, char** argv, char** envp) {
UnalignedMemoryLoadStoreTSOEmulation = "\e[31mHalf-Barriers\e[0m";
}
const auto FFInfo = SIGBUSTest::CalculateFirstFaultInformation();
if (FFInfo.RMWFaultAlignment >= 64) {
SplitLock16BEmulationType = "\e[32mHardware cacheline unaligned atomics\e[0m";
SplitLock16BConfigurationType = "\e[32mTear-free\e[0m";
} else {
SplitLock16BEmulationType = TSOFacts.LSE ? "\e[31mTearing CAS loops\e[0m" : "\e[31mTearing LL/SC loops\e[0m";
SplitLock16BConfigurationType = StrictInProcessSplitLocks() ? "In-process mutex" : "Tearing";
}
if (FFInfo.LoadStoreFaultAlignment != 1) {
UnalignedMemoryLoadStoreAlignmentGranularity = fmt::format("\e[32m{}-byte\e[0m", FFInfo.LoadStoreFaultAlignment);
} else {
UnalignedMemoryLoadStoreAlignmentGranularity = TSOFacts.LSE2 ? "\e[32m16-byte\e[0m" : "\e[31mNatural alignment\e[0m";
}
if (FFInfo.LoadStoreFaultAlignment != FFInfo.RMWFaultAlignment) {
if (FFInfo.LoadStoreFaultAlignment != 1) {
UnalignedRMWAlignmentGranularity = fmt::format("\e[32m{}-byte\e[0m", FFInfo.RMWFaultAlignment);
} else {
UnalignedRMWAlignmentGranularity = TSOFacts.LSE2 ? "\e[32m16-byte\e[0m" : "\e[31mNatural alignment\e[0m";
}
}
fprintf(stdout, "Hardware Features:\n");
fprintf(stdout, "\tMemory atomics emulation method: %s\n", TSOFacts.LSE ? "\e[32mLSE\e[0m" : "\e[31mLL/SC\e[0m");
fprintf(stdout, "\tUnaligned atomic memory granularity: %s\n", TSOFacts.LSE2 ? "\e[32m16-byte\e[0m" : "\e[31mNatural alignment\e[0m");
///< TODO: Once TME is supported by hardware this can change.
fprintf(stdout, "\tUnaligned atomic memory granularity: %s\n", UnalignedMemoryLoadStoreAlignmentGranularity.c_str());
if (FFInfo.LoadStoreFaultAlignment != FFInfo.RMWFaultAlignment) {
fprintf(stdout, "\tUnaligned atomic RMW granularity: %s\n", UnalignedRMWAlignmentGranularity.c_str());
}
fprintf(stdout, "\tUnaligned memory loadstore emulation: %s\n", UnalignedMemoryLoadStoreTSOEmulation);
fprintf(stdout, "\t16-Byte split-lock atomic emulation: %s\n", TSOFacts.LSE ? "\e[31mTearing CAS loops\e[0m" : "\e[31mTearing LL/SC loops\e[0m");
fprintf(stdout, "\t16-Byte split-lock atomic emulation: %s\n", SplitLock16BEmulationType);
fprintf(stdout, "\t64-Byte split-lock atomic emulation: %s\n", TSOFacts.LSE ? "\e[31mTearing CAS loops\e[0m" : "\e[31mTearing LL/SC loops\e[0m");
fprintf(stdout, "\tGPR memory model emulation: %s\n", GPRMemoryTSOEmulation);
fprintf(stdout, "\tMemcpy memory model emulation: %s\n", MemcpyMemoryTSOEmulation);
fprintf(stdout, "\tVector memory model emulation: %s\n", VectorMemoryTSOEmulation);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(HalfBarrierTSOEnabled, HALFBARRIERTSOENABLED);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
fprintf(stderr, "Strict: %d\n", StrictInProcessSplitLocks());
fprintf(stdout, "\nConfiguration:\n");
fprintf(stdout, "\tTSO Emulation: %s\n", TSOEnabled() ? "Enabled" : "Disabled");
fprintf(stdout, "\tMemcpy TSO Emulation: %s\n", TSOEnabled() && MemcpySetTSOEnabled() ? "Enabled" : "Disabled");
fprintf(stdout, "\tVector TSO Emulation: %s\n", TSOEnabled() && VectorTSOEnabled() ? "Enabled" : "Disabled");
fprintf(stdout, "\tHalf-barrier unaligned TSO emulation: %s\n", TSOEnabled() && HalfBarrierTSOEnabled() ? "Enabled" : "Disabled");
fprintf(stdout, "\t16-Byte strict split-lock emulation: %s\n", StrictInProcessSplitLocks() ? "In-process mutex" : "Tearing");
fprintf(stdout, "\t16-Byte strict split-lock emulation: %s\n", SplitLock16BConfigurationType);
fprintf(stdout, "\t64-Byte strict split-lock emulation: %s\n", StrictInProcessSplitLocks() ? "In-process mutex" : "Tearing");
}
#ifdef ARCHITECTURE_arm64
if (Options.is_set_by_user("identification_reg_info")) {
auto Features = FEX::GetCPUFeaturesFromIDRegisters();
fextl::string features {};
+30 -14
View File
@@ -32,7 +32,6 @@
#include <sys/personality.h>
#include <sys/prctl.h>
#include <sys/random.h>
#include <linux/prctl.h>
#define PAGE_START(x) ((x) & ~(uintptr_t)(4095))
#define PAGE_OFFSET(x) ((x) & 4095)
@@ -646,6 +645,11 @@ public:
if (Is64BitMode()) {
AuxVariables.emplace_back(auxv_t {4, 0x38}); // AT_PHENT
// 64-bit vsyscall entry points are hardcoded to a single page at 0xffffffffff600000.
// FEX can't actually map anything there so it is a hardcoded quirk, similar to how the kernel traps these executions.
// Just track it as a mapped anonymous executable page.
Handler->AddVirtualPage(Thread, 0xFFFFFFFFFF600000ULL, FEXCore::Utils::FEX_PAGE_SIZE, PROT_READ | PROT_EXEC);
} else {
AuxVariables.emplace_back(auxv_t {4, 0x20}); // AT_PHENT
@@ -767,26 +771,38 @@ public:
// Ensure we don't read past the end into garbage data
stat_buffer[std::clamp(bytes_read, 0L, static_cast<ssize_t>(sizeof(stat_buffer)) - 1)] = '\0';
uint64_t start_code, end_code, start_stack, start_data, end_data, start_brk, arg_start, arg_end, env_start, env_end;
// See man proc_pid_stat
int items_read = sscanf(stat_buffer,
"%*d %*s %*c %*d %*d " // 1 to 5
"%*d %*d %*d %*u %*u " // 6 to 10
"%*u %*u %*u %*u %*u " // 11 to 15
"%*d %*d %*d %*d %*d " // 16 to 20
"%*d %*u %*u %*d %*u " // 21 to 25
"%llu %llu %llu %*u %*u " // 26 to 30
"%*u %*u %*u %*u %*u " // 31 to 35
"%*u %*u %*d %*d %*u " // 36 to 40
"%*u %*u %*u %*d %llu " // 40 to 45
"%llu %llu %llu %llu %llu " // 46 to 50
"%llu", // 51
&map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
&map.arg_start, &map.arg_end, &map.env_start, &map.env_end);
"%*d %*s %*c %*d %*d " // 1 to 5
"%*d %*d %*d %*u %*u " // 6 to 10
"%*u %*u %*u %*u %*u " // 11 to 15
"%*d %*d %*d %*d %*d " // 16 to 20
"%*d %*u %*u %*d %*u " // 21 to 25
"%lu %lu %lu %*u %*u " // 26 to 30
"%*u %*u %*u %*u %*u " // 31 to 35
"%*u %*u %*d %*d %*u " // 36 to 40
"%*u %*u %*u %*d %lu " // 40 to 45
"%lu %lu %lu %lu %lu " // 46 to 50
"%lu", // 51
&start_code, &end_code, &start_stack, &start_data, &end_data, &start_brk, &arg_start, &arg_end, &env_start, &env_end);
if (items_read != 10) {
return false;
}
map.start_code = start_code;
map.end_code = end_code;
map.start_stack = start_stack;
map.start_data = start_data;
map.end_data = end_data;
map.start_brk = start_brk;
map.arg_start = arg_start;
map.arg_end = arg_end;
map.env_start = env_start;
map.env_end = env_end;
map.brk = reinterpret_cast<uint64_t>(sbrk(0));
// The kernel will leave these values unchanged, see implementation in sys.c
@@ -63,7 +63,7 @@ $end_info$
#include <utility>
#include <sys/sysinfo.h>
#include <sys/signal.h>
#include <signal.h>
namespace FEX::Logging {
static bool SilentLog {};
@@ -597,9 +597,14 @@ int main(int argc, char** argv, char** const envp) {
SyscallHandler->DefaultProgramBreak(BRKInfo.Base, BRKInfo.Size);
// Request code cache generation
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Request code cache generation
FEXServerClient::PopulateCodeCache(FEXServerClient::GetServerFD(), Loader.GetMainElfFD(), FEXCore::Config::Get_MULTIBLOCK());
if (VDSOMapping) {
// Finalize code cache for libVDSO-guest.so. This needs to be done explicitly since VDSO doesn't use LoadLib.
SyscallHandler->TriggerGuestLibWrapperCodeCacheLoad(*ParentThread->Thread, reinterpret_cast<uintptr_t>(VDSOMapping.VDSOBase));
}
}
// Pull RIP and stack pointer from loader and set the thread data to it.
+12 -2
View File
@@ -6,10 +6,20 @@ target_link_libraries(FEXOfflineCompiler PRIVATE
cpp-optparse
FEXCore
JemallocLibs
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt)
if (MINGW)
patch_library_wine(FEXOfflineCompiler)
target_include_directories(FEXOfflineCompiler PRIVATE
"${CMAKE_SOURCE_DIR}/Source/Windows/include/"
"${CMAKE_SOURCE_DIR}/Source/"
"${CMAKE_SOURCE_DIR}/Source/Windows/"
)
target_link_libraries(FEXOfflineCompiler PRIVATE CommonWindows ntdll_ex)
else()
target_link_libraries(FEXOfflineCompiler PRIVATE ${PTHREAD_LIB} LinuxEmulation)
endif()
LinkerGC(FEXOfflineCompiler)
install(TARGETS FEXOfflineCompiler RUNTIME
+557 -33
View File
@@ -1,13 +1,20 @@
// SPDX-License-Identifier: MIT
#ifndef _WIN32
#include "../FEXInterpreter/ELFCodeLoader.h"
#endif
#include <DummyHandlers.h>
#ifndef _WIN32
#include <PortabilityInfo.h>
#include <Thunks.h>
#endif
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeCache.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <Common/ArgumentLoader.h>
#include <Common/Config.h>
@@ -16,14 +23,42 @@
#include <OptionParser.h>
#ifndef _WIN32
#include <elf.h>
#endif
#include <fmt/printf.h>
#include <libgen.h>
#include <fcntl.h>
#include <fstream>
#include <optional>
#include <ranges>
#ifndef _WIN32
#include <sys/mman.h>
#else
#include <Common/CPUFeatures.h>
#include <Common/Handle.h>
#include <Common/ImageTracker.h>
#include <Common/InvalidationTracker.h>
#include <Common/JITGuardPage.h>
#include <Common/Logging.h>
#include <Common/Module.h>
#include <Common/OvercommitTracker.h>
#include <Common/PortabilityInfo.h>
static std::unique_ptr<FEX::Windows::OvercommitTracker> OvercommitTracker;
#endif
static FEXCore::Core::InternalThreadState* Thread = nullptr;
#ifdef _WIN32
class AOTSyscallHandler : public FEXCore::HLE::SyscallHandler {
#else
class AOTSyscallHandler : public FEXCore::HLE::SyscallHandler, public FEX::HLE::SyscallMmapInterface {
#endif
public:
AOTSyscallHandler(FEXCore::HLE::SyscallOSABI SyscallOSABI) {
AOTSyscallHandler(FEXCore::Context::Context& CTX, FEXCore::HLE::SyscallOSABI SyscallOSABI)
: CTX(CTX) {
OSABI = SyscallOSABI;
}
@@ -32,24 +67,39 @@ public:
return 0;
}
FEXCore::Context::Context& CTX;
#ifdef _WIN32
FEX::Windows::ImageTracker ImageTracker {CTX, true};
const std::unordered_map<DWORD, FEXCore::Core::InternalThreadState*> ThreadsUnused;
FEX::Windows::InvalidationTracker InvalidationTracker {CTX, ThreadsUnused};
#else
FEXCore::ExecutableFileInfo FileInfo;
std::map<uint64_t, uint64_t> FileRanges;
#endif
uintptr_t VAFileStart = 0;
// These are no-ops implementations of the SyscallHandler API
std::optional<FEXCore::ExecutableFileSectionInfo> LookupExecutableFileSection(FEXCore::Core::InternalThreadState*, uint64_t Address) override {
#ifndef _WIN32
auto It = FileRanges.upper_bound(Address - VAFileStart);
LOGMAN_THROW_A_FMT(It != FileRanges.begin(), "Could not find associated file mapping");
--It;
LOGMAN_THROW_A_FMT(VAFileStart + It->first + It->second > Address, "Could not find associated file mapping for {:#x}", Address);
return FEXCore::ExecutableFileSectionInfo {FileInfo, VAFileStart, VAFileStart + It->first, VAFileStart + It->first + It->second};
#else
return ImageTracker.LookupExecutableFileSection(Address);
#endif
}
FEXCore::HLE::ExecutableRangeInfo QueryGuestExecutableRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Address) override {
#ifndef _WIN32
return {0, UINT64_MAX, true};
#else
return InvalidationTracker.QueryExecutableRange(Address);
#endif
}
#ifndef _WIN32
void* GuestMmap(FEXCore::Core::InternalThreadState*, void* addr, size_t Size, int prot, int Flags, int fd, off_t offset) override {
// Force writeable to allow applying relocations
auto Ret = mmap(addr, Size, prot | PROT_WRITE, Flags, fd, offset);
@@ -63,6 +113,19 @@ public:
uint64_t GuestMunmap(FEXCore::Core::InternalThreadState*, void* addr, uint64_t length) override {
return munmap(addr, length);
}
void AddVirtualPage(FEXCore::Core::InternalThreadState* Thread, uint64_t addr, size_t length, int prot) override {
LogMan::Msg::AFmt("Can't Track mmap through here");
FEX_UNREACHABLE;
}
#else
void MarkOvercommitRange(uint64_t Start, uint64_t Length) override {
OvercommitTracker->MarkRange(Start, Length);
}
void UnmarkOvercommitRange(uint64_t Start, uint64_t Length) override {
OvercommitTracker->UnmarkRange(Start, Length);
}
#endif
};
static void MsgHandler(LogMan::DebugLevels Level, const char* Message) {
@@ -86,8 +149,18 @@ struct std::hash<FEXCore::ExecutableFileInfo> {
}
};
// Windows requires O_BINARY, whereas on Linux it's implicit
#ifndef O_BINARY
#define O_BINARY 0
#endif
// Placeholder data to ensure the compile thread doesn't de-reference nullptr data
static FEXCore::Core::CPUState::gdt_segment gdt[32] {};
#if !defined(_WIN32) || defined(_M_ARM64EC)
static constexpr size_t DefaultCS {FEXCore::Core::CPUState::DEFAULT_USER_CS};
#else
static constexpr size_t DefaultCS {4};
#endif
static FEXCore::Core::InternalThreadState* SetupCompileThread(FEXCore::Context::Context& CTX, bool Is64Bit) {
auto Thread = CTX.CreateThread(0, 0);
@@ -95,13 +168,11 @@ static FEXCore::Core::InternalThreadState* SetupCompileThread(FEXCore::Context::
auto Frame = Thread->CurrentFrame;
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_GDT] = &gdt[0];
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_LDT] = &gdt[0];
Frame->State.cs_idx = FEXCore::Core::CPUState::DEFAULT_USER_CS << 3;
Frame->State.cs_idx = DefaultCS << 3;
auto GDT = FEXCore::Core::CPUState::GetSegmentFromIndex(Frame->State, Frame->State.cs_idx);
FEXCore::Core::CPUState::SetGDTBase(GDT, 0);
FEXCore::Core::CPUState::SetGDTLimit(GDT, 0xFFFFFU);
Frame->State.cs_cached =
FEXCore::Core::CPUState::CalculateGDTBase(*FEXCore::Core::CPUState::GetSegmentFromIndex(Frame->State, Frame->State.cs_idx));
Frame->State.cs_cached = FEXCore::Core::CPUState::CalculateGDTBase(*GDT);
if (Is64Bit) {
GDT->L = 1; // L = Long Mode = 64-bit
@@ -114,20 +185,256 @@ static FEXCore::Core::InternalThreadState* SetupCompileThread(FEXCore::Context::
return Thread;
}
// Returns filename of generated cache on success
static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInfo& Binary, fextl::set<uintptr_t> BlockList, std::string_view OutDir) {
uint64_t CodeCacheConfigId = 0; // TODO: Make unique to active configuration
#ifdef _WIN32
static bool RelocateMappedImage(HMODULE Module) {
const auto* NtHeaders = reinterpret_cast<FEX::Windows::ArchImageNtHeaders*>(RtlImageNtHeader(Module));
if (!NtHeaders) {
return false;
}
const auto BaseAddress = reinterpret_cast<uintptr_t>(Module);
const auto PreferredBase = NtHeaders->OptionalHeader.ImageBase;
const auto Delta = static_cast<intptr_t>(BaseAddress - PreferredBase);
// Wine will automatically relocate all DLLs to their mapped address, but PE relocations must still be applied so
// FEXCore can correctly transform them into FEX relocations
if (Delta == 0) {
return true;
}
ULONG RelocSize = 0;
auto* RelocBlock =
reinterpret_cast<IMAGE_BASE_RELOCATION*>(RtlImageDirectoryEntryToData(Module, true, IMAGE_DIRECTORY_ENTRY_BASERELOC, &RelocSize));
if (!RelocBlock || RelocSize == 0) {
return true;
}
// Reprotect all sections as RW to apply relocations, saving their prior protections
struct SectionPatchState {
void* Address;
SIZE_T Size;
DWORD PreviousProtection;
};
std::vector<SectionPatchState> SectionStates;
SectionStates.reserve(NtHeaders->FileHeader.NumberOfSections);
auto* SectionHeader = IMAGE_FIRST_SECTION(NtHeaders);
const auto* SectionHeaderEnd = SectionHeader + NtHeaders->FileHeader.NumberOfSections;
for (; SectionHeader != SectionHeaderEnd; ++SectionHeader) {
if (SectionHeader->SizeOfRawData == 0) {
continue;
}
const auto SecAddr = reinterpret_cast<void*>(BaseAddress + SectionHeader->VirtualAddress);
const SIZE_T SecSize = SectionHeader->Misc.VirtualSize;
DWORD OldProt = 0;
if (!VirtualProtect(SecAddr, SecSize, PAGE_READWRITE, &OldProt)) {
for (const auto& State : SectionStates) {
DWORD Ignored;
VirtualProtect(State.Address, State.Size, State.PreviousProtection, &Ignored);
}
return false;
}
SectionStates.push_back({SecAddr, SecSize, OldProt});
}
// Apply relocations to all sections
bool RelocSuccess = true;
const uintptr_t RelocEnd = reinterpret_cast<uintptr_t>(RelocBlock) + RelocSize;
const uint32_t ImageSize = NtHeaders->OptionalHeader.SizeOfImage;
while (reinterpret_cast<uintptr_t>(RelocBlock) < RelocEnd && RelocBlock->SizeOfBlock) {
if (RelocBlock->VirtualAddress >= ImageSize) {
RelocSuccess = false;
break;
}
const auto Count = (RelocBlock->SizeOfBlock - sizeof(IMAGE_BASE_RELOCATION)) / sizeof(USHORT);
const auto PageAddress = BaseAddress + RelocBlock->VirtualAddress;
RelocBlock = LdrProcessRelocationBlock(PageAddress, Count, reinterpret_cast<USHORT*>(RelocBlock + 1), Delta);
if (!RelocBlock) {
RelocSuccess = false;
break;
}
}
// Restore sections to previous protection states
for (const auto& State : SectionStates) {
DWORD Ignored;
VirtualProtect(State.Address, State.Size, State.PreviousProtection, &Ignored);
}
if (!RelocSuccess) {
return false;
}
LogMan::Msg::IFmt("Relocated image {:X} -> {:X}", PreferredBase, BaseAddress);
return true;
}
#ifdef ARCHITECTURE_arm64ec
static void* MapView(HANDLE SectionHandle) {
return MapViewOfFile(SectionHandle, FILE_MAP_EXECUTE | FILE_MAP_READ, 0, 0, 0);
}
#else
static void* MapView(HANDLE SectionHandle) {
void* BaseAddress = nullptr;
SIZE_T ViewSize = 0;
LARGE_INTEGER Offset {};
// Map images in the lower 32-bits for WOW64 so relocations can be correctly applied
const ULONG_PTR ZeroBits = 0x7fffffff;
NTSTATUS Status =
NtMapViewOfSection(SectionHandle, GetCurrentProcess(), &BaseAddress, ZeroBits, 0, &Offset, &ViewSize, ViewShare, 0, PAGE_EXECUTE_READ);
if (Status < 0) {
return nullptr;
}
return BaseAddress;
}
#endif
// Returns the base address of the mapped image
static std::optional<uint64_t> TryMapImage(FEX::Windows::InvalidationTracker& InvalidationTracker, FEX::Windows::ImageTracker& ImageTracker,
const FEXCore::CodeMapFileId& ID, FEXCore::ExecutableFileInfo& Info) {
{
FEX::Windows::ScopedHandle File {CreateFileA(Info.Filename.c_str(), GENERIC_READ | SYNCHRONIZE, FILE_SHARE_READ | FILE_SHARE_DELETE,
nullptr, OPEN_EXISTING, FILE_ATTRIBUTE_NORMAL, nullptr)};
if (!File) {
LogMan::Msg::EFmt("Couldn't find image: {}", Info.Filename);
return std::nullopt;
}
FEX::Windows::ScopedHandle Section {CreateFileMappingA(*File, nullptr, SEC_IMAGE | PAGE_EXECUTE_READ, 0, 0, nullptr)};
if (!Section) {
LogMan::Msg::EFmt("Couldn't create section for image: {}", Info.Filename);
return std::nullopt;
}
void* Mapping = MapView(*Section);
if (!Mapping) {
LogMan::Msg::EFmt("Couldn't map section for image: {}", Info.Filename);
return std::nullopt;
}
if (!RelocateMappedImage(reinterpret_cast<HMODULE>(Mapping))) {
LogMan::Msg::EFmt("Failed to apply image relocations");
UnmapViewOfFile(Mapping);
return std::nullopt;
}
uint64_t BaseAddress = reinterpret_cast<uint64_t>(Mapping);
LogMan::Msg::IFmt("Mapped image: {} @ {:X}", Info.Filename, BaseAddress);
InvalidationTracker.HandleImageMap(FEX::Windows::BaseName(Info.Filename), BaseAddress);
ImageTracker.HandleImageMap(Info.Filename, BaseAddress, false /* unused during cache generation */);
return BaseAddress;
}
}
static LONG ExceptionHandler(_EXCEPTION_POINTERS* ExceptionInfo) {
if (ExceptionInfo->ExceptionRecord->ExceptionCode == EXCEPTION_ACCESS_VIOLATION) {
const auto FaultAddress = static_cast<uint64_t>(ExceptionInfo->ExceptionRecord->ExceptionInformation[1]);
if (OvercommitTracker->HandleAccessViolation(FaultAddress)) {
return EXCEPTION_CONTINUE_EXECUTION;
}
#ifdef ARCHITECTURE_arm64ec
ARM64_NT_CONTEXT ArmContext {};
auto* Context = &ArmContext;
#else
auto* Context = ExceptionInfo->ContextRecord;
#endif
if (FEX::Windows::JITGuardPage::HandleJITGuardPage(Thread, reinterpret_cast<void*>(FaultAddress), Context->X,
reinterpret_cast<__uint128_t*>(Context->V), &Context->Pc)) {
#ifdef ARCHITECTURE_arm64ec
auto* ECContext = reinterpret_cast<ARM64EC_NT_CONTEXT*>(ExceptionInfo->ContextRecord);
ECContext->X0 = Context->X0;
ECContext->X19 = Context->X19;
ECContext->X20 = Context->X20;
ECContext->X21 = Context->X21;
ECContext->X22 = Context->X22;
ECContext->X25 = Context->X25;
ECContext->X26 = Context->X26;
ECContext->X27 = Context->X27;
ECContext->Fp = Context->Fp;
ECContext->Lr = Context->Lr;
ECContext->Sp = Context->Sp;
ECContext->Pc = Context->Pc;
for (size_t i = 0; i < 8; ++i) {
memcpy(&reinterpret_cast<__uint128_t*>(ECContext->V)[8 + i], &reinterpret_cast<__uint128_t*>(Context->V)[8 + i], sizeof(uint64_t));
}
#endif
return EXCEPTION_CONTINUE_EXECUTION;
}
}
return EXCEPTION_CONTINUE_SEARCH;
}
struct winsize {
int ws_col;
};
#endif
// Returns filename of generated cache on success
static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInfo& Binary, uint64_t CodeCacheConfigId,
fextl::set<uintptr_t> BlockList, std::string_view OutDir) {
#ifndef _WIN32
ELFCodeLoader Loader(Binary.Filename.c_str(), -1, "", fextl::vector<fextl::string> {Binary.Filename.c_str()},
fextl::vector<fextl::string> {}, nullptr, nullptr, true /* skip interpreter */);
if (!Loader.ELFWasLoaded()) {
fmt::print("Invalid or unsupported ELF file.\n");
return std::nullopt;
}
const bool Is64Bit = Loader.Is64BitMode();
#elif defined(_M_ARM64EC)
const bool Is64Bit = true;
#else
const bool Is64Bit = false;
#endif
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Is64Bit ? "1" : "0");
// Load HostFeatures
#ifndef _WIN32
auto HostFeatures = FEX::FetchHostFeatures();
#else
const auto NtDll = GetModuleHandle("ntdll.dll");
const bool IsWine = !!GetProcAddress(NtDll, "wine_get_version");
auto HostFeatures = FEX::Windows::CPUFeatures::FetchHostFeatures(
IsWine, Is64Bit ? FEXCore::HostFeatures::HostTypeEnum::Arm64ec : FEXCore::HostFeatures::HostTypeEnum::Wow64);
#endif
auto CTX = FEXCore::Context::Context::CreateNewContext(HostFeatures);
CTX->GetCodeCache().InitiateCacheGeneration();
#ifdef _WIN32
OvercommitTracker = std::make_unique<FEX::Windows::OvercommitTracker>(IsWine);
auto SyscallOSABI = Is64Bit ? FEXCore::HLE::SyscallOSABI::OS_LINUX64 : FEXCore::HLE::SyscallOSABI::OS_LINUX32;
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(SyscallOSABI);
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(*CTX, SyscallOSABI);
SyscallHandler->VAFileStart =
TryMapImage(SyscallHandler->InvalidationTracker, SyscallHandler->ImageTracker, Binary.FileId, Binary).value_or(0);
if (!SyscallHandler->VAFileStart) {
return std::nullopt;
}
// Register exception handler for OvercommitTracker
AddVectoredExceptionHandler(1, ExceptionHandler);
#else
Loader.CalculateHWCaps(CTX.get());
auto SyscallOSABI = Is64Bit ? FEXCore::HLE::SyscallOSABI::OS_LINUX64 : FEXCore::HLE::SyscallOSABI::OS_LINUX32;
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(*CTX, SyscallOSABI);
// Populate relocations from ELF file
{
@@ -137,10 +444,12 @@ static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInf
SyscallHandler->FileInfo.Relocations = Binary.Relocations;
}
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Is64Bit ? "1" : "0");
// Load HostFeatures
auto HostFeatures = FEX::FetchHostFeatures();
if (!Is64Bit) {
const auto PageSize = sysconf(_SC_PAGESIZE);
// Block upper address space
FEXCore::Allocator::SetupHooks(PageSize > 0 ? PageSize : FEXCore::Utils::FEX_PAGE_SIZE);
}
#endif
if (!std::filesystem::exists(Binary.Filename)) {
fmt::print("File {} does not exist\n", Binary.Filename);
@@ -148,28 +457,21 @@ static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInf
return /*EXIT_FAILURE*/ std::nullopt;
}
auto CTX = FEXCore::Context::Context::CreateNewContext(HostFeatures);
Loader.CalculateHWCaps(CTX.get());
auto SignalDelegation = std::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
CTX->SetSignalDelegator(SignalDelegation.get());
CTX->SetSyscallHandler(SyscallHandler.get());
#ifndef _WIN32
auto ThunkHandler = FEX::HLE::CreateThunkHandler();
CTX->SetThunkHandler(ThunkHandler.get());
#endif
if (!CTX->InitCore()) {
return std::nullopt;
}
if (!Is64Bit) {
const auto PageSize = sysconf(_SC_PAGESIZE);
// Block upper address space
FEXCore::Allocator::SetupHooks(PageSize > 0 ? PageSize : FEXCore::Utils::FEX_PAGE_SIZE);
}
auto Thread = SetupCompileThread(*CTX, Is64Bit);
Thread = SetupCompileThread(*CTX, Is64Bit);
#ifndef _WIN32
{
auto ElfBase = Loader.LoadMainElfFile(nullptr, SyscallHandler.get(), Thread);
if (!ElfBase.has_value()) {
@@ -195,13 +497,21 @@ static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInf
}
}
}
CTX->GetCodeCache().InitiateCacheGeneration();
#endif
{
std::vector<std::unique_ptr<ELFCodeLoader>> LoaderMem;
// Refuse to continue if the block list contains any out-of-bounds blocks.
// This often indicates a corrupted code map.
{
auto [min_val, max_val] = std::ranges::minmax_element(BlockList, std::less {});
auto MinBound = SyscallHandler->LookupExecutableFileSection(Thread, *min_val + SyscallHandler->VAFileStart);
auto MaxBound = SyscallHandler->LookupExecutableFileSection(Thread, *max_val + SyscallHandler->VAFileStart);
LOGMAN_THROW_A_FMT(MinBound && MaxBound, "Cached blocks offsets {:#x}-{:#x} out of bounds for library {} ({:016x} @ {:#x})!",
*min_val, *max_val, Binary.Filename, Binary.FileId, SyscallHandler->VAFileStart);
}
fmt::print(stderr, "Compiling code...\n");
FEX_CONFIG_OPT(MaxInst, MAXINST);
for (auto Addr : BlockList) {
if (!CTX->CheckIfBlockIsCacheable(*Thread, Addr + SyscallHandler->VAFileStart, MaxInst)) {
@@ -213,13 +523,17 @@ static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInf
auto Filename = fmt::format("{}{}-{:016x}", OutDir, FEXCore::CodeMap::GetBaseFilename(Binary, false), CodeCacheConfigId);
auto FilenameNew = Filename + ".new";
int fd = open(FilenameNew.c_str(), O_CREAT | O_WRONLY, 0644);
int fd = open(FilenameNew.c_str(), O_CREAT | O_WRONLY | O_BINARY, 0644);
{
auto Entry = SyscallHandler->LookupExecutableFileSection(Thread, SyscallHandler->VAFileStart).value();
#ifndef _WIN32
CTX->GetCodeCache().SaveData(*Thread, fd, Entry, 0 /* TODO: Use static base address information if available */);
#else
CTX->GetCodeCache().SaveData(*Thread, fd, Entry, SyscallHandler->VAFileStart);
#endif
}
std::filesystem::rename(FilenameNew.c_str(), Filename.c_str());
close(fd);
std::filesystem::rename(FilenameNew.c_str(), Filename.c_str());
return Filename;
}
}
@@ -291,10 +605,11 @@ static int GenerateCache(int argc, const char** argv) {
const auto PortableInfo = FEX::ReadPortabilityInformation();
char* envp[] = {nullptr};
FEXCore::Config::Shutdown();
FEX::Config::LoadConfig("", envp, PortableInfo);
auto NumBlocks = Data.at(ProgramName).size();
auto GeneratedCache = GenerateSingleCache(ProgramName, Data.at(ProgramName), OutDir);
auto GeneratedCache = GenerateSingleCache(ProgramName, 0 /* TODO: Config id */, Data.at(ProgramName), OutDir);
if (GeneratedCache) {
fmt::print("Successfully populated cache {} ({} blocks) via {}\n\n", GeneratedCache.value(), NumBlocks,
std::filesystem::path {CodeMapPath}.filename().string());
@@ -302,20 +617,229 @@ static int GenerateCache(int argc, const char** argv) {
return GeneratedCache ? 0 : 1;
}
/**
* Writes aggregated code map data into a single code map file that is ready to be used for cache generation
*/
static void WriteNewCodeMap(const FEXCore::ExecutableFileInfo& File, const std::string& OutputName, const fextl::set<uintptr_t>& Blocks,
bool IsExecutable, const std::set<FEXCore::ExecutableFileInfo>& Dependencies) {
fmt::print("Writing {} blocks to {}\n", Blocks.size(), OutputName);
struct CodeMapOpener : FEXCore::CodeMapOpener {
CodeMapOpener(const std::string& Filename) {
FD = open(Filename.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_BINARY, 0644);
}
int OpenCodeMapFile() override {
return FD;
}
int FD;
};
CodeMapOpener CodeMapOpener(OutputName);
FEXCore::CodeMapWriter OutputCodeMap(CodeMapOpener, true);
if (IsExecutable) {
// List the main executable and all used libraries
OutputCodeMap.AppendSetMainExecutable(File);
for (auto& Dependency : Dependencies) {
OutputCodeMap.AppendLibraryLoad(Dependency);
}
} else {
// List only the library itself
OutputCodeMap.AppendLibraryLoad(File);
}
for (auto& Block : Blocks) {
OutputCodeMap.AppendBlock(FEXCore::ExecutableFileSectionInfo {File, 0}, Block);
}
}
struct ParsedContentsAndDependencies {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
std::set<FEXCore::CodeMapFileId> Dependencies;
};
/**
* Discovers any pending code maps, parses their contents into a runtime data structure, and deletes them
*/
static std::map<FEXCore::CodeMapFileId, ParsedContentsAndDependencies> ImportPendingCodeMaps(const std::string& NewCodeMapDirectory) {
// TODO: Handle nomb code maps
std::map<FEXCore::CodeMapFileId, ParsedContentsAndDependencies> Result;
for (auto& Entry : std::filesystem::directory_iterator(NewCodeMapDirectory)) {
if (!Entry.is_regular_file()) {
continue;
}
const auto Name = Entry.path().filename().string();
if (!Name.ends_with(".bin")) {
continue;
}
if (std::filesystem::file_size(Entry.path()) == 0) {
fmt::println("Found zero-size code map {}, deleting", Name);
std::filesystem::remove(Entry.path());
continue;
}
fmt::print("Importing new code map {}\n", Name);
std::ifstream Incoming(Entry.path(), std::ios_base::binary);
std::set<FEXCore::CodeMapFileId> Dependencies;
std::optional<FEXCore::CodeMapFileId> ExecutableFileId;
for (auto& [FileId, Contents] : FEXCore::CodeMap::ParseCodeMap(Incoming)) {
auto& [Filename, Blocks, IsExecutable, _] =
Result.emplace(std::piecewise_construct, std::forward_as_tuple(FileId), std::tuple {}).first->second;
Filename = std::move(Contents.Filename);
Blocks.merge(std::move(Contents.Blocks));
IsExecutable = Contents.IsExecutable;
if (IsExecutable) {
LOGMAN_THROW_A_FMT(!ExecutableFileId, "Expected a unique executable identifiers per code map");
ExecutableFileId = FileId;
} else {
Dependencies.insert(FileId);
}
}
// Every imported code map should have had exactly one executable marker
LOGMAN_THROW_A_FMT(ExecutableFileId, "Could not find an executable identifer in the code map");
Result.at(*ExecutableFileId).Dependencies = std::move(Dependencies);
// Delete imported code map
Incoming.close();
std::filesystem::remove(Entry.path());
}
return Result;
}
/**
* Checks and processes new code maps generated by FEX
*
* Processed code maps are merged into the reference ("ready") code maps
*/
static void AggregateCodeMaps(const std::string& NewCodeMapDirectory, const std::string& ReadyCodeMapDirectory) {
auto IncomingCodeMap = ImportPendingCodeMaps(NewCodeMapDirectory);
for (auto& [FileId, Contents] : IncomingCodeMap) {
// For each referenced binary, add the newly referenced offsets to that binary's reference code map
const FEXCore::ExecutableFileInfo File {nullptr, FileId, Contents.Filename};
const auto BinaryName = std::string {FEXCore::CodeMap::GetBaseFilename(File, false)};
auto OutputName = fmt::format("{}/{}", ReadyCodeMapDirectory, BinaryName);
if (auto ReferenceCodeMap = std::ifstream(OutputName, std::ios_base::binary)) {
auto PreviousBlocks = FEXCore::CodeMap::ParseCodeMap(ReferenceCodeMap).at(File.FileId).Blocks;
auto NumPreviousBlocks = PreviousBlocks.size();
Contents.Blocks.merge(std::move(PreviousBlocks));
if (Contents.Blocks.size() == NumPreviousBlocks) {
// No new blocks => skip updating
continue;
} else {
fmt::println(" Found {} new blocks ({} total) in code map {} for {}", Contents.Blocks.size() - NumPreviousBlocks,
Contents.Blocks.size(), BinaryName, File.Filename);
}
}
// Update code map
std::set<FEXCore::ExecutableFileInfo> Dependencies;
for (auto& Dependency : Contents.Dependencies) {
Dependencies.emplace(nullptr, Dependency, IncomingCodeMap.at(Dependency).Filename);
}
WriteNewCodeMap(File, OutputName, Contents.Blocks, Contents.IsExecutable, Dependencies);
}
}
static int ProcessAll() {
const auto CacheDirectory = FEX::Config::GetCacheDirectory();
const std::string NewCodeMapDirectory = fmt::format("{}codemap/new", CacheDirectory);
const std::string ReadyCodeMapDirectory = fmt::format("{}codemap/ready", CacheDirectory);
// Import new code maps and aggregate them into ready code maps
std::filesystem::create_directories(ReadyCodeMapDirectory);
AggregateCodeMaps(NewCodeMapDirectory, ReadyCodeMapDirectory);
// Generate caches
fextl::string OutDir = CacheDirectory + "cache/";
std::filesystem::create_directories(OutDir);
// Iterate over all executables (.exe).
// These determine the emulator configuration to use when compiling dependencies.
for (auto& Entry : std::filesystem::directory_iterator(ReadyCodeMapDirectory)) {
std::ifstream CodeMap(Entry.path(), std::ios_base::binary);
auto Parsed = FEXCore::CodeMap::ParseCodeMap(CodeMap);
auto ExecutableIt = std::ranges::find_if(Parsed, [](const auto& Entry) { return Entry.second.IsExecutable; });
if (ExecutableIt == Parsed.end()) {
// Skip libraries; they're only processed as dependencies of a main executable
continue;
}
fmt::println("\nChecking caches for executable {}", ExecutableIt->second.Filename);
// TODO: Compute the cache config id from the active FEX configuration
uint64_t CodeCacheConfigId = 0;
auto GetCacheFilename = [&](const FEXCore::ExecutableFileInfo& File) {
return fmt::format("{}{}-{:016x}", OutDir, FEXCore::CodeMap::GetBaseFilename(File, false), CodeCacheConfigId);
};
// Check the main binary and all of its dependencies
for (auto& [FileId, Contents] : Parsed) {
const FEXCore::ExecutableFileInfo File {nullptr, FileId, Contents.Filename};
std::error_code ec;
const auto BinaryName = FEXCore::CodeMap::GetBaseFilename(File, false);
const auto MergedCodeMapFilename = fmt::format("{}/{}", ReadyCodeMapDirectory, BinaryName);
const auto LastCodeMapUpdate = std::filesystem::last_write_time(MergedCodeMapFilename, ec);
if (ec) {
// No reference code map exists for this dependency yet, so there's nothing to generate a cache from
continue;
}
if (std::filesystem::last_write_time(GetCacheFilename(File), ec) > LastCodeMapUpdate && !ec) {
fmt::println(" Cache up to date: {}", BinaryName);
continue;
}
// TODO: Also check for matching FEX version from cache header
fmt::println(" {} cache: {}", ec ? "Generating" : "Updating outdated", BinaryName);
// Defer to GenerateCache
const auto FileIdArg = fmt::format("{:016x}", FileId);
std::vector<const char*> GenerateArgs {
"generate", "--fileid", FileIdArg.c_str(), "--outdir", OutDir.c_str(), MergedCodeMapFilename.c_str(),
};
if (GenerateCache(GenerateArgs.size(), GenerateArgs.data()) != 0) {
fmt::println("ERROR: Cache generation failed for {}", BinaryName);
}
}
}
return 0;
}
int main(int argc, char** argv) {
#ifndef _WIN32
LogMan::Throw::InstallHandler(AssertHandler);
LogMan::Msg::InstallHandler(MsgHandler);
#else
FEX::Windows::Logging::Init();
#endif
std::vector<const char*> Args {argv + 1, argv + argc};
auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
Args[0] = CommandName.c_str();
if (!Args.empty()) {
Args[0] = CommandName.c_str();
}
if (argc >= 2 && argv[1] == std::string_view {"generate"}) {
return GenerateCache(argc - 1, Args.data());
} else if (argc >= 2 && argv[1] == std::string_view {"process-all"}) {
return ProcessAll();
} else {
fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
fmt::print("Commands:\n");
fmt::print(" generate\tTrigger cache generation from combined code map\n");
fmt::print(" process-all\tProcess all new code maps and update all caches\n");
return EXIT_FAILURE;
}
}
+1 -1
View File
@@ -21,7 +21,7 @@
#include <optional>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/types.h>
+16 -6
View File
@@ -12,6 +12,7 @@
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <fmt/ranges.h>
#include <inttypes.h>
#include <atomic>
#include <cassert>
@@ -64,6 +65,9 @@ static std::string NewCodeMapDirectory;
// Path to directory for processed code maps (suitable for cache generation)
static std::string ReadyCodeMapDirectory;
// Path to FEXOfflineCompiler executable (inferred from FEXServer install location)
const std::string OfflineCompilerPath = (std::filesystem::read_symlink("/proc/self/exe").parent_path() / "FEXOfflineCompiler").string();
void SetWatchFD(int FD) {
WatchFD = FD;
}
@@ -94,8 +98,8 @@ void CheckRaiseFDLimit() {
if (MaxFDs.rlim_cur == MaxFDs.rlim_max) {
fprintf(stderr, "[FEXMountDaemon] Our open FD limit is already set to max and we are wanting to increase it\n");
fprintf(stderr, "[FEXMountDaemon] FEXMountDaemon will now no longer be able to track new instances of FEX\n");
fprintf(stderr, "[FEXMountDaemon] Current limit is %zd(hard %zd) FDs and we are at %zd\n", MaxFDs.rlim_cur, MaxFDs.rlim_max,
GetNumFilesOpen());
fprintf(stderr, "[FEXMountDaemon] Current limit is %" PRIuMAX "(hard %" PRIuMAX ") FDs and we are at %zu\n", (uintmax_t)MaxFDs.rlim_cur,
(uintmax_t)MaxFDs.rlim_max, GetNumFilesOpen());
fprintf(stderr, "[FEXMountDaemon] Ask your administrator to raise your kernel's hard limit on open FDs\n");
return;
}
@@ -109,7 +113,8 @@ void CheckRaiseFDLimit() {
NewLimit.rlim_cur = std::min(NewLimit.rlim_cur, NewLimit.rlim_max);
if (setrlimit(RLIMIT_NOFILE, &NewLimit) != 0) {
fprintf(stderr, "[FEXMountDaemon] Couldn't raise FD limit to %zd even though our hard limit is %zd\n", NewLimit.rlim_cur, NewLimit.rlim_max);
fprintf(stderr, "[FEXMountDaemon] Couldn't raise FD limit to %" PRIu64 " even though our hard limit is %" PRIu64 "\n",
(uintmax_t)NewLimit.rlim_cur, (uintmax_t)NewLimit.rlim_max);
} else {
// Set the new limit
MaxFDs = NewLimit;
@@ -223,7 +228,12 @@ bool InitializeServerSocket(bool abstract) {
if (abstract) {
ServerSocketName = FEXServerClient::GetServerSocketName();
} else {
ServerSocketName = FEXServerClient::GetServerSocketPath();
ServerSocketName = FEXServerClient::GetServerSocketPath(false);
if (ServerSocketName.size() > sizeof(sockaddr_un::sun_path) - 1) {
LogMan::Msg::EFmt("Socket path '{}' too large for Unix domain sockets. Moving to tmp", ServerSocketName);
ServerSocketName = FEXServerClient::GetServerSocketPath(true);
}
// Unlink the socket file if it exists
// We are being asked to create a daemon, not error check
// We don't care if this failed or not
@@ -470,8 +480,8 @@ int32_t EmbedSubprocess(const char* path, char* const* args) {
* Spawn a FEXOfflineCompiler instance to generate a code cache from the given code map
*/
static int RunOfflineCompiler(const char* CodeMap) {
const char* ExecveArgs[] = {"FEXOfflineCompiler", "generate", CodeMap, nullptr};
return EmbedSubprocess("FEXOfflineCompiler", const_cast<char* const*>(&ExecveArgs[0]));
const char* ExecveArgs[] = {OfflineCompilerPath.c_str(), "generate", CodeMap, nullptr};
return EmbedSubprocess(OfflineCompilerPath.c_str(), const_cast<char* const*>(&ExecveArgs[0]));
};
void HandleSocketData(fasio::tcp_socket& Socket) {
+1 -1
View File
@@ -9,7 +9,7 @@
#include <fcntl.h>
#include <filesystem>
#include <sys/poll.h>
#include <poll.h>
#include <sys/stat.h>
#include <sys/wait.h>
#include <thread>
@@ -35,6 +35,7 @@ $end_info$
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <atomic>
#include <cstring>
@@ -1171,7 +1172,7 @@ GdbServer::HandledPacketType GdbServer::CommandMultiLetterV(const fextl::string&
}
if (packet.starts_with("vKill")) {
tgkill(::getpid(), ::getpid(), SIGKILL);
FHU::Syscalls::tgkill(::getpid(), ::getpid(), SIGKILL);
}
// TODO: vRun
@@ -649,17 +649,19 @@ void SignalDelegator::HandleGuestSignal(FEX::HLE::ThreadStateObject* ThreadObjec
"capacity size. This will "
"likely crash! Asserting now!");
// Peek into sigset_t implementation details, extracting the __val member for glibc and __bits for musl
auto& [sigmask_val] = _context->uc_sigmask;
static_assert(sizeof(sigmask_val[0]) == sizeof(uint64_t), "Unknown sigset_t layout");
ThreadObject->SignalInfo.DeferredSignalFrames.emplace_back(ThreadStateObject::DeferredSignalState {
.Info = SigInfo,
.Signal = Signal,
.SigMask = _context->uc_sigmask.__val[0],
.SigMask = sigmask_val[0],
});
uint64_t NewMask = GetNewSigMask(Signal);
// Update our host signal mask so we don't hit race conditions with signals
// This allows us to maintain the expected signal mask through the guest signal handling and then all the way back again
memcpy(&_context->uc_sigmask, &NewMask, sizeof(uint64_t));
sigmask_val[0] = GetNewSigMask(Signal);
// Now update the faulting page permissions so it will fault on write.
mprotect(reinterpret_cast<void*>(&Thread->InterruptFaultPage), sizeof(Thread->InterruptFaultPage), PROT_NONE);
@@ -60,6 +60,7 @@ $end_info$
#include <syscall.h>
#include <sys/mman.h>
#include <sys/utsname.h>
#include <thread>
#include <unistd.h>
namespace FEX::HLE {
@@ -899,9 +900,19 @@ uint64_t UnimplementedSyscallSafe(FEXCore::Core::CpuStateFrame* Frame, uint64_t
}
void SyscallHandler::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
TM.LockBeforeFork();
Thread->CTX->LockBeforeFork(Thread);
VMATracking.Mutex.lock();
while (true) {
TM.LockBeforeFork();
Thread->CTX->LockBeforeFork(Thread);
if (std::try_lock(CodeCachePatchingMutex, VMATracking.Mutex) == -1) {
break;
}
// Lock failed: Another thread has temporarily acquired these mutexes.
// Release them to a void a deadlock and retry later
CTX->UnlockAfterFork(Thread, false);
TM.UnlockAfterFork(Thread, false);
std::this_thread::sleep_for(std::chrono::milliseconds {10});
};
}
void SyscallHandler::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread, bool Child) {
@@ -910,8 +921,10 @@ void SyscallHandler::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThr
FM.SetProtectedCodeMapFD(-1);
VMATracking.Mutex.StealAndDropActiveLocks();
CodeCachePatchingMutex.StealAndDropActiveLocks();
} else {
VMATracking.Mutex.unlock();
CodeCachePatchingMutex.unlock();
}
CTX->UnlockAfterFork(LiveThread, Child);
@@ -108,6 +108,8 @@ public:
// does a guest munmap as if done via a guest syscall
virtual uint64_t GuestMunmap(FEXCore::Core::InternalThreadState* Thread, void* addr, uint64_t length) = 0;
virtual void AddVirtualPage(FEXCore::Core::InternalThreadState* Thread, uint64_t addr, size_t length, int prot) = 0;
};
class SyscallHandler : public FEXCore::HLE::SyscallHandler,
@@ -254,6 +256,10 @@ public:
std::optional<LateApplyExtendedVolatileMetadata> TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t addr, size_t length,
int prot, int flags, int fd, off_t offset,
std::optional<FEXCore::ExecutableFileSectionInfo>& CachedSection);
void AddVirtualPage(FEXCore::Core::InternalThreadState* Thread, uint64_t addr, size_t length, int prot) override;
using SyscallMmapInterface::AddVirtualPage;
void TrackMunmap(FEXCore::Core::InternalThreadState* Thread, void* addr, size_t length);
void TrackMremap(FEXCore::Core::InternalThreadState* Thread, uint64_t OldAddress, size_t OldSize, size_t NewSize, int flags, uint64_t NewAddress);
void TrackShmat(FEXCore::Core::InternalThreadState* Thread, int shmid, uint64_t shmaddr, int shmflg, uint64_t Length);
@@ -261,10 +267,14 @@ public:
void TrackMprotect(FEXCore::Core::InternalThreadState* Thread, void* addr, size_t len, int prot);
void TrackMadvise(FEXCore::Core::InternalThreadState* Thread, uintptr_t Base, uintptr_t Size, int advice);
void InvalidateCodeRangeIfNecessary(FEXCore::Core::InternalThreadState* Thread, uint64_t Base, uint64_t Length) {
void InvalidateCodeRangeIfNecessary(FEXCore::Core::InternalThreadState* Thread, uint64_t Base, uint64_t Length, bool CheckPendingVMAResources) {
if (SMCChecks != FEXCore::Config::CONFIG_SMC_NONE) {
TM.InvalidateGuestCodeRange(Thread, Base, Length);
}
if (CheckPendingVMAResources && Thread) {
auto lk = FEXCore::GuardSignalDeferringSection(VMATracking.Mutex, Thread);
VMATracking.FlushPendingResourceDeletions();
}
}
void InvalidateCodeRangeIfNecessaryOnRemap(FEXCore::Core::InternalThreadState* Thread, uint64_t OldAddress, uint64_t NewAddress,
@@ -292,6 +302,7 @@ public:
std::optional<FEXCore::ExecutableFileSectionInfo>
LookupExecutableFileSection(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestAddr) final override;
void TriggerGuestLibWrapperCodeCacheLoad(FEXCore::Core::InternalThreadState&, uint64_t AnyAddr);
int OpenCodeMapFile() override;
FEXCore::HLE::ExecutableRangeInfo QueryGuestExecutableRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Address) override;
@@ -358,6 +369,9 @@ private:
std::mutex FutexMutex;
std::mutex SyscallMutex;
// std::mutex CodeCachePatchingMutex;
FEXCore::ForkableUniqueMutex CodeCachePatchingMutex;
FEX::CodeLoader* LocalLoader {};
bool NeedToCheckXID {true};
@@ -437,7 +437,6 @@ namespace x64 {
REGISTER_SYSCALL_IMPL_X64(getsockopt, SyscallPassthrough5<SYSCALL_DEF(getsockopt)>);
REGISTER_SYSCALL_IMPL_X64(wait4, SyscallPassthrough4<SYSCALL_DEF(wait4)>);
REGISTER_SYSCALL_IMPL_X64(semop, SyscallPassthrough3<SYSCALL_DEF(semop)>);
REGISTER_SYSCALL_IMPL_X64(gettimeofday, SyscallPassthrough2<SYSCALL_DEF(gettimeofday)>);
REGISTER_SYSCALL_IMPL_X64(getrlimit, SyscallPassthrough2<SYSCALL_DEF(getrlimit)>);
REGISTER_SYSCALL_IMPL_X64(getrusage, SyscallPassthrough2<SYSCALL_DEF(getrusage)>);
REGISTER_SYSCALL_IMPL_X64(sysinfo, SyscallPassthrough1<SYSCALL_DEF(sysinfo)>);
@@ -33,7 +33,7 @@ $end_info$
#include <stdint.h>
#include <sched.h>
#include <sys/personality.h>
#include <sys/poll.h>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/resource.h>
#include <sys/syscall.h>
@@ -30,7 +30,21 @@ $end_info$
#include <Linux/Utils/ELFParser.h>
namespace FEX::HLE {
// SMC interactions
static void HandleSegfaultForCodeCacheFinalization(FEXCore::Core::InternalThreadState& Thread, FEXCore::MappedCodeCacheFile& Code,
uintptr_t FaultAddress) {
FEXCORE_PROFILE_SCOPED("Load code cache page");
size_t PageIdx = (reinterpret_cast<std::byte*>(FaultAddress) - Code.CodeBuffer.data()) / FEXCore::Utils::FEX_PAGE_SIZE;
auto RangeToFinalize = Thread.CTX->GetCodeCache().SelectCodeRangeToFinalize(Code, PageIdx, PageIdx + 1);
if (!RangeToFinalize.empty()) {
Thread.CTX->GetCodeCache().FinalizeCodePages(Code, RangeToFinalize);
}
}
// Handles segfaults from:
// - call-ret shadow stack overflow
// - guest-side self-modifying code (SMC)
// - lazy loading of mapped code cache pages
bool SyscallHandler::HandleSegfault(FEXCore::Core::InternalThreadState* Thread, int Signal, void* info, void* ucontext) {
const auto FaultAddress = (uintptr_t)((siginfo_t*)info)->si_addr;
@@ -46,13 +60,25 @@ bool SyscallHandler::HandleSegfault(FEXCore::Core::InternalThreadState* Thread,
// Can't use the deferred signal lock in the SIGSEGV handler.
auto lk = FEXCore::MaskSignalsAndLockMutex<std::shared_lock>(_SyscallHandler->VMATracking.Mutex);
auto VMATracking = &_SyscallHandler->VMATracking;
auto& VMATracking = _SyscallHandler->VMATracking;
// If the write spans two pages, they will be flushed one at a time (generating two faults)
auto Entry = VMATracking->FindVMAEntry(FaultAddress);
auto Entry = VMATracking.FindVMAEntry(FaultAddress);
// If an untracked address, or the mapping wasn't writable, it can't be handled here
if (Entry == VMATracking->VMAs.end() || !Entry->second.Prot.Writable) {
if (Entry == VMATracking.VMAs.end()) {
// Not a guest page; check mapped code cache pages
auto* Code = VMATracking.FindMappedCodeCacheByHostAddress(FaultAddress);
if (!Code) {
// Untracked address; not handled here
return false;
}
std::lock_guard lk(_SyscallHandler->CodeCachePatchingMutex);
HandleSegfaultForCodeCacheFinalization(*Thread, *Code, FaultAddress);
return true;
}
// If the mapping wasn't writable, it can't be handled here
if (!Entry->second.Prot.Writable) {
return false;
}
@@ -170,7 +196,7 @@ void SyscallHandler::MarkGuestExecutableRange(FEXCore::Core::InternalThreadState
}
void SyscallHandler::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
InvalidateCodeRangeIfNecessary(Thread, Start, Length);
InvalidateCodeRangeIfNecessary(Thread, Start, Length, false);
}
static FEXCore::ExecutableFileSectionInfo BuildSectionInfo(const VMATracking::MappedResource& Resource, uint64_t Base, uint64_t Size) {
@@ -233,41 +259,56 @@ static ReadELFHeadersResult ReadELFHeaders(int FD, std::span<std::byte> HeaderDa
return ReadELFHeadersResult {std::move(Parser.phdrs), std::move(Relocations), HasCodeRelocations};
}
static void LoadCodeCache(FEXCore::Core::InternalThreadState& Thread, FEXCore::ExecutableFileSectionInfo& Section, uint64_t CodeCacheConfigId) {
static fextl::unique_ptr<FEXCore::MappedCodeCacheFile>
LoadCodeCache(FEXCore::Core::InternalThreadState& Thread, VMATracking::VMATracking& VMATracking,
const FEXCore::ExecutableFileInfo& FileInfo, uint64_t CodeCacheConfigId, uint64_t FileStartVA) {
auto& CodeCache = Thread.CTX->GetCodeCache();
auto CacheFilename = fextl::fmt::format("{}cache/{}-{:016x}", FEX::Config::GetCacheDirectory(),
FEXCore::CodeMap::GetBaseFilename(Section.FileInfo, false), CodeCacheConfigId);
FEXCore::CodeMap::GetBaseFilename(FileInfo, false), CodeCacheConfigId);
int CacheFD = open(CacheFilename.c_str(), O_RDONLY);
if (CacheFD == -1) {
LogMan::Msg::IFmt("Cache file does not exist: {}", CacheFilename);
return;
return nullptr;
}
struct stat buf;
if (fstat(CacheFD, &buf) != 0) {
LogMan::Msg::EFmt("Invalid cache file: {}", CacheFilename);
close(CacheFD);
return;
return nullptr;
}
auto CacheFileSize = buf.st_size;
auto CacheFileSize = static_cast<std::size_t>(buf.st_size);
auto MappedCache = (std::byte*)FEXCore::Allocator::mmap(nullptr, CacheFileSize, PROT_READ, MAP_PRIVATE, CacheFD, 0);
LOGMAN_THROW_A_FMT(MappedCache, "Failed to map code cache into memory");
if (!Thread.CTX->GetCodeCache().LoadData(&Thread, MappedCache, Section)) {
// TODO: Delete this cache file
}
FEXCore::Allocator::munmap(MappedCache, CacheFileSize);
close(CacheFD);
if (!MappedCache || MappedCache == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to map code cache into memory");
return nullptr;
}
auto Result = CodeCache.LoadCache(std::span {MappedCache, CacheFileSize}, FileInfo, FileStartVA);
if (!Result) {
FEXCore::Allocator::munmap(MappedCache, CacheFileSize);
return nullptr;
}
// NOTE: This is synchronized by acquiring VMATracking.Mutex at call site
CodeCache.RegisterMappedCodeBuffer(*Result);
return Result;
}
void* SyscallHandler::GuestMmap(bool Is64Bit, FEXCore::Core::InternalThreadState* Thread, void* addr, size_t length, int prot, int flags,
int fd, off_t offset) {
LOGMAN_THROW_A_FMT(Is64Bit || (length >> 32) == 0, "values must fit to 32 bits");
uint64_t Result {};
uint64_t Result;
size_t Size = FEXCore::AlignUp(length, FEXCore::Utils::FEX_PAGE_SIZE);
std::optional<LateApplyExtendedVolatileMetadata> LateMetadata = std::nullopt;
std::optional<FEXCore::ExecutableFileSectionInfo> CachedSection;
bool PendingResourceDeletion;
{
// NOTE: Frontend calls this with a nullptr Thread during initialization, but
@@ -290,9 +331,10 @@ void* SyscallHandler::GuestMmap(bool Is64Bit, FEXCore::Core::InternalThreadState
}
LateMetadata = TrackMmap(Thread, Result, length, prot, flags, fd, offset, CachedSection);
PendingResourceDeletion = VMATracking.HasPendingResourceDeletions();
}
InvalidateCodeRangeIfNecessary(Thread, Result, Size);
InvalidateCodeRangeIfNecessary(Thread, Result, Size, PendingResourceDeletion);
if (LateMetadata) {
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
@@ -300,7 +342,8 @@ void* SyscallHandler::GuestMmap(bool Is64Bit, FEXCore::Core::InternalThreadState
}
if (EnableCodeCaching && CachedSection) {
LoadCodeCache(*Thread, *CachedSection, CodeCacheConfigId);
Thread->CTX->GetCodeCache().EnableLoadedSection(
Thread, *static_cast<const VMATracking::ExecutableFileState&>(CachedSection->FileInfo).MappedCache, *CachedSection);
}
return reinterpret_cast<void*>(Result);
@@ -310,8 +353,9 @@ uint64_t SyscallHandler::GuestMunmap(bool Is64Bit, FEXCore::Core::InternalThread
LOGMAN_THROW_A_FMT(Is64Bit || (reinterpret_cast<uintptr_t>(addr) >> 32) == 0, "values must fit to 32 bits: {}", fmt::ptr(addr));
LOGMAN_THROW_A_FMT(Is64Bit || (length >> 32) == 0, "values must fit to 32 bits");
uint64_t Result {};
uint64_t Result;
uint64_t Size = FEXCore::AlignUp(length, FEXCore::Utils::FEX_PAGE_SIZE);
bool PendingResourceDeletion;
{
// Frontend calls this with nullptr Thread during initialization.
@@ -331,8 +375,9 @@ uint64_t SyscallHandler::GuestMunmap(bool Is64Bit, FEXCore::Core::InternalThread
}
}
TrackMunmap(Thread, addr, length);
PendingResourceDeletion = VMATracking.HasPendingResourceDeletions();
}
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), Size);
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), Size, PendingResourceDeletion);
if (length) {
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
@@ -366,6 +411,26 @@ uint64_t SyscallHandler::GuestMremap(bool Is64Bit, FEXCore::Core::InternalThread
return Result;
}
void SyscallHandler::TriggerGuestLibWrapperCodeCacheLoad(FEXCore::Core::InternalThreadState& Thread, uint64_t AnyAddr) {
if (!EnableCodeCaching) {
return;
}
// TODO: Instead of deferring the entire cache load, only delay applicance of sha256 relocations!
auto lk = FEXCore::GuardSignalDeferringSection<std::shared_lock>(VMATracking.Mutex, &Thread);
auto VMAEntry = VMATracking.FindVMAEntry(reinterpret_cast<uint64_t>(AnyAddr));
for (auto* VMA = VMAEntry->second.Resource->FirstVMA; VMA; VMA = VMA->ResourceNextVMA) {
if (!VMA->Prot.Executable) {
continue;
}
auto SectionInfo = BuildSectionInfo(*VMAEntry->second.Resource, VMA->Base, VMA->Length);
LoadCodeCache(Thread, VMATracking, SectionInfo.FileInfo, CodeCacheConfigId, SectionInfo.FileStartVA);
}
}
int SyscallHandler::OpenCodeMapFile() {
// Query from FEXServer whether this is the first instance of this executable; if it is, also enable code dumping!
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
@@ -413,7 +478,7 @@ uint64_t SyscallHandler::GuestMprotect(FEXCore::Core::InternalThreadState* Threa
TrackMprotect(Thread, addr, len, prot);
}
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), len);
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), len, false);
// Prepare for delayed code cache load after ld/Wine is done applying relocations.
// Hooking into mprotect is a reliable heuristic that matches behavior of ld (for ELF) and Wine (for PE).
@@ -437,17 +502,21 @@ uint64_t SyscallHandler::GuestMprotect(FEXCore::Core::InternalThreadState* Threa
}
// Trigger delayed cache load. This must be done separately since
// LoadCodeCache will call interfaces that acquire the VMATracking mutex.
// EnableLoadedSection will call interfaces that acquire the VMATracking mutex.
for (auto& CachedSection : CachedSections) {
LoadCodeCache(*Thread, CachedSection, CodeCacheConfigId);
auto Cache = static_cast<const VMATracking::ExecutableFileState&>(CachedSection.FileInfo).MappedCache.get();
if (Cache) {
Thread->CTX->GetCodeCache().EnableLoadedSection(Thread, *Cache, CachedSection);
}
}
return Result;
}
uint64_t SyscallHandler::GuestShmat(bool Is64Bit, FEXCore::Core::InternalThreadState* Thread, int shmid, const void* shmaddr, int shmflg) {
uint64_t Result {};
uint64_t Length {};
uint64_t Result;
uint64_t Length;
bool PendingResourceDeletion;
{
auto lk = FEXCore::GuardSignalDeferringSection(VMATracking.Mutex, Thread);
@@ -472,15 +541,17 @@ uint64_t SyscallHandler::GuestShmat(bool Is64Bit, FEXCore::Core::InternalThreadS
Length = stat.shm_segsz;
TrackShmat(Thread, shmid, Result, shmflg, Length);
PendingResourceDeletion = VMATracking.HasPendingResourceDeletions();
}
InvalidateCodeRangeIfNecessary(Thread, Result, Length);
InvalidateCodeRangeIfNecessary(Thread, Result, Length, PendingResourceDeletion);
return Result;
}
uint64_t SyscallHandler::GuestShmdt(bool Is64Bit, FEXCore::Core::InternalThreadState* Thread, const void* shmaddr) {
uint64_t Result {};
uint64_t Length {};
uint64_t Result;
uint64_t Length;
bool PendingResourceDeletion;
{
auto lk = FEXCore::GuardSignalDeferringSection(VMATracking.Mutex, Thread);
if (Is64Bit) {
@@ -496,9 +567,10 @@ uint64_t SyscallHandler::GuestShmdt(bool Is64Bit, FEXCore::Core::InternalThreadS
}
Length = TrackShmdt(Thread, reinterpret_cast<uintptr_t>(shmaddr));
PendingResourceDeletion = VMATracking.HasPendingResourceDeletions();
}
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uintptr_t>(shmaddr), Length);
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uintptr_t>(shmaddr), Length, PendingResourceDeletion);
return Result;
}
@@ -537,7 +609,7 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
if (PathLength != -1 && S_ISREG(buf.st_mode) && (buf.st_mode & S_IXUSR)) {
// ELF files that are mapped multiple times get a separate MappedResource for each base virtual address
if ((prot & PROT_READ) && Inserted) {
Resource->MappedFile = fextl::make_unique<FEXCore::ExecutableFileInfo>();
Resource->MappedFile = fextl::make_unique<VMATracking::ExecutableFileState>();
Resource->MappedFile->Filename = fextl::string(Tmp, PathLength);
Resource->MappedFile->FileId = CTX->GetCodeCache().ComputeCodeMapId(Resource->MappedFile->Filename, fd);
@@ -623,7 +695,18 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
// Load code cache if present.
// FEXServer was requested to generate library caches on program launch.
if (EnableCodeCaching && Resource && Resource->MappedFile && VMATracking::VMAProt::fromProt(prot).Executable) {
if (Thread) {
if (!Resource->MappedFile->AttemptedCacheLoad) {
Resource->MappedFile->MappedCache = LoadCodeCache(*Thread, VMATracking, *Resource->MappedFile, CodeCacheConfigId, Resource->FirstVMA->Base);
Resource->MappedFile->AttemptedCacheLoad = true;
}
if (!Resource->MappedFile->MappedCache) {
// No cache present
} else if (Resource->MappedFile->Filename.ends_with("-guest.so")) {
// For guest library wrappers, cache loading must be delayed until LoadLib is called.
// Before that, we can't patch up the SHA256 function identifiers.
LogMan::Msg::IFmt("Delaying code cache load for {}", Resource->MappedFile->Filename);
} else if (Thread) {
if (!Resource->RequiresDelayedCacheLoad) {
CachedSection.emplace(BuildSectionInfo(*Resource, addr, Size));
} else {
@@ -638,6 +721,12 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
return VolatileMetadata;
}
void SyscallHandler::AddVirtualPage(FEXCore::Core::InternalThreadState* Thread, uint64_t addr, size_t length, int prot) {
auto lk = FEXCore::GuardSignalDeferringSectionWithFallback(VMATracking.Mutex, Thread);
VMATracking.TrackVMARange(CTX, nullptr, addr, 0, length, VMATracking::VMAFlags::fromFlags(MAP_ANONYMOUS | MAP_PRIVATE),
VMATracking::VMAProt::fromProt(prot));
}
void SyscallHandler::TrackMunmap(FEXCore::Core::InternalThreadState* Thread, void* addr, size_t length) {
uint64_t Size = FEXCore::AlignUp(length, FEXCore::Utils::FEX_PAGE_SIZE);
VMATracking.DeleteVMARange(CTX, reinterpret_cast<uintptr_t>(addr), Size);
@@ -7,10 +7,20 @@ desc: VMA Tracking
$end_info$
*/
#include "FEXCore/Utils/MathUtils.h"
#include "LinuxSyscalls/Syscalls.h"
#include <sys/shm.h>
namespace FEX::HLE::VMATracking {
ExecutableFileState::~ExecutableFileState() {
if (MappedCache && MappedCache->MappedFile.data()) {
auto ret = FEXCore::Allocator::munmap(MappedCache->MappedFile.data(),
FEXCore::AlignUp(MappedCache->MappedFile.size_bytes(), FEXCore::Utils::FEX_PAGE_SIZE));
LOGMAN_THROW_A_FMT(ret == 0, "Error unmapping cache for {}: {} {}", Filename, errno, strerror(errno));
}
}
/// Helpers ///
auto VMAProt::fromProt(int Prot) -> VMAProt {
return VMAProt {
@@ -235,7 +245,13 @@ void VMATracking::DeleteVMARange(FEXCore::Context::Context* CTX, uintptr_t Base,
// If linked to a Mapped Resource, remove from linked list and possibly delete the Mapped Resource
if (Current->Resource) {
if (ListRemove(Current) && Current->Resource != PreservedMappedResource) {
MappedResources.erase(Current->Resource->Iterator);
auto Iter = Current->Resource->Iterator;
// Defer deletion if the resource has mapped code cache data, so its code buffer
// outlives code cache invalidation (which runs after the VMA lock is released).
if (Current->Resource->MappedFile && Current->Resource->MappedFile->MappedCache) {
PendingResourceDeletions.push_back(std::move(*Current->Resource));
}
MappedResources.erase(Iter);
}
}
@@ -277,6 +293,11 @@ void VMATracking::DeleteVMARange(FEXCore::Context::Context* CTX, uintptr_t Base,
}
}
void VMATracking::FlushPendingResourceDeletions() {
Mutex.check_lock_owned_by_self_as_write();
PendingResourceDeletions.clear();
}
// Change flags of mappings in a range and split the mappings if needed
void VMATracking::ChangeProtectionFlags(uintptr_t Base, uintptr_t Length, VMAProt NewProt) {
Mutex.check_lock_owned_by_self_as_write();
@@ -542,7 +563,11 @@ uintptr_t VMATracking::DeleteSHMRegion(FEXCore::Context::Context* CTX, uintptr_t
do {
if (Entry->second.Resource == Resource) {
if (ListRemove(&Entry->second)) {
MappedResources.erase(Entry->second.Resource->Iterator);
auto Iter = Entry->second.Resource->Iterator;
if (Entry->second.Resource->MappedFile && Entry->second.Resource->MappedFile->MappedCache) {
PendingResourceDeletions.push_back(std::move(*Entry->second.Resource));
}
MappedResources.erase(Iter);
}
Entry = VMAs.erase(Entry);
} else {
@@ -552,4 +577,18 @@ uintptr_t VMATracking::DeleteSHMRegion(FEXCore::Context::Context* CTX, uintptr_t
return ShmLength;
}
FEXCore::MappedCodeCacheFile* VMATracking::FindMappedCodeCacheByHostAddress(uintptr_t HostAddr) const {
for (auto& [_, Resource] : MappedResources) {
if (Resource.MappedFile && Resource.MappedFile->MappedCache) {
auto* Code = Resource.MappedFile->MappedCache.get();
auto BufferStart = reinterpret_cast<uintptr_t>(Code->CodeBuffer.data());
if (HostAddr >= BufferStart && HostAddr < BufferStart + Code->CodeBuffer.size_bytes()) {
return Code;
}
}
}
return nullptr;
}
} // namespace FEX::HLE::VMATracking
@@ -4,12 +4,18 @@
#include <cstdint>
#include <tuple>
#include "FEXCore/Core/CodeCache.h"
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <elf.h>
namespace FEXCore {
struct ExecutableFileInfo;
struct MappedCodeCacheFile;
} // namespace FEXCore
namespace FEX::HLE::VMATracking {
///// VMA (Virtual Memory Area) tracking /////
@@ -32,6 +38,13 @@ struct MRID {
struct VMAEntry;
struct ExecutableFileState : FEXCore::ExecutableFileInfo {
~ExecutableFileState();
bool AttemptedCacheLoad = false;
fextl::unique_ptr<FEXCore::MappedCodeCacheFile> MappedCache;
};
/**
* Meta data associated to one system resource.
*
@@ -43,7 +56,7 @@ struct VMAEntry;
struct MappedResource {
using ContainerType = fextl::multimap<MRID, MappedResource>;
fextl::unique_ptr<FEXCore::ExecutableFileInfo> MappedFile;
fextl::unique_ptr<ExecutableFileState> MappedFile;
// Pointer to lowest memory range this file is mapped to
VMAEntry* FirstVMA;
uint64_t Length; // 0 if not fixed size
@@ -136,8 +149,22 @@ struct VMATracking {
return MappedResources.equal_range(mrid);
}
// Find any MappedCodeCacheFile that contains the given host code address.
// - Mutex must be shared_locked before calling
FEXCore::MappedCodeCacheFile* FindMappedCodeCacheByHostAddress(uintptr_t HostAddr) const;
bool HasPendingResourceDeletions() const {
return !PendingResourceDeletions.empty();
}
// Flush pending MappedResource deletions. This must be called after code
// invalidation related to unmapped/remapped memory to avoid memory leaks.
// - Mutex must be unique_locked before calling
void FlushPendingResourceDeletions();
private:
MappedResource::ContainerType MappedResources;
fextl::vector<MappedResource> PendingResourceDeletions;
};
@@ -30,7 +30,7 @@ $end_info$
#include <optional>
#include <sys/stat.h>
#include <bits/types/sigset_t.h>
#include <signal.h>
#include <linux/seccomp.h>
namespace FEX::HLE {
@@ -22,6 +22,16 @@ $end_info$
namespace FEX::HLE::x64 {
void RegisterTime(FEX::HLE::SyscallHandler* Handler) {
using namespace FEXCore::IR;
REGISTER_SYSCALL_IMPL_X64(gettimeofday, [](FEXCore::Core::CpuStateFrame* Frame, timeval* tv, struct timezone* tz) -> uint64_t {
FaultSafeUserMemAccess::VerifyIsWritableOrNull(tv, sizeof(*tv));
FaultSafeUserMemAccess::VerifyIsWritableOrNull(tz, sizeof(*tz));
// Passed through glibc to ensure vdso is used if possible.
uint64_t Result = ::gettimeofday(tv, tz);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X64(time, [](FEXCore::Core::CpuStateFrame* Frame, time_t* tloc) -> uint64_t {
FaultSafeUserMemAccess::VerifyIsWritableOrNull(tloc, sizeof(time_t));
uint64_t Result = ::time(tloc);
+2
View File
@@ -259,6 +259,8 @@ void ThunkHandler_impl::LoadLib(std::string_view Name) {
LogMan::Msg::DFmt("Loaded {} syms", i);
}
_SyscallHandler->TriggerGuestLibWrapperCodeCacheLoad(*ThreadObject->Thread, ThreadObject->Thread->CurrentFrame->State.rip);
}
/**
@@ -21,6 +21,10 @@ struct VDSOMapping {
size_t VDSOSize {};
void* X86GeneratedCodePtr {};
size_t X86GeneratedCodeSize {};
explicit operator bool() const {
return VDSOBase != nullptr;
}
};
struct VDSOEntrypoints {
Loaded 100 of 449 files, more files were not shown because too many files have changed in this diff. Show more