Compare commits

..
96 Commits
Author SHA1 Message Date
Ryan Houdek 8c956e6ce1 Docs: Update for release FEX-2202 2022-02-05 22:48:35 -08:00
Ryan Houdek 832d013c92 Merge pull request #1513 from Sonicadvance1/reduce_flags_memory_usage
FEXCore: Defer a significant number of ALU flag calculation
2022-02-04 16:22:24 -08:00
Mai M 1c24206117 Merge pull request #1550 from Sonicadvance1/fix_weirdo_crc32
OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
2022-02-04 01:01:50 -05:00
Ryan Houdek f1979c15a2 unittests: Adds new CRC32 unittests
The instruction decode tables for crc32 introduced some dumb.
`F2h` and `F2h && 66h` prefixes both work for crc32.
This is a failure on Intel's part for sticking crc32 in to the vector
table.

MOVBE without any prefixes also does the same garbage where prefix `66h`
acts as an operand prefix size ONLY.
2022-02-03 21:11:50 -08:00
Ryan Houdek 556a1dab24 OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
This table is particularly terrible. CRC32 is the first instruction in
this table that needs either prefix `72h` OR `66h && F2h`

For 8bit CRC32, this ignores the 66h operand size override prefix.
  - But our table decoding didn't handle this
For 16bit/32bit/64bit CRC32 this behaviour changes depending on 66h
prefix AND REX.W
  - 66h prefix is ignored when REX.W is set, always 64bit but it falls
    down the other table path

This is an absolutely weird edge case that nobody should hit, but here
we are.
2022-02-03 21:11:50 -08:00
Mai M caffad8562 Merge pull request #1549 from Sonicadvance1/implement_pcmpgtq
OpcodeDispatcher: Implements PCMPGTQ
2022-02-03 21:46:13 -05:00
Mai M 5978143141 Merge pull request #1547 from Sonicadvance1/remove_system_xxhash
CMake: Always use local xxhash to statically link
2022-02-03 21:45:58 -05:00
Ryan Houdek 594c70b5e0 OpcodeDispatcher: Implements PCMPGTQ
I thought we already had this implemented but I guess it was missed.

Required for SSE 4.2
2022-02-03 18:36:29 -08:00
Ryan Houdek 655e6989ca FEXCore: Defer a significant number of ALU flag calculation
This was mainly an optimization around memory usage. ALU ops tend to
bloat the IR quite heavily, but I also noticed a 2-4% uplift in
performance of some applications. So a nice side effect.

Should let us more aggressively target reducing our IR intrusive
allocator size since this is quite reduced.

In a pedantic heavy ALU op code block this reduces the number of IR ops
from 14,756 IR ops to 2,016 prior to optimization.
After optimization both had reduced down to 50 IR ops, proving the
output IR was the same.
2022-02-03 01:37:59 -08:00
Ryan Houdek afeb228a89 CMake: Always use local xxhash to statically link
Dynamically linking xxhash is causing problems with pressure-vessel.

With this in place we only have the typical C++ dependencies
```
$ ldd ./Bin/FEXLoader
        linux-vdso.so.1 (0x00007fff44d9d000)
        libstdc++.so.6 => /lib/x86_64-linux-gnu/libstdc++.so.6 (0x00007f4c4d884000)
        libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007f4c4d7a0000)
        libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1 (0x00007f4c4d786000)
        libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007f4c4d55e000)
        /lib64/ld-linux-x86-64.so.2 (0x00007f4c4e0fa000)
```
2022-02-03 01:31:43 -08:00
Mai M d308a438ea Merge pull request #1546 from Sonicadvance1/fix_fexconfig
Fixes FEXConfig build
2022-02-01 21:13:08 -05:00
Ryan Houdek 260fc8ba52 Fixes FEXConfig build
Oops. This was added late and didn't test it.
2022-02-01 15:44:06 -08:00
Mai M ade0d0f241 Merge pull request #1543 from Sonicadvance1/fixes_for_1423
Linux: Fixes for older build environments
2022-02-01 16:52:40 -05:00
Mai M 11a5105547 Merge pull request #1544 from Sonicadvance1/allow_disable_interpreter
Adds an option to disable the IR interpreter
2022-02-01 16:52:22 -05:00
Ryan Houdek 10ad5db686 Adds an option to disable the IR interpreter
By default we won't build with the interpeter to reduce user confusion.
The interpreter isn't really useful to end users so remove it.

Completely removes it from building except for the fallback operations.

This also removes the selection from FEXConfig to remove selection
confusion there.

File Stats:
FEXLoader Size with Interpreter:    3422768 bytes
FEXLoader Size without Interpreter: 3301944 bytes
Size difference:                    96.4699915%
Bytes removed:                      120824 bytes
4k pages removed:                   29.498046875 -> 30 rounded up

VM Stats (Reported from bloaty):
Memory Size with Interpreter:    6.50Mi
Memory Size without Interpreter: 6.38Mi
Size difference:                 98.1538462%
2022-02-01 13:00:29 -08:00
Ryan Houdek 68c441575d Linux: Fixes for older build environments
Should resolve the new building issues from #1423
2022-02-01 12:17:09 -08:00
Ryan Houdek 334a8ef87c Merge pull request #1542 from Sonicadvance1/fix_pressure_vessel_hangs
Fix pressure vessel hangs
2022-01-31 08:57:20 -08:00
Ryan Houdek b7a76af72f Merge pull request #1541 from Sonicadvance1/implement_crc
OpcodeDispatcher: Implements CRC32 instruction
2022-01-31 08:57:01 -08:00
Ryan Houdek 4c92b562b8 Merge pull request #1540 from Sonicadvance1/remove_extract
OpcodeDispatcher: Removes extraneous extract in VFCMP
2022-01-31 08:56:47 -08:00
Stefanos Kornilios Mitsis Poiitidis 9d08451903 Merge pull request #1536 from Sonicadvance1/fix_orbitals
Softfloat: Stop doing special handling for FREM
2022-01-31 16:51:18 +02:00
Stefanos Kornilios Mitsis Poiitidis c252f8bfc5 Merge pull request #1539 from Sonicadvance1/fix_wrong_offsets
IR: Fixes some wrong offsets in passes
2022-01-31 15:40:33 +02:00
Ryan Houdek dc7ec6377b Linux: Safely handle Filemanagement mutex on fork
If an application is forking heavily with threaded file accesses
happening then the mutex can end up in an unknown state.

On fork make sure to lock the mutex then immediately unlock after fork
occurs.

This final step resolves hanging that pressure-vessel hits on startup.
Since it is doing a ton of file opening and forking during
initialization.
2022-01-30 18:15:57 -08:00
Ryan Houdek ce6f4edaaa FileManagement: Use ScopedSignalMaskWithMutex
When using mutexes in syscall helpers we need to be extra careful around
signals.
2022-01-30 18:15:57 -08:00
Ryan Houdek 983c35ea3b Allocator: Use ScopedSignalMaskWithMutex
Instead of just a basic mutex, also mask the signals.
This fixes the problem where we can end up receiving a signal in the
middle of memory allocation. Thus leaving the locked mutex in a broken
state.

This more closely matches the Linux kernel behaviour.
Since if you're in the middle of a memory allocating syscall, you won't
get signaled.
2022-01-30 18:15:57 -08:00
Ryan Houdek 70aaa1117a FEXHeaderUtils: Adds ScopedSignalMaskWithMutex
This class allows a scoped region lock a mutex and mask signals.

This is necessary for thread and signal safety coming up
2022-01-30 18:15:57 -08:00
Ryan Houdek 59e9859087 unittests: Implements CRC32 unit tests 2022-01-30 15:38:26 -08:00
Ryan Houdek d9453ff639 OpcodeDispatcher: Implements CRC32 instruction
Now that the rest of the code matches behaviour, we just need to pass
this through.

Easy enough and get Horizon Zero Dawn running.
2022-01-30 15:38:26 -08:00
Ryan Houdek 70754991d1 CPUID: Fill out CPUID for SSE4.2 feature
Currently force disabled until the rest of SSE 4.2 is enabled
This is to remind us in the future that SSE4.2 can only be enabled in
CPUID with CRC32 instruction support.
2022-01-30 15:38:26 -08:00
Ryan Houdek 57ebfceb48 HostFeatures: Check for CRC32 op support
Available with CRC32 bit on Arm64 or SSE4.2 on x86-64
2022-01-30 15:38:26 -08:00
Ryan Houdek 9e224d2bb0 x86 JIT: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek e43bd04901 JITArm64: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 48762e03a6 Interpreter: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 858924309e IR: Implements CRC32 op 2022-01-29 23:33:53 -08:00
Ryan Houdek cab02d1e65 OpcodeDispatcher: Removes extraneous extract in VFCMP
We don't need to extract the element to compare it.
2022-01-28 22:22:34 -08:00
Ryan Houdek 2a64f80567 IR: Fixes some wrong offsets in passes
GPR and FPR ending offsets were off by one here. Just a quick fix.
2022-01-28 22:19:51 -08:00
Ryan Houdek 174ddea99d Softfloat: Stop doing special handling for FREM
This isn't correct and breaks games.
This makes the FREM and REM1 implementation the same.
While not 100% correct, it is still better than before.
New issues will be created to handle the differences in the future.

Fixes #1374.
Also fixes most of the HL2 issues, just not the seam issue.
2022-01-28 19:42:57 -08:00
Ryan Houdek 6fb0e3c0cf Softfloat: Allow x87 fallback for all ops 2022-01-28 19:42:25 -08:00
Ryan Houdek 2b044bbdf4 Disable fprem unittests
These are about to be broken
2022-01-28 19:39:03 -08:00
Mai M ea76de0fd2 Merge pull request #1533 from Sonicadvance1/revise_posix_tests
unittests: Revise POSIX tests known failures and disabled
2022-01-25 15:19:24 -05:00
Ryan Houdek 9c8642e0dc unittests: Revise POSIX tests known failures and disabled
Some of these behaviours have changed now, particularly around signal
handling.

Some things still fail now of course. But most everything is now
documented as to why it is failing or disabled.

Fixes #955
2022-01-25 11:41:45 -08:00
Ryan Houdek 13f35f7b79 Merge pull request #1530 from Sonicadvance1/rootfs_fetcher_fixes
FEXRootFSFetcher: Fixes some edge case behaviours
2022-01-25 10:29:55 -08:00
Ryan Houdek e2798e370e Merge pull request #1518 from Sonicadvance1/fix_signed_branch
JIT: Fixes signed displacement wraparound on 32-bit
2022-01-25 10:29:46 -08:00
Ryan Houdek d9149548b5 Merge pull request #1531 from Sonicadvance1/fix_sockopt
Linux: Fixes 32-bit getsockopt and setsockopt
2022-01-25 09:09:16 -08:00
Ryan Houdek 4823933f79 Linux: Fixes 32-bit getsockopt and setsockopt
On Set, we have four options that need to be converted.
On Get, we have two options that need to be converted.

This fixes a crash that Tomb Raider 2013 was having on launch.
2022-01-24 17:14:50 -08:00
Ryan Houdek ee04067424 FEXRootFSFetcher: Fixes some edge case behaviours
Makes curl do its continue feature to give the users the best chance of
downloading a rootfs. We don't need to restart the full file transfer on
failure. Helps people with slower connections.

On failure to download, asks the user if they want to retry the download
rather than just exiting with a weird error about hash failure.

Once the image is downloaded, now changes options depending on if
squashfuse or unsquashfs works.

Prevents the user from selecting a bad option and getting unexpected
behaviour. Ideally we would do a squashfs mount test as well for
platforms that don't have working FUSE, like termux. This is harder to
get right and its for an unsupported platform, so I'm not going to
invest more time with it.

Fixes #1525
Fixes #1526
Fixes #1527
2022-01-23 22:54:05 -08:00
Ryan Houdek e4aef26ef5 FEXRootFSFetcher: Adds helper namespace for tool checking
Location to check if curl, squashfuse, and unsquashfs are working.

unsquashfs is a bit more complex where it needs to parse the help output
to see if zstd is supported
2022-01-23 22:44:31 -08:00
Ryan Houdek a41dc8eafa FEXRootFSFetcher: Fix pipe redirecting
In the case of launching without stdout/stderr then redirection could
have these constants be a redirected FD that sits in the same fd number.

Use -2 to indicate no redirection.
Use -1 to indicate closing traditional stderr/stdout
The rest will indicate if stdout and stderr should be replaced as
normal.
Making sure not to close the incoming fds if they matched the
stdout/stderr FD numbers.
2022-01-23 22:41:39 -08:00
Ryan Houdek f41cd8deff OpcodeDispatcher: Renamed GetDynamicPC to GetRelocatedPC
For clarity.
2022-01-23 18:52:53 -08:00
Ryan Houdek 2a0c3cce30 Core: Have GetDynamicPC mask based on operating size
This ensures on 32-bit we overflow correctly under relocation.
2022-01-23 18:50:16 -08:00
Ryan Houdek e817f5d98c unittests: Adds 32-bit tests for signed displacement wraparound
A bit meta since it needs to JIT some minor code but easy enough.
Ensures something like #1517 won't happen again.
2022-01-23 18:38:44 -08:00
Ryan Houdek 8b8cda9b80 JIT: Fixes signed displacement wraparound on 32-bit
This cropped up mostly with multiblock and `jmp <signed displacement>`
This also happened with non multiblock `jcc <signed displacement>`

Due to how IR relocations occur, this needs to happen fairly late but
isn't a big deal.

Fixes #1517
2022-01-23 18:38:43 -08:00
Ryan Houdek 8e3893df07 Merge pull request #1523 from lioncash/vixl-update
Externals: Update vixl
2022-01-20 15:13:36 -08:00
lioncash 51b335914c github: Synchronize submodules before checking them out
Ensures that we don't get stale remotes.
2022-01-20 17:57:45 -05:00
lioncash 8835d57ae3 Arm64Emitter: Adjust XRegister to Register
With the updated API, we need to make use of Register as opposed to
XRegister in our arrays.
2022-01-20 16:42:04 -05:00
lioncash eba1b65fb0 Externals: Update vixl to updated branch
Now we have access to some SVE goodies.
2022-01-20 16:42:02 -05:00
Ryan Houdek a3a138ef7e Merge pull request #1520 from lioncash/vixl
External: Point vixl submodule towards FEX's fork
2022-01-14 13:40:18 -08:00
lioncash 62d9a494cd External: Point vixl submodule towards FEX's fork
This allows it to be managed by all organization members
2022-01-14 12:21:19 -05:00
Stefanos Kornilios Mitsis Poiitidis 6744a06a53 Merge pull request #1519 from Sonicadvance1/aarch64_single_instruction_opt
AArch64: Single instruction optimization for AESKeyGenAssist
2022-01-14 15:41:57 +02:00
Ryan Houdek 140e9824b7 Merge pull request #1516 from lioncash/fmt
externals: Update fmt to 8.1.1
2022-01-14 01:55:50 -08:00
Ryan Houdek bad84f61fa AArch64: Single instruction optimization for AESKeyGenAssist
No need to do adr when loads can do a 1MB offset loadstore
2022-01-14 01:49:21 -08:00
lioncash 2296126af3 externals: Update fmt to 8.1.1
Brings along a bunch of enhancements and ensures we always build against
the latest version.

Also fixes up a few issues that arose due to changes in fmt
2022-01-13 14:48:35 -05:00
Stefanos Kornilios Mitsis Poiitidis 0a8717d9a8 Merge pull request #1515 from Sonicadvance1/fix_ptest
OpcodeDispatcher: Fixes ptest flags calculation.
2022-01-13 10:38:38 +02:00
Ryan Houdek 4e2220c27f unittests: Adds ptest unit test to ensure correct flag setting
ptest wasn't correctly setting OF, SF, AF, and PF to zero until now.
Do a unit test to ensure correct behaviour here
2022-01-11 16:45:56 -08:00
Ryan Houdek c87e11cee9 OpcodeDispatcher: Fixes ptest flags calculation.
We were missing four flags that require setting zero.
2022-01-11 16:45:11 -08:00
Ryan Houdek 7768f6965a Merge pull request #1501 from Sonicadvance1/finish_siginfo_32bit
Linux: Handles the remaining 32-bit siginfo_t usage
2022-01-11 00:02:55 -08:00
Ryan Houdek 023aaaae0c Merge pull request #1499 from Sonicadvance1/resolve_rootfs_path_in_interpreter
FEXLoader: Resolve the absolute path to rootfs if possible
2022-01-11 00:02:27 -08:00
Stefanos Kornilios Mitsis Poiitidis a2aa9f3fc1 Merge pull request #1512 from Sonicadvance1/fix_ssa_id_print
IR: Fixes SSA ID printing
2022-01-11 09:09:06 +02:00
Ryan Houdek e46ec9a0ce IR: Fixes SSA ID printing
These should print as decimal. They were ending up as hex
2022-01-10 18:09:03 -08:00
Ryan Houdek 784cbdd973 Merge pull request #1500 from Sonicadvance1/rootfsfetch_check_curl
FEXRootFSFetcher: Check if curl is installed and fail before running
2022-01-10 16:23:05 -08:00
Ryan Houdek 6022715a9b Merge pull request #1510 from Sonicadvance1/fix_asan_cpuid
CPUID: Fixes ASAN problem with reading midr
2022-01-10 02:12:10 -08:00
Ryan Houdek 9eb5ba5ad1 Merge pull request #1509 from Sonicadvance1/fix_logserver_sync
SocketLogging: Fixes MsgHandler not syncing with Assert level
2022-01-10 02:12:01 -08:00
Ryan Houdek 82e5977709 Merge pull request #1506 from Sonicadvance1/fix_apitest_syscalls
APITests: Fixes InterruptableConditionVariable test to use the syscal…
2022-01-10 02:11:44 -08:00
Ryan Houdek 7c08b67dff Merge pull request #1504 from Sonicadvance1/fix_unittest_rootfs_define
unittests: Fixes ROOTFS needing to be defined prior to cmake
2022-01-10 02:11:35 -08:00
Ryan Houdek 609587f9ee Merge pull request #1503 from Sonicadvance1/implement_bcd_tests
unittests: Adds a BCD unit test
2022-01-10 02:11:07 -08:00
Ryan Houdek 5dda3a1599 Merge pull request #1497 from Sonicadvance1/fix_alternative_links
Linux: Fixes emulatedpath with symlink following
2022-01-10 02:10:54 -08:00
Ryan Houdek 73aaa4c3a6 CPUID: Fixes ASAN problem with reading midr
Needs to be a string_view for the MIDR for the StrConv helper to work in
this instance.
There is no null terminator character when reading from the file is why.
2022-01-10 01:21:56 -08:00
Ryan Houdek 3ba5371d36 SocketLogging: Fixes MsgHandler not syncing with Assert level
AssertHandler by default synchronizes but MsgHandler with Assert level
should also synchronize.

Fixes an issue where LogMan::Msg::AFmt wasn't syncing so the
FEXLogServer would never see the messages.
2022-01-10 01:16:53 -08:00
Ryan Houdek bf581decde Merge pull request #1507 from Sonicadvance1/fix_warnings
Fixes some of the warnings that cropped up
2022-01-10 01:12:39 -08:00
Ryan Houdek 250504502a Fixes some of the warnings that cropped up 2022-01-10 00:46:10 -08:00
Ryan Houdek a0e826feaf APITests: Fixes InterruptableConditionVariable test to use the syscall wrappers.
Fixes a build error on old Ubuntu
2022-01-09 22:00:57 -08:00
Ryan Houdek eb17edec05 Merge pull request #1502 from Sonicadvance1/fix_fexlog_server_message
FEXLogServer: Stop duplicating and dropping messages
2022-01-09 03:22:24 -08:00
Ryan Houdek 228aed98c7 unittests: Fixes ROOTFS needing to be defined prior to cmake
cmake will bake in the environment variable in to the build scripts.
Instead have the guest_test_runner fetch it at runtime.

This means if you forget to set ROOTFS prior to running cmake, you can
now set it afterwards and rerun with just ctest instead of a cmake
dance.

Fixes #315
2022-01-09 01:56:13 -08:00
Ryan Houdek 6ba4aec88e unittests: Adds a BCD unit test
Nothing really fantastical found here. Just that sub-precision results
weren't rounded correctly on store

Fixes #770
2022-01-09 01:36:26 -08:00
Ryan Houdek b8d6b2cd4a F80: Ensures BCDStore rounds to the current rounding mode
BCD storing will round any subprecision results depending on the current
rounding mode.
2022-01-09 01:35:28 -08:00
Ryan Houdek d4e2f42f90 FEXLogServer: Stop duplicating and dropping messages
In the case that multiple messages appearing in a single packet then we
were repeating the first message and dropping any subsequent messages.

Fixes #1496
2022-01-09 00:22:12 -08:00
Ryan Houdek 3055c23365 Linux: Handles the remaining 32-bit siginfo_t usage
Just need to translate them between 32-bit and 64-bit versions.

Fixes #1254
2022-01-09 00:03:17 -08:00
Ryan Houdek cb7feaefbb Types: Allows passing 64-bit host siginfo_t to 32-bit siginfo_t
Needed for waitid
2022-01-09 00:00:13 -08:00
Ryan Houdek 5a5a498ed6 FEXRootFSFetcher: Check if curl is installed and fail before running
Before doing anything that requires curl, actually check if it is
installed.
Then instruct the user to install curl before using.

Doesn't try installing curl itself since we don't have a clean way to
execute sudo from potentially GUI.

Fixes #1498
2022-01-08 21:35:29 -08:00
Ryan Houdek 285ed8f1e0 FEXRootFSFetcher: Adds new Exec function with stdout,stderr redirection
Just so we can test for applications without spamming terminal
2022-01-08 21:34:59 -08:00
Ryan Houdek a68da468a5 FEXRootFSFetcher: ExecAndWaitForResponse sign extend program result
Only the lower 8bits of the execve result is the program result.
Makes sure to sign extend it so -1 is a true -1 instead of 255
2022-01-08 21:33:24 -08:00
Ryan Houdek d59aa6874e FEXLoader: Resolve the absolute path to rootfs if possible
If the user passes in an absolute path then check to see if it exists in
the rootfs before executing.

Useful for launching applications directly out of the rootfs with
FEXInterpreter.

In the case that the absolute path doesn't exist in the rootfs then
fallback to the host system as usual
2022-01-07 03:42:41 -08:00
Ryan Houdek 19fd89d2bf Linux: Fixes emulatedpath with symlink following
Some syscalls support `AT_SYMLINK_NOFOLLOW` In these instances we need
to follow the symlink on a couple of syscalls.

Fixes executing wine using the basic wine path
eg:
FEXBash "wine dxcapsviewer.exe"
2022-01-07 02:26:26 -08:00
Mai M e36beb8dbe Merge pull request #1495 from Sonicadvance1/add_tune_arch
CMake: Adds TUNE_ARCH option
2022-01-05 18:09:26 -05:00
Ryan Houdek 70f447b265 CMake: Adds TUNE_ARCH option
I forgot about this option working for tuning arch on AArch64. This will
be used in PPA releases in the future. Will leave the previous option
since it can be used in testing.
2022-01-05 13:48:35 -08:00
Ryan Houdek a0026c92a8 Merge pull request #1494 from Seas0/main
ThunkLibs: Add meta data to libvulkan_device
2022-01-04 23:37:54 -08:00
Seas0 8fc5f66a5b ThunkLibs: Add meta data to libvulkan_device 2022-01-05 14:28:26 +08:00
99 changed files with 3992 additions and 1259 deletions

No files matched your search

+4 -2
View File
@@ -31,7 +31,9 @@ jobs:
- name : submodule checkout
# Need to update submodules
run: git submodule update --init --depth 1
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
@@ -49,7 +51,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -1
View File
@@ -1,7 +1,7 @@
[submodule "External/vixl"]
shallow = true
path = External/vixl
url = https://github.com/Sonicadvance1/vixl.git
url = https://github.com/FEX-Emu/vixl.git
[submodule "External/cpp-optparse"]
path = External/cpp-optparse
url = https://github.com/Sonicadvance1/cpp-optparse
+18 -6
View File
@@ -19,6 +19,7 @@ option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -26,6 +27,7 @@ set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global d
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set (OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version in the format of <MMYY>{.<REV>}")
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
@@ -38,6 +40,11 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -268,13 +275,9 @@ endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
pkg_check_modules(XXHASH libxxhash>=0.8.0 QUIET)
if (NOT XXHASH_FOUND)
message(STATUS "xxHash not found. Using Externals")
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
endif()
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
@@ -335,6 +338,15 @@ if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
endif()
endif()
if (NOT TUNE_ARCH STREQUAL "generic")
check_cxx_compiler_flag("-march=${TUNE_ARCH}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
add_compile_options("-march=${TUNE_ARCH}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH}' but the compiler doesn't support this")
endif()
endif()
if (TUNE_CPU STREQUAL "native")
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
+18 -13
View File
@@ -100,19 +100,7 @@ set (SRCS
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp
Interface/Core/Interpreter/InterpreterFallbacks.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -152,6 +140,23 @@ set (SRCS
Utils/Threads.cpp
)
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp)
+4 -4
View File
@@ -23,7 +23,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} JIT_0x{:x}_{:x}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
fmt::print(fp.get(), "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
@@ -31,7 +31,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} {}_{:x}\n", HostAddr, CodeSize, Name, HostAddr);
fmt::print(fp.get(), "{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
@@ -39,7 +39,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} {}\n", HostAddr, CodeSize, Name);
fmt::print(fp.get(), "{} {:x} {}\n", HostAddr, CodeSize, Name);
}
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
@@ -47,7 +47,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} FEXJIT\n", HostAddr, CodeSize);
fmt::print(fp.get(), "{} {:x} FEXJIT\n", HostAddr, CodeSize);
}
} // namespace FEXCore
+264 -8
View File
@@ -57,35 +57,131 @@ struct X80SoftFloat {
// Ops
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
faddp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_add(lhs, rhs);
#endif
}
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fsubp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_sub(lhs, rhs);
#endif
}
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fmulp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_mul(lhs, rhs);
#endif
}
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fdivp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_div(lhs, rhs);
#endif
}
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
X80SoftFloat Rem = extF80_rem(lhs, rhs);
if (SignBit(Rem)) {
Rem = extF80_add(Rem, rhs);
}
else {
Rem.Sign = SignBit(lhs);
}
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Rem;
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem1;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
@@ -93,15 +189,47 @@ struct X80SoftFloat {
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Tmp = lhs;
Tmp.Exponent = 0x3FFF;
Tmp.Sign = lhs.Sign;
return Tmp;
#endif
}
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
int32_t TrueExp = lhs.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
#endif
}
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
@@ -112,61 +240,189 @@ struct X80SoftFloat {
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fscale; # st0 = st0 * 2^(rdint(st1))
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
return Result;
#endif
}
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
f2xm1; # st0 = 2^st(0) - 1
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
#endif
}
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st(1)
fldt %[lhs]; # st(0)
fyl2x; # st(1) * log2l(st(0))
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = Src2_d * log2l(Src1_d);
return Tmp;
#endif
}
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs];
fldt %[rhs];
fpatan;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
#endif
}
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fptan;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = tanl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsin;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = sinl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fcos;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = cosl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsqrt;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
return extF80_sqrt(lhs);
#endif
}
operator float() const {
+6 -1
View File
@@ -434,7 +434,12 @@ namespace JSON {
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
if (Core > MaxCoreNumber) {
#ifdef INTERPRETER_ENABLED
constexpr uint32_t MinCoreNumber = 0;
#else
constexpr uint32_t MinCoreNumber = 1;
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, std::to_string(FEXCore::Config::CONFIG_IRJIT));
}
@@ -535,7 +535,6 @@ uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
}
bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -50,7 +50,7 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
@@ -113,7 +113,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
+5 -2
View File
@@ -125,7 +125,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
std::string_view MIDRView(&Data.at(0), 18);
if (FEXCore::StrConv::Conv(MIDRView, &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
@@ -403,6 +404,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
uint32_t CoreCount = Cores();
// XXX: Enable once the rest of the SSE4.2 instructions are emulated
uint32_t SupportsSSE42 = CTX->HostFeatures.SupportsCRC && false ? 1 : 0;
Res.eax = FAMILY_IDENTIFIER;
@@ -432,7 +435,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(0 << 17) | // Process-context identifiers
(0 << 18) | // Prefetching from memory mapped device
(1 << 19) | // SSE4.1
(0 << 20) | // SSE4.2
(SupportsSSE42 << 20) | // SSE4.2
(0 << 21) | // X2APIC
(1 << 22) | // MOVBE
(1 << 23) | // POPCNT
+2
View File
@@ -493,9 +493,11 @@ namespace FEXCore::Context {
// Create CPU backend
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
State->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, State, CompileThread);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
State->PassManager->InsertRegisterAllocationPass(DoSRA);
+11 -12
View File
@@ -880,20 +880,19 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
case 0x38: { // F38 Table!
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->LastEscapePrefix == 0xF2) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F2;
} else if (DecodeInst->LastEscapePrefix == 0xF3) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F3;
} else if (DecodeInst->LastEscapePrefix == 0x66) {
// Operand size
Prefix = PF_38_66;
if (DecodeInst->Flags & DecodeFlags::FLAG_OPERAND_SIZE) {
Prefix |= PF_38_66;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REPNE_PREFIX) {
Prefix |= PF_38_F2;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REP_PREFIX) {
Prefix |= PF_38_F3;
}
uint16_t LocalOp = (Prefix << 8) | ReadByte();
+1 -1
View File
@@ -35,7 +35,7 @@ public:
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
private:
+2 -1
View File
@@ -970,7 +970,8 @@ GdbServer::HandledPacketType GdbServer::handleThreadOp(const std::string &packet
GdbServer::HandledPacketType GdbServer::handleBreakpoint(const std::string &packet) {
auto ss = std::istringstream(packet);
bool Set{};
// Don't do anything with set breakpoints yet
[[maybe_unused]] bool Set{};
uint64_t Addr;
uint64_t Type;
Set = ss.get() == 'Z';
@@ -53,6 +53,7 @@ HostFeatures::HostFeatures() {
#ifdef _M_ARM_64
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// Only supported when FEAT_AFP is supported
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
@@ -78,6 +79,7 @@ HostFeatures::HostFeatures() {
#ifdef _M_X86_64
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
+1
View File
@@ -15,6 +15,7 @@ class HostFeatures final {
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
bool SupportsAES{};
bool SupportsCRC{};
bool SupportsCLZERO{};
bool SupportsAtomics{};
bool SupportsRCPC{};
@@ -38,7 +38,13 @@ DEF_OP(Constant) {
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
GD = Data->CurrentEntry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
GD = (Data->CurrentEntry + Op->Offset) & Mask;
}
DEF_OP(InlineConstant) {
@@ -835,9 +841,8 @@ DEF_OP(Bfi) {
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for BFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
uint64_t SourceMask = (1ULL << Op->Width) - 1;
if (Op->Width == 64)
SourceMask = ~0ULL;
@@ -848,9 +853,8 @@ DEF_OP(Bfe) {
DEF_OP(Sbfe) {
auto Op = IROp->C<IR::IROp_Sbfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for SBFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for SBFE: {}", IROp->Size);
int64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t ShiftLeftAmount = (64 - (Op->Width + Op->lsb));
uint64_t ShiftRightAmount = ShiftLeftAmount + Op->lsb;
@@ -889,11 +893,10 @@ DEF_OP(Select) {
DEF_OP(VExtractToGPR) {
auto Op = IROp->C<IR::IROp_VExtractToGPR>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractToGPR: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractToGPR: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
@@ -298,6 +298,63 @@ namespace AES {
}
}
namespace CRC32 {
// CRC32 per byte lookup table.
constexpr std::array<uint32_t, 256> CRC32CTable = []() consteval {
std::array<uint32_t, 256> Table{};
// Clang 11.x doesn't support bitreverse as a consteval
// constexpr uint32_t Polynomial = 0x1EDC6F41;
constexpr uint32_t PolynomialRev = 0x82F63B78; //__builtin_bitreverse32(Polynomial);
for (size_t Char = 0; Char < std::size(Table); ++Char) {
uint32_t CurrentChar = Char;
for (size_t i = 0; i < 8; ++i) {
if (CurrentChar & 1) {
CurrentChar = (CurrentChar >> 1) ^ PolynomialRev;
}
else {
CurrentChar >>= 1;
}
}
Table[Char] = CurrentChar;
}
return Table;
}();
uint32_t crc32cb(uint32_t Accumulator, uint8_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ data] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32ch(uint32_t Accumulator, uint16_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cw(uint32_t Accumulator, uint32_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cx(uint32_t Accumulator, uint64_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 32) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 40) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 48) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 56) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
}
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
@@ -429,6 +486,33 @@ DEF_OP(AESKeyGenAssist) {
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
uint32_t Src1 = *GetSrc<uint32_t*>(Data->SSAData, Op->Src1);
uint8_t *Src2 = GetSrc<uint8_t*>(Data->SSAData, Op->Src2);
uint32_t Tmp{};
switch (Op->SrcSize) {
case 1:
Tmp = CRC32::crc32cb(Src1, *(uint8_t*)Src2);
break;
case 2:
Tmp = CRC32::crc32ch(Src1, *(uint16_t*)Src2);
break;
case 4:
Tmp = CRC32::crc32cw(Src1, *(uint32_t*)Src2);
break;
case 8:
Tmp = CRC32::crc32cx(Src1, *(uint64_t*)Src2);
break;
default:
LOGMAN_MSG_A_FMT("Unknown CRC32C size: {}", Op->SrcSize);
break;
}
memcpy(GDP, &Tmp, sizeof(Tmp));
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -323,7 +323,7 @@ DEF_OP(F80BCDLOAD) {
DEF_OP(F80BCDSTORE) {
auto Op = IROp->C<IR::IROp_F80BCDStore>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src1 = X80SoftFloat::FRNDINT(*GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]));
bool Negative = Src1.Sign;
// Clear the Sign bit
@@ -227,6 +227,8 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
static X80SoftFloat handle(X80SoftFloat Src1) {
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(Src1);
// Clear the Sign bit
Src1.Sign = 0;
@@ -0,0 +1,209 @@
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/F80Ops.h"
#include <cstddef>
#include <cstdint>
namespace FEXCore::CPU {
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
return {FABI_UNKNOWN, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float)) {
return {FABI_F80_F32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double)) {
return {FABI_F80_F64, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t)) {
return {FABI_F80_I16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t)) {
return {FABI_VOID_U16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t)) {
return {FABI_F80_I32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat)) {
return {FABI_F32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat)) {
return {FABI_F64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat)) {
return {FABI_I16_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat)) {
return {FABI_I32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat)) {
return {FABI_I64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_I64_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat)) {
return {FABI_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_F80_F80_F80, (void*)fn};
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags]);
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
default:
break;
}
return false;
}
}
@@ -283,6 +283,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
@@ -321,205 +322,6 @@ void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
return {FABI_UNKNOWN, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float)) {
return {FABI_F80_F32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double)) {
return {FABI_F80_F64, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t)) {
return {FABI_F80_I16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t)) {
return {FABI_VOID_U16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t)) {
return {FABI_F80_I32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat)) {
return {FABI_F32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat)) {
return {FABI_F64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat)) {
return {FABI_I16_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat)) {
return {FABI_I32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat)) {
return {FABI_I64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_I64_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat)) {
return {FABI_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_F80_F80_F80, (void*)fn};
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags]);
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
default:
break;
}
return false;
}
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
volatile void *StackEntry = alloca(0);
@@ -298,6 +298,7 @@ namespace FEXCore::CPU {
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
///< F80 ops
DEF_OP(F80LOADFCW);
@@ -1402,10 +1402,9 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractElement: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractElement: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
@@ -65,7 +65,13 @@ DEF_OP(EntrypointOffset) {
auto Constant = Entry + Op->Offset;
auto Dst = GetReg<RA_64>(Node);
LoadConstant(Dst, Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(Dst, Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -54,7 +54,7 @@ DEF_OP(AESDecLast) {
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Label Constant;
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
// Do a "regular" AESE step
@@ -63,8 +63,7 @@ DEF_OP(AESKeyGenAssist) {
aese(VTMP1.V16B(), VTMP2.V16B());
// Do a table shuffle to undo ShiftRows
adr(TMP1.X(), &Constant);
ldr(VTMP3, MemOperand(TMP1.X()));
ldr(VTMP3, &ConstantLiteral);
// Now EOR in the RCON
if (Op->RCON) {
@@ -80,14 +79,29 @@ DEF_OP(AESKeyGenAssist) {
}
b(&PastConstant);
bind(&Constant);
dc32(0x0B0E0104);
dc32(0x040B0E01);
dc32(0x0306090C);
dc32(0x0C030609);
place(&ConstantLiteral);
bind(&PastConstant);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (Op->SrcSize) {
case 1:
crc32cb(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 2:
crc32ch(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 4:
crc32cw(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 8:
crc32cx(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_64>(Op->Src2.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -97,7 +111,7 @@ void Arm64JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
+6 -1
View File
@@ -582,7 +582,12 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -305,7 +305,7 @@ private:
DEF_OP(GetHostFlag);
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
@@ -442,6 +442,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
@@ -45,7 +45,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
mov(GetDst<RA_64>(Node), Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
mov(GetDst<RA_64>(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -45,6 +45,36 @@ DEF_OP(AESKeyGenAssist) {
vaeskeygenassist(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), Op->RCON);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (IROp->Size) {
case 4:
mov(TMP1, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
break;
case 8:
mov(TMP1, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", IROp->Size);
}
switch (Op->SrcSize) {
case 1:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt8());
break;
case 2:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt16());
break;
case 4:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt32());
break;
case 8:
crc32(GetDst<RA_64>(Node).cvt64(), TMP1.cvt64());
break;
}
}
#undef DEF_OP
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -54,7 +84,7 @@ void X86JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
+6 -1
View File
@@ -553,7 +553,12 @@ bool X86JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, u
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -434,6 +434,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
+236 -240
View File
@@ -78,8 +78,11 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
LogMan::Msg::DFmt("Unhandled OSABI syscall");
}
// Calculate flags early.
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
auto NewRIP = GetDynamicPC(Op, -Op->InstSize);
auto NewRIP = GetRelocatedPC(Op, -Op->InstSize);
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), NewRIP);
const auto& GPRIndicesRef = *GPRIndexes;
@@ -114,6 +117,9 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
}
void OpDispatchBuilder::ThunkOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
const uint32_t RSPOffset = GPROffset(X86State::REG_RSP);
const uint8_t GPRSize = CTX->GetGPRSize();
uint8_t *sha256 = (uint8_t *)(Op->PC + 2);
@@ -163,6 +169,12 @@ void OpDispatchBuilder::RETOp(OpcodeArgs) {
// ABI Optimization: Flags don't survive calls or rets
if (CTX->Config.ABILocalFlags) {
_InvalidateFlags(~0UL); // all flags
// Deferred flags are invalidated now
InvalidateDeferredFlags();
}
else {
// Calculate flags early.
CalculateDeferredFlags();
}
auto Constant = _Constant(GPRSize);
@@ -203,6 +215,9 @@ void OpDispatchBuilder::IRETOp(OpcodeArgs) {
return;
}
// Calculate flags early.
CalculateDeferredFlags();
const uint32_t RSPOffset = GPROffset(X86State::REG_RSP);
const uint8_t GPRSize = CTX->GetGPRSize();
@@ -376,6 +391,9 @@ void OpDispatchBuilder::SecondaryALUOp(OpcodeArgs) {
template<uint32_t SrcIndex>
void OpDispatchBuilder::ADCOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
uint8_t Size = GetDstSize(Op);
@@ -404,6 +422,9 @@ void OpDispatchBuilder::ADCOp(OpcodeArgs) {
template<uint32_t SrcIndex>
void OpDispatchBuilder::SBBOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
auto Size = GetDstSize(Op);
@@ -706,14 +727,20 @@ void OpDispatchBuilder::CALLOp(OpcodeArgs) {
// ABI Optimization: Flags don't survive calls or rets
if (CTX->Config.ABILocalFlags) {
_InvalidateFlags(~0UL); // all flags
// Deferred flags are invalidated now
InvalidateDeferredFlags();
}
else {
// Calculate flags early.
CalculateDeferredFlags();
}
auto ConstantPC = GetDynamicPC(Op);
auto ConstantPC = GetRelocatedPC(Op);
OrderedNode *JMPPCOffset = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *NewRIP = _Add(ConstantPC, JMPPCOffset);
auto ConstantPCReturn = GetDynamicPC(Op);
auto ConstantPCReturn = GetRelocatedPC(Op);
auto ConstantSize = _Constant(GPRSize);
auto OldSP = _LoadContext(GPRSize, RSPOffset, GPRClass);
@@ -733,12 +760,15 @@ void OpDispatchBuilder::CALLAbsoluteOp(OpcodeArgs) {
const uint32_t RSPOffset = GPROffset(X86State::REG_RSP);
const uint8_t GPRSize = CTX->GetGPRSize();
// Calculate flags early.
CalculateDeferredFlags();
BlockSetRIP = true;
const uint8_t Size = GetSrcSize(Op);
OrderedNode *JMPPCOffset = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto ConstantPCReturn = GetDynamicPC(Op);
auto ConstantPCReturn = GetRelocatedPC(Op);
auto ConstantSize = _Constant(Size);
auto OldSP = _LoadContext(GPRSize, RSPOffset, GPRClass);
@@ -986,6 +1016,9 @@ OrderedNode *OpDispatchBuilder::SelectCC(uint8_t OP, OrderedNode *TrueValue, Ord
}
void OpDispatchBuilder::SETccOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
@@ -995,6 +1028,9 @@ void OpDispatchBuilder::SETccOp(OpcodeArgs) {
}
void OpDispatchBuilder::CMOVOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1004,6 +1040,9 @@ void OpDispatchBuilder::CMOVOp(OpcodeArgs) {
}
void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
BlockSetRIP = true;
auto TakeBranch = _Constant(1);
@@ -1011,9 +1050,24 @@ void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
auto SrcCond = SelectCC(Op->OP & 0xF, TakeBranch, DoNotTakeBranch);
// Jump instruction only uses up to 32-bit signed displacement
LOGMAN_THROW_A_FMT(Op->Src[0].IsLiteral(), "Src1 needs to be literal here");
int64_t TargetOffset = Op->Src[0].Data.Literal.Value;
uint64_t InstRIP = Op->PC + Op->InstSize;
uint64_t Target = InstRIP + TargetOffset;
uint64_t Target = Op->PC + Op->InstSize + Op->Src[0].Data.Literal.Value;
if (CTX->GetGPRSize() == 4) {
// If the GPRSize is 4 then we need to be careful about PC wrapping
if (TargetOffset < 0 && -TargetOffset > InstRIP) {
// Invert the signed value if we are underflowing
TargetOffset = 0x1'0000'0000ULL + TargetOffset;
}
else if (TargetOffset >= 0 && Target >= 0x1'0000'0000ULL) {
// We are overflowing, wrap around
TargetOffset = TargetOffset - 0x1'0000'0000ULL;
}
Target &= 0xFFFFFFFFU;
}
auto TrueBlock = JumpTargets.find(Target);
auto FalseBlock = JumpTargets.find(Op->PC + Op->InstSize);
@@ -1034,7 +1088,7 @@ void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
SetTrueJumpTarget(CondJump, JumpTarget);
SetCurrentCodeBlock(JumpTarget);
auto NewRIP = GetDynamicPC(Op, Op->Src[0].Data.Literal.Value);
auto NewRIP = GetRelocatedPC(Op, TargetOffset);
// Store the new RIP
_ExitFunction(NewRIP);
@@ -1052,7 +1106,7 @@ void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
SetCurrentCodeBlock(JumpTarget);
// Leave block
auto RIPTargetConst = GetDynamicPC(Op);
auto RIPTargetConst = GetRelocatedPC(Op);
// Store the new RIP
_ExitFunction(RIPTargetConst);
@@ -1061,6 +1115,9 @@ void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
}
void OpDispatchBuilder::CondJUMPRCXOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
BlockSetRIP = true;
uint8_t JcxGPRSize = CTX->GetGPRSize();
JcxGPRSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) ? (JcxGPRSize >> 1) : JcxGPRSize;
@@ -1094,7 +1151,7 @@ void OpDispatchBuilder::CondJUMPRCXOp(OpcodeArgs) {
SetTrueJumpTarget(CondJump, JumpTarget);
SetCurrentCodeBlock(JumpTarget);
auto NewRIP = GetDynamicPC(Op, Op->Src[0].Data.Literal.Value);
auto NewRIP = GetRelocatedPC(Op, Op->Src[0].Data.Literal.Value);
// Store the new RIP
_ExitFunction(NewRIP);
@@ -1112,7 +1169,7 @@ void OpDispatchBuilder::CondJUMPRCXOp(OpcodeArgs) {
SetCurrentCodeBlock(JumpTarget);
// Leave block
auto RIPTargetConst = GetDynamicPC(Op);
auto RIPTargetConst = GetRelocatedPC(Op);
// Store the new RIP
_ExitFunction(RIPTargetConst);
@@ -1121,6 +1178,9 @@ void OpDispatchBuilder::CondJUMPRCXOp(OpcodeArgs) {
}
void OpDispatchBuilder::LoopOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
bool CheckZF = Op->OP != 0xE2;
bool ZFTrue = Op->OP == 0xE1;
@@ -1170,7 +1230,7 @@ void OpDispatchBuilder::LoopOp(OpcodeArgs) {
SetTrueJumpTarget(CondJump, JumpTarget);
SetCurrentCodeBlock(JumpTarget);
auto NewRIP = GetDynamicPC(Op, Op->Src[1].Data.Literal.Value);
auto NewRIP = GetRelocatedPC(Op, Op->Src[1].Data.Literal.Value);
// Store the new RIP
_ExitFunction(NewRIP);
@@ -1188,7 +1248,7 @@ void OpDispatchBuilder::LoopOp(OpcodeArgs) {
SetCurrentCodeBlock(JumpTarget);
// Leave block
auto RIPTargetConst = GetDynamicPC(Op);
auto RIPTargetConst = GetRelocatedPC(Op);
// Store the new RIP
_ExitFunction(RIPTargetConst);
@@ -1197,15 +1257,36 @@ void OpDispatchBuilder::LoopOp(OpcodeArgs) {
}
void OpDispatchBuilder::JUMPOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
BlockSetRIP = true;
// Jump instruction only uses up to 32-bit signed displacement
LOGMAN_THROW_A_FMT(Op->Src[0].IsLiteral(), "Src1 needs to be literal here");
int64_t TargetOffset = Op->Src[0].Data.Literal.Value;
uint64_t InstRIP = Op->PC + Op->InstSize;
uint64_t TargetRIP = InstRIP + TargetOffset;
if (CTX->GetGPRSize() == 4) {
// If the GPRSize is 4 then we need to be careful about PC wrapping
if (TargetOffset < 0 && -TargetOffset > InstRIP) {
// Invert the signed value if we are underflowing
TargetOffset = 0x1'0000'0000ULL + TargetOffset;
}
else if (TargetOffset >= 0 && TargetRIP >= 0x1'0000'0000ULL) {
// We are overflowing, wrap around
TargetOffset = TargetOffset - 0x1'0000'0000ULL;
}
TargetRIP &= 0xFFFFFFFFU;
}
// This is just an unconditional relative literal jump
if (Multiblock) {
LOGMAN_THROW_A_FMT(Op->Src[0].IsLiteral(), "Src1 needs to be literal here");
uint64_t Target = Op->PC + Op->InstSize + Op->Src[0].Data.Literal.Value;
auto JumpBlock = JumpTargets.find(Target);
auto JumpBlock = JumpTargets.find(TargetRIP);
if (JumpBlock != JumpTargets.end()) {
_Jump(GetNewJumpBlock(Target));
_Jump(GetNewJumpBlock(TargetRIP));
}
else {
// If the block isn't a jump target then we need to create an exit block
@@ -1215,18 +1296,15 @@ void OpDispatchBuilder::JUMPOp(OpcodeArgs) {
auto JumpTarget = CreateNewCodeBlockAfter(GetCurrentBlock());
SetJumpTarget(Jump, JumpTarget);
SetCurrentCodeBlock(JumpTarget);
_ExitFunction(GetDynamicPC(Op, Op->Src[0].Data.Literal.Value));
_ExitFunction(GetRelocatedPC(Op, TargetOffset));
}
return;
}
// Fallback
{
// This source is a literal
auto RIPOffset = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto RIPTargetConst = GetDynamicPC(Op);
auto NewRIP = _Add(RIPOffset, RIPTargetConst);
auto RIPTargetConst = GetRelocatedPC(Op);
auto NewRIP = _Add(_Constant(TargetOffset), RIPTargetConst);
// Store the new RIP
_ExitFunction(NewRIP);
@@ -1234,6 +1312,9 @@ void OpDispatchBuilder::JUMPOp(OpcodeArgs) {
}
void OpDispatchBuilder::JUMPAbsoluteOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
BlockSetRIP = true;
// This is just an unconditional jump
// This uses ModRM to determine its location
@@ -1446,6 +1527,9 @@ void OpDispatchBuilder::FLAGControlOp(OpcodeArgs) {
break;
}
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Result{};
switch (Type) {
case OpType::Clear: {
@@ -1730,6 +1814,9 @@ void OpDispatchBuilder::SHRImmediateOp(OpcodeArgs) {
}
void OpDispatchBuilder::SHLDOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1775,6 +1862,9 @@ void OpDispatchBuilder::SHLDOp(OpcodeArgs) {
}
GenerateFlags_ShiftLeft(Op, Res, Dest, Shift);
// Calculate flags early.
CalculateDeferredFlags();
auto Jump = _Jump();
auto NextJumpTarget = CreateNewCodeBlockAfter(JumpTarget);
SetJumpTarget(Jump, NextJumpTarget);
@@ -1782,7 +1872,6 @@ void OpDispatchBuilder::SHLDOp(OpcodeArgs) {
SetCurrentCodeBlock(NextJumpTarget);
}
void OpDispatchBuilder::SHLDImmediateOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1823,6 +1912,10 @@ void OpDispatchBuilder::SHLDImmediateOp(OpcodeArgs) {
}
void OpDispatchBuilder::SHRDOp(OpcodeArgs) {
// Calculate flags early.
// This instruction conditionally generates flags so we need to insure sane state going in.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1865,6 +1958,9 @@ void OpDispatchBuilder::SHRDOp(OpcodeArgs) {
Res = _Bfe(Size, 0, Res);
}
GenerateFlags_ShiftRight(Op, Res, Dest, Shift);
// Calculate deferred flags immediately.
// This block is ending so it needs to serialize
CalculateDeferredFlags();
auto Jump = _Jump();
auto NextJumpTarget = CreateNewCodeBlockAfter(JumpTarget);
@@ -2176,31 +2272,7 @@ void OpDispatchBuilder::BEXTRBMIOp(OpcodeArgs) {
// Finally store the result.
StoreResult(GPRClass, Op, Dest, -1);
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
SetRFLAG<X86State::RFLAG_AF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_SF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_CF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_OF_LOC>(_Constant(0));
// PF
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(_Constant(0));
}
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Dest, _Constant(0),
_Constant(1), _Constant(0));
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroOp);
GenerateFlags_BEXTR(Op, Dest);
}
void OpDispatchBuilder::BLSIBMIOp(OpcodeArgs) {
@@ -2213,84 +2285,22 @@ void OpDispatchBuilder::BLSIBMIOp(OpcodeArgs) {
// ...and we're done. Painless!
StoreResult(GPRClass, Op, Result, -1);
// Now for the flags:
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
//
// AF and PF are documented as being in an undefined state after
// a BLSI operation, however, we choose to reliably clear them.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(GetSrcBitSize(Op) - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
GenerateFlags_BLSI(Op, Result);
}
void OpDispatchBuilder::BLSMSKBMIOp(OpcodeArgs) {
// Equivalent to: (Src - 1) ^ Src
auto Zero = _Constant(0);
auto One = _Constant(1);
auto* Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _Xor(_Sub(Src, One), Src);
StoreResult(GPRClass, Op, Result, -1);
// Now for the flags.
SetRFLAG<X86State::RFLAG_ZF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
GenerateFlags_BLSMSK(Op, Src);
}
void OpDispatchBuilder::BLSRBMIOp(OpcodeArgs) {
// Equivalent to: (Src - 1) & Src
auto Zero = _Constant(0);
auto One = _Constant(1);
auto* Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -2298,38 +2308,7 @@ void OpDispatchBuilder::BLSRBMIOp(OpcodeArgs) {
StoreResult(GPRClass, Op, Result, -1);
// Now for flags.
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(GetSrcBitSize(Op) - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
GenerateFlags_BLSR(Op, Result, Src);
}
void OpDispatchBuilder::BMI2Shift(OpcodeArgs) {
@@ -2387,40 +2366,7 @@ void OpDispatchBuilder::BZHI(OpcodeArgs) {
StoreResult(GPRClass, Op, Result, -1);
// Now for the flags
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_UGT,
MaskedIndex, Bounds,
One, Zero);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
GenerateFlags_BZHI(Op, Result, MaskedIndex);
}
void OpDispatchBuilder::RORX(OpcodeArgs) {
@@ -2466,6 +2412,9 @@ void OpDispatchBuilder::PEXT(OpcodeArgs) {
}
void OpDispatchBuilder::ADXOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
// Handles ADCX and ADOX
const bool IsADCX = Op->OP == 0x1F6;
@@ -2500,6 +2449,9 @@ void OpDispatchBuilder::ADXOp(OpcodeArgs) {
}
void OpDispatchBuilder::RCROp1Bit(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
const auto Size = GetSrcBitSize(Op);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2543,6 +2495,9 @@ void OpDispatchBuilder::RCROp1Bit(OpcodeArgs) {
}
void OpDispatchBuilder::RCROp8x1Bit(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
const auto Size = GetSrcBitSize(Op);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2574,6 +2529,9 @@ void OpDispatchBuilder::RCROp(OpcodeArgs) {
return;
}
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2636,6 +2594,9 @@ void OpDispatchBuilder::RCROp(OpcodeArgs) {
}
void OpDispatchBuilder::RCRSmallerOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2682,6 +2643,9 @@ void OpDispatchBuilder::RCRSmallerOp(OpcodeArgs) {
}
void OpDispatchBuilder::RCLOp1Bit(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
const auto Size = GetSrcBitSize(Op);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2714,6 +2678,9 @@ void OpDispatchBuilder::RCLOp(OpcodeArgs) {
return;
}
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2778,6 +2745,9 @@ void OpDispatchBuilder::RCLOp(OpcodeArgs) {
}
void OpDispatchBuilder::RCLSmallerOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
@@ -2843,6 +2813,9 @@ void OpDispatchBuilder::BTOp(OpcodeArgs) {
const uint32_t Size = GetDstBitSize(Op);
const uint32_t Mask = Size - 1;
// Calculate flags early.
CalculateDeferredFlags();
if (Op->Src[SrcIndex].IsGPR()) {
Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
} else {
@@ -2899,6 +2872,9 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
const uint32_t Size = GetDstBitSize(Op);
const uint32_t Mask = Size - 1;
// Calculate flags early.
CalculateDeferredFlags();
if (Op->Src[SrcIndex].IsGPR()) {
Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
} else {
@@ -2974,6 +2950,9 @@ void OpDispatchBuilder::BTSOp(OpcodeArgs) {
const uint32_t Size = GetDstBitSize(Op);
const uint32_t Mask = Size - 1;
// Calculate flags early.
CalculateDeferredFlags();
if (Op->Src[SrcIndex].IsGPR()) {
Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
} else {
@@ -3045,6 +3024,9 @@ void OpDispatchBuilder::BTCOp(OpcodeArgs) {
const uint32_t Size = GetDstBitSize(Op);
const uint32_t Mask = Size - 1;
// Calculate flags early.
CalculateDeferredFlags();
if (Op->Src[SrcIndex].IsGPR()) {
Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
} else {
@@ -3322,19 +3304,8 @@ void OpDispatchBuilder::PopcountOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
Src = _Popcount(Src);
StoreResult(GPRClass, Op, Src, -1);
// Set ZF
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(Zero);
GenerateFlags_POPCOUNT(Op, Src);
}
void OpDispatchBuilder::XLATOp(OpcodeArgs) {
@@ -3495,6 +3466,7 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
if (Size < 32) {
Result = _Bfe(Size, 0, Result);
}
GenerateFlags_SUB(Op, Result, Dest, OneConst, false);
}
@@ -3532,16 +3504,17 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
OrderedNode *TailDest = _LoadContext(GPRSize, GPROffset(X86State::REG_RDI), GPRClass);
TailDest = _Add(TailDest, PtrDir);
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RDI), TailDest);
}
else {
// Calculate deffered flags.
// This block is ending and it needs flag status
CalculateDeferredFlags();
// Create all our blocks
auto LoopHead = CreateNewCodeBlockAfter(GetCurrentBlock());
auto LoopTail = CreateNewCodeBlockAfter(LoopHead);
auto LoopEnd = CreateNewCodeBlockAfter(LoopTail);
// At the time this was written, our RA can't handle accessing nodes across blocks.
// So we need to re-load and re-calculate essential values each iteration of the loop.
@@ -3619,6 +3592,9 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
auto PtrDir = _Select(FEXCore::IR::COND_EQ, DF, _Constant(0), SizeConst, NegSizeConst);
if (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) {
// Calculate flags early. because end of block
CalculateDeferredFlags();
// Create all our blocks
auto LoopHead = CreateNewCodeBlockAfter(GetCurrentBlock());
auto LoopTail = CreateNewCodeBlockAfter(LoopHead);
@@ -3736,6 +3712,9 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RSI), Dest_RSI);
}
else {
// Calculate flags early.
CalculateDeferredFlags();
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
// read DF once
@@ -3779,6 +3758,9 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
GenerateFlags_SUB(Op, Result, Src2, Src1);
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *TailCounter = _LoadContext(GPRSize, GPROffset(X86State::REG_RCX), GPRClass);
// Decrement counter
@@ -3845,6 +3827,9 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RSI), TailDest_RSI);
}
else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
// XXX: Theoretically LODS could be optimized to
// RSI += {-}(RCX * Size)
// RAX = [RSI - Size]
@@ -3930,7 +3915,6 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
OrderedNode* Result = _Sub(Src1, Src2);
if (Size < 4)
Result = _Bfe(Size * 8, 0, Result);
GenerateFlags_SUB(Op, Result, Src1, Src2);
auto SizeConst = _Constant(Size);
@@ -3947,6 +3931,9 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RDI), TailDest_RDI);
}
else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
// read DF once
@@ -3990,6 +3977,9 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
GenerateFlags_SUB(Op, Result, Src1, Src2);
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *TailCounter = _LoadContext(GPRSize, GPROffset(X86State::REG_RCX), GPRClass);
OrderedNode *TailDest_RDI = _LoadContext(GPRSize, GPROffset(X86State::REG_RDI), GPRClass);
@@ -4221,20 +4211,15 @@ void OpDispatchBuilder::BSFOp(OpcodeArgs) {
auto Result = _FindLSB(Src);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// If Src was zero then the destination doesn't get modified
auto SelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
Dest, Result);
// ZF is set to 1 if the source was zero
auto ZFSelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
OneConst, ZeroConst);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, SelectOp, DstSize, -1);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFSelectOp);
GenerateFlags_BITSELECT(Op, Src);
}
void OpDispatchBuilder::BSROp(OpcodeArgs) {
@@ -4247,20 +4232,15 @@ void OpDispatchBuilder::BSROp(OpcodeArgs) {
auto Result = _FindMSB(Src);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// If Src was zero then the destination doesn't get modified
auto SelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
Dest, Result);
// ZF is set to 1 if the source was zero
auto ZFSelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
OneConst, ZeroConst);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, SelectOp, DstSize, -1);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFSelectOp);
GenerateFlags_BITSELECT(Op, Src);
}
void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
@@ -4396,6 +4376,9 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
}
void OpDispatchBuilder::CMPXCHGPairOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
// REX.W used to determine if it is 16byte or 8byte
// Unlike CMPXCHG, the destination can only be a memory location
@@ -4485,16 +4468,20 @@ void OpDispatchBuilder::BeginFunction(uint64_t RIP, std::vector<FEXCore::Fronten
auto Block = GetNewJumpBlock(RIP);
SetCurrentCodeBlock(Block);
IRHeader.first->Blocks = Block->Wrapped(DualListData.ListBegin());
LOGMAN_THROW_A_FMT(IsDeferredFlagsStored(), "Something failed to calculate flags and now we began with invalid state");
}
void OpDispatchBuilder::Finalize() {
// Calculate flags early.
// This usually doesn't emit any IR but in the case of hitting the block instruction limit it will
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
// Node 0 is invalid node
OrderedNode *RealNode = reinterpret_cast<OrderedNode*>(GetNode(1));
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
FEXCore::IR::IROp_Header *IROp =
#endif
[[maybe_unused]] const FEXCore::IR::IROp_Header *IROp =
RealNode->Op(DualListData.DataBegin());
LOGMAN_THROW_A_FMT(IROp->Op == OP_IRHEADER, "First op in function must be our header");
@@ -4657,7 +4644,7 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
}
else if (Operand.IsRIPRelative()) {
if (CTX->Config.Is64BitMode) {
Src = GetDynamicPC(Op, Operand.Data.RIPLiteral.Value.s);
Src = GetRelocatedPC(Op, Operand.Data.RIPLiteral.Value.s);
}
else {
// 32bit this isn't RIP relative but instead absolute
@@ -4732,7 +4719,7 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
return Src;
}
OrderedNode *OpDispatchBuilder::GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset) {
OrderedNode *OpDispatchBuilder::GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset) {
const uint8_t GPRSize = CTX->GetGPRSize();
return _EntrypointOffset(Op->PC + Op->InstSize + Offset - Entry, GPRSize);
}
@@ -4801,7 +4788,7 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
}
else if (Operand.IsRIPRelative()) {
if (CTX->Config.Is64BitMode) {
MemStoreDst = GetDynamicPC(Op, Operand.Data.RIPLiteral.Value.s);
MemStoreDst = GetRelocatedPC(Op, Operand.Data.RIPLiteral.Value.s);
}
else {
// 32bit this isn't RIP relative but instead absolute
@@ -5074,6 +5061,9 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
break;
}
// Calculate flags early.
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
if (setRIP) {
@@ -5081,7 +5071,7 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
BlockSetRIP = setRIP;
// We want to set RIP to the next instruction after HLT/INT3
auto NewRIP = GetDynamicPC(Op);
auto NewRIP = GetRelocatedPC(Op);
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), NewRIP);
}
@@ -5094,7 +5084,7 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
SetFalseJumpTarget(CondJump, FalseBlock);
SetCurrentCodeBlock(FalseBlock);
auto NewRIP = GetDynamicPC(Op);
auto NewRIP = GetRelocatedPC(Op);
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), NewRIP);
_Break(Reason, Literal);
@@ -5114,14 +5104,7 @@ void OpDispatchBuilder::TZCNT(OpcodeArgs) {
Src = _FindTrailingZeros(Src);
StoreResult(GPRClass, Op, Src, -1);
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, 0, Src));
GenerateFlags_TZCNT(Op, Src);
}
void OpDispatchBuilder::LZCNT(OpcodeArgs) {
@@ -5129,15 +5112,7 @@ void OpDispatchBuilder::LZCNT(OpcodeArgs) {
auto Res = _CountLeadingZeroes(Src);
StoreResult(GPRClass, Op, Res, -1);
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, GetSrcBitSize(Op) - 1, Src));
GenerateFlags_LZCNT(Op, Src);
}
void OpDispatchBuilder::MOVBEOp(OpcodeArgs) {
@@ -5190,12 +5165,26 @@ void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RDX), CounterHigh);
}
void OpDispatchBuilder::CRC32(OpcodeArgs) {
// Destination GPR size is always 4 or 8 bytes depending on widening
uint8_t DstSize = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REX_WIDENING ? 8 : 4;
OrderedNode *Dest = LoadSource_WithOpSize(GPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
// Incoming memory is 8, 16, 32, or 64
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, 1);
auto Result = _CRC32(Dest, Src, GetSrcSize(Op));
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, DstSize, -1);
}
void OpDispatchBuilder::UnimplementedOp(OpcodeArgs) {
// Ensure flags are calculated on invalid op.
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
// We don't actually support this instruction
// Multiblock may hit it though
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), GetDynamicPC(Op, -Op->InstSize));
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), GetRelocatedPC(Op, -Op->InstSize));
_Break(FEXCore::IR::Break_Unimplemented, 0);
BlockSetRIP = true;
@@ -5204,11 +5193,14 @@ void OpDispatchBuilder::UnimplementedOp(OpcodeArgs) {
}
void OpDispatchBuilder::InvalidOp(OpcodeArgs) {
// Ensure flags are calculated on invalid op.
CalculateDeferredFlags();
const uint8_t GPRSize = CTX->GetGPRSize();
// We don't actually support this instruction
// Multiblock may hit it though
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), GetDynamicPC(Op, -Op->InstSize));
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), GetRelocatedPC(Op, -Op->InstSize));
_Break(FEXCore::IR::Break_InvalidInstruction, 0);
BlockSetRIP = true;
}
@@ -6079,11 +6071,11 @@ constexpr uint16_t PF_F2 = 3;
#undef OPD
#undef OPDReg
#define OPD(prefix, opcode) ((prefix << 8) | opcode)
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
@@ -6136,6 +6128,7 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<2, 4, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<2, 8, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<4, 8, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 8>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMIN, 1>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMIN, 4>},
{OPD(PF_38_66, 0x3A), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VUMIN, 2>},
@@ -6156,8 +6149,11 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_66, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66, 0xF6), 1, &OpDispatchBuilder::ADXOp},
{OPD(PF_38_F3, 0xF6), 1, &OpDispatchBuilder::ADXOp},
+565 -20
View File
@@ -14,6 +14,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <fmt/format.h>
#include <map>
#include <stddef.h>
#include <utility>
@@ -35,6 +36,37 @@ enum class SelectionFlag {
};
public:
enum class FlagsGenerationType : uint8_t {
TYPE_NONE,
TYPE_ADC,
TYPE_SBB,
TYPE_SUB,
TYPE_ADD,
TYPE_MUL,
TYPE_UMUL,
TYPE_LOGICAL,
TYPE_LSHL,
TYPE_LSHLI,
TYPE_LSHR,
TYPE_LSHRI,
TYPE_ASHR,
TYPE_ASHRI,
TYPE_ROR,
TYPE_RORI,
TYPE_ROL,
TYPE_ROLI,
TYPE_FCMP,
TYPE_BEXTR,
TYPE_BLSI,
TYPE_BLSMSK,
TYPE_BLSR,
TYPE_POPCOUNT,
TYPE_BZHI,
TYPE_TZCNT,
TYPE_LZCNT,
TYPE_BITSELECT,
};
SelectionFlag flagsOp{};
uint8_t flagsOpSize{};
OrderedNode* flagsOpDest{};
@@ -87,6 +119,9 @@ public:
// cmp qword [rdi-8], 0
// jne .label
if (LastOp && !BlockSetRIP) {
// Calculate flags first
CalculateDeferredFlags();
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end()) {
@@ -101,6 +136,11 @@ public:
return true;
}
}
if (LastOp) {
LOGMAN_THROW_A_FMT(IsDeferredFlagsStored(), "FinishOp: Deferred flags weren't generated at end of block");
}
BlockSetRIP = false;
return false;
@@ -500,6 +540,8 @@ public:
void MPSADBWOp(OpcodeArgs);
void CRC32(OpcodeArgs);
void UnimplementedOp(OpcodeArgs);
void InvalidOp(OpcodeArgs);
@@ -519,7 +561,7 @@ private:
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
@@ -557,23 +599,515 @@ private:
OrderedNode *SelectCC(uint8_t OP, OrderedNode *TrueValue, OrderedNode *FalseValue);
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High);
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High);
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
/**
* @name Deferred RFLAG calculation and generation.
*
* Only handles the six flags that ALU ops typically generate.
* Specifically: CF, PF, AF, ZF, SF, OF
* These six flags are heavily generated through basic ALU ops and balloon the IR if not early eliminated.
* This tracking structure only tracks single blocks and requires RFLAGS calculation at block-ending ops.
* Some flags generating ALU ops only touch part of the registers, In these cases it will do calculation up front.
* This means we still need our IR passes to eliminate all redundant flags accesses but this light OpcodeDispatcher optimization
* doesn't take it to that level.
* @{ */
// Deferred flag generation tracking structure.
// This structure is used to track RFlags from ALU ops for invalidation.
//
// Future ideas: Use an invalidation mask to do partial generation of flags.
// Particularly for the instructions that don't do the full set of flags calculations.
// These instructions currently calculate the deferred RFLAGS immediately then overwrite rflags state.
// RCLSE IR pass will catch and remove redundant rflags stores like this currently.
struct DeferredFlagData {
// What type of flags to generate
FlagsGenerationType Type {FlagsGenerationType::TYPE_NONE};
// Source size of the op
uint8_t SrcSize;
// Every flag generation type has a result
OrderedNode *Res{};
union {
// UMUL, BEXTR, BLSI, BLSMSK, POPCOUNT, TZCNT, LZCNT, BITSELECT
struct {
} NoSource;
// MUL, BLSR, BZHI
struct {
OrderedNode *Src1;
} OneSource;
// Logical, LSHL, LSHR, ASHR, ROR, ROL
struct {
OrderedNode *Src1;
OrderedNode *Src2;
} TwoSource;
// ADC, SBB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
OrderedNode *Src3;
} ThreeSource;
// LSHLI, LSHRI, ASHRI, RORI, ROLI
struct {
OrderedNode *Src1;
uint64_t Imm;
} OneSrcImmediate;
// ADD, SUB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
bool UpdateCF;
} TwoSrcImmediate;
} Sources{};
};
DeferredFlagData CurrentDeferredFlags{};
/**
* @brief Takes the current deferred flag state and stores the result in to RFLAGS.
*
* Once executed there will no longer be any deferred flag state and RFLAGS will have the correct flags in it.
* Necessary to do when leaving a IR block, or if an instruction is doing a partial overwrite of the flags.
*/
void CalculateDeferredFlags(uint32_t FlagsToCalculateMask = ~0U);
/**
* @brief Invalidates the current deferred flags structure.
*
* If the emulated instruction is going to overwrite all of the flags but isn't tracked using the deferred flag system
* then use this function to stop tracking the current active deferred flags.
*/
void InvalidateDeferredFlags() {
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
/**
* @brief Checks if there is any deferred flag state active.
*
* @return True if RFLAGs contains the flags. False if deferred flags is tracking the data.
*/
bool IsDeferredFlagsStored() const {
return CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE;
}
/**
* @name These functions are used by the deferred flag handling while it is calculating and storing flags in to RFLAGs.
* @{ */
void CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High);
void CalculcateFlags_UMUL(OrderedNode *High);
void CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_BEXTR(OrderedNode *Src);
void CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BLSMSK(OrderedNode *Src);
void CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src);
void CalculcateFlags_POPCOUNT(OrderedNode *Src);
void CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src);
void CalculcateFlags_TZCNT(OrderedNode *Src);
void CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BITSELECT(OrderedNode *Src);
/** @} */
/**
* @name These functions generated deferred RFLAGs tracking.
*
* Depending on the operation it may force a RFLAGs calculation before storing the new deferred state.
* @{ */
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADC,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SBB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SUB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADD,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_MUL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = High,
},
},
};
}
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_UMUL,
.SrcSize = GetSrcSize(Op),
.Res = High,
};
}
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LOGICAL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_RORI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
}
};
}
void GenerateFlags_FCMP(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_FCMP,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
}
};
}
void GenerateFlags_BEXTR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BEXTR,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSI,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSMSK(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSMSK,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_POPCOUNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_POPCOUNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BZHI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Result, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BZHI,
.SrcSize = GetSrcSize(Op),
.Res = Result,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_TZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_TZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_LZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BITSELECT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BITSELECT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
/** @} */
/** @} */
OrderedNode * GetX87Top();
enum class X87Tag {
@@ -613,11 +1147,22 @@ private:
else
return _LoadMem(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
};
void InstallOpcodeHandlers(Context::OperatingMode Mode);
}
template <>
struct fmt::formatter<FEXCore::IR::OpDispatchBuilder::FlagsGenerationType> : fmt::formatter<int> {
using Base = fmt::formatter<int>;
// Pass-through the underlying value, so IDs can
// be formatted like any integral value.
template <typename FormatContext>
auto format(const FEXCore::IR::OpDispatchBuilder::FlagsGenerationType& ID, FormatContext& ctx) {
return Base::format(static_cast<int>(ID), ctx);
}
};
@@ -41,8 +41,17 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
// Calculate flags early.
// Could use InvalidateDeferredFlags() if we had masked invalidation.
// This is only a partial overwrite of flags since OF isn't stored here.
CalculateDeferredFlags();
NumFlags = 5;
}
else {
// We are overwriting all RFLAGS. Invalidate the deferred flag state.
InvalidateDeferredFlags();
}
auto OneConst = _Constant(1);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
@@ -52,6 +61,9 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
}
OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Original = _Constant(2);
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
@@ -68,8 +80,185 @@ OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
return Original;
}
void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
if (CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE) {
// Nothing to do
return;
}
switch (CurrentDeferredFlags.Type) {
case FlagsGenerationType::TYPE_ADC:
CalculcateFlags_ADC(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SBB:
CalculcateFlags_SBB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SUB:
CalculcateFlags_SUB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_ADD:
CalculcateFlags_ADD(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_MUL:
CalculcateFlags_MUL(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_UMUL:
CalculcateFlags_UMUL(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LOGICAL:
CalculcateFlags_Logical(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHL:
CalculcateFlags_ShiftLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHLI:
CalculcateFlags_ShiftLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_LSHR:
CalculcateFlags_ShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHRI:
CalculcateFlags_ShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ASHR:
CalculcateFlags_SignShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ASHRI:
CalculcateFlags_SignShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROR:
CalculcateFlags_RotateRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_RORI:
CalculcateFlags_RotateRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROL:
CalculcateFlags_RotateLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ROLI:
CalculcateFlags_RotateLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_FCMP:
CalculcateFlags_FCMP(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_BEXTR:
CalculcateFlags_BEXTR(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSI:
CalculcateFlags_BLSI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSMSK:
CalculcateFlags_BLSMSK(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSR:
CalculcateFlags_BLSR(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_POPCOUNT:
CalculcateFlags_POPCOUNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BZHI:
CalculcateFlags_BZHI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_TZCNT:
CalculcateFlags_TZCNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LZCNT:
CalculcateFlags_LZCNT(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BITSELECT:
CalculcateFlags_BITSELECT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_NONE:
default: ERROR_AND_DIE_FMT("Unhandled flags type {}", CurrentDeferredFlags.Type);
}
// Done calculating
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = SrcSize * 8;
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -79,7 +268,7 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(Size - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -140,9 +329,7 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
const auto SrcSize = GetSrcSize(Op);
void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -212,7 +399,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
void OpDispatchBuilder::CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -222,7 +409,7 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -263,15 +450,13 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *FinalAnd = _And(XorOp1, XorOp2);
FinalAnd = _Bfe(1, GetSrcSize(Op) * 8 - 1, FinalAnd);
FinalAnd = _Bfe(1, SrcSize * 8 - 1, FinalAnd);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(FinalAnd);
}
}
void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
const auto SrcSize = GetSrcSize(Op);
void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -339,7 +524,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High) {
// PF/AF/ZF/SF
// Undefined
{
@@ -354,7 +539,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
// CF and OF are set if the result of the operation can't be fit in to the destination register
// If the value can fit then the top bits will be zero
auto SignBit = _Sbfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto SignBit = _Sbfe(1, SrcSize * 8 - 1, Res);
auto SelectOp = _Select(FEXCore::IR::COND_EQ, High, SignBit, _Constant(0), _Constant(1));
@@ -363,7 +548,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_UMUL(OrderedNode *High) {
// AF/SF/PF/ZF
// Undefined
{
@@ -385,7 +570,7 @@ void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, Ord
}
}
void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// AF
{
// Undefined
@@ -395,7 +580,7 @@ void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op,
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -430,11 +615,11 @@ auto oldflag = GetRFLAG(FEXCore::X86State::flag);\
auto newval = _Select(FEXCore::IR::COND_EQ, cond, _Constant(0), oldflag, newflag);\
SetRFLAG<FEXCore::X86State::flag>(newval);
void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
auto Size = _Constant(GetSrcSize(Op) * 8);
auto Size = _Constant(SrcSize * 8);
auto ShiftAmt = _Sub(Size, Src2);
auto LastBit = _And(_Lshr(Src1, ShiftAmt), _Constant(1));
COND_FLAG_SET(Src2, RFLAG_CF_LOC, LastBit);
@@ -466,7 +651,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
// SF
{
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val = _Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -474,12 +659,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
{
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
// When Shift > 1 then OF is undefined
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -514,7 +699,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
// SF
{
auto val =_Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val =_Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -522,12 +707,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
{
// Only defined when Shift is 1 else undefined
// OF flag is set if a sign change occurred
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -562,7 +747,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, LshrOp);
@@ -574,14 +759,14 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
}
}
void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
// CF
{
// Extract the last bit shifted in to CF
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - Shift, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, SrcSize * 8 - Shift, Src1));
}
// PF
@@ -610,20 +795,20 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::Dec
// SF
{
auto LshrOp = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto LshrOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
// OF
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto SourceBit = _Bfe(1, GetSrcSize(Op) * 8 - 1, Src1);
auto SourceBit = _Bfe(1, SrcSize * 8 - 1, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, LshrOp));
}
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -659,7 +844,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -674,7 +859,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -710,7 +895,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -721,13 +906,13 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// Only defined when Shift is 1 else undefined
// Is set to the MSB of the original value
if (Shift == 1) {
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - 1, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src1));
}
}
}
void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
auto NewCF = _Bfe(1, OpSize - 1, Res);
@@ -755,8 +940,8 @@ void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
//auto Size = _Constant(GetSrcSize(Res) * 8);
@@ -785,10 +970,10 @@ void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp O
}
}
void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
auto NewCF = _Bfe(1, OpSize - Shift, Src1);
@@ -807,10 +992,10 @@ void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::D
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
// CF
{
@@ -827,4 +1012,251 @@ void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::De
}
}
void OpDispatchBuilder::CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
}
void OpDispatchBuilder::CalculcateFlags_BEXTR(OrderedNode *Src) {
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// CF and OF are defined as being set to zero
//
SetRFLAG<X86State::RFLAG_CF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_OF_LOC>(_Constant(0));
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
SetRFLAG<X86State::RFLAG_AF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_SF_LOC>(_Constant(0));
// PF
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(_Constant(0));
}
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Src, _Constant(0),
_Constant(1), _Constant(0));
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src) {
// Now for the flags:
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
//
// AF and PF are documented as being in an undefined state after
// a BLSI operation, however, we choose to reliably clear them.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Src, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Src, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_BLSMSK(OrderedNode *Src) {
// Now for the flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_ZF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_POPCOUNT(OrderedNode *Src) {
// Set ZF
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(Zero);
}
void OpDispatchBuilder::CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for the flags
auto Bounds = _Constant(SrcSize * 8- 1);
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_UGT,
Src, Bounds,
One, Zero);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_TZCNT(OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, 0, Src));
}
void OpDispatchBuilder::CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src));
}
void OpDispatchBuilder::CalculcateFlags_BITSELECT(OrderedNode *Src) {
// OF, SF, AF, PF, CF all undefined
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// ZF is set to 1 if the source was zero
auto ZFSelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
OneConst, ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFSelectOp);
}
}
@@ -241,6 +241,8 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 2>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 8>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 2>(OpcodeArgs);
@@ -1187,13 +1189,6 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
OrderedNode *Src2{};
if constexpr (Scalar) {
Src2 = _VExtractElement(GetDstSize(Op), Size, Dest, 0);
}
else {
Src2 = Dest;
}
uint8_t CompType = Op->Src[1].Data.Literal.Value;
OrderedNode *Result{};
@@ -1201,30 +1196,30 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
//auto ALUOp = _VCMPGT(Size, ElementSize, Dest, Src);
switch (CompType) {
case 0x00: case 0x08: case 0x10: case 0x18: // EQ
Result = _VFCMPEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPEQ(Size, ElementSize, Dest, Src);
break;
case 0x01: case 0x09: case 0x11: case 0x19: // LT, GT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
break;
case 0x02: case 0x0A: case 0x12: case 0x1A: // LE, GE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
break;
case 0x03: case 0x0B: case 0x13: case 0x1B: // Unordered
Result = _VFCMPUNO(Size, ElementSize, Src2, Src);
Result = _VFCMPUNO(Size, ElementSize, Dest, Src);
break;
case 0x04: case 0x0C: case 0x14: case 0x1C: // NEQ
Result = _VFCMPNEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPNEQ(Size, ElementSize, Dest, Src);
break;
case 0x05: case 0x0D: case 0x15: case 0x1D: // NLT, NGT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x06: case 0x0E: case 0x16: case 0x1E: // NLE, NGE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x07: case 0x0F: case 0x17: case 0x1F: // Ordered
Result = _VFCMPORD(Size, ElementSize, Src2, Src);
Result = _VFCMPORD(Size, ElementSize, Dest, Src);
break;
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
@@ -1432,18 +1427,7 @@ void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
(1 << FCMP_FLAG_LT) |
(1 << FCMP_FLAG_UNORDERED));
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
GenerateFlags_FCMP(Op, Res, Src1, Src2);
flagsOp = SelectionFlag::FCMP;
flagsOpDest = Src1;
@@ -2205,6 +2189,9 @@ template
void OpDispatchBuilder::VectorVariableBlend<8>(OpcodeArgs);
void OpDispatchBuilder::PTestOp(OpcodeArgs) {
// Invalidate deferred flags early
InvalidateDeferredFlags();
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -2233,8 +2220,14 @@ void OpDispatchBuilder::PTestOp(OpcodeArgs) {
Test2 = _Select(FEXCore::IR::COND_EQ,
Test2, ZeroConst, OneConst, ZeroConst);
// Careful, these flags are different between {V,}PTEST and VTESTP{S,D}
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(Test1);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Test2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(ZeroConst);
}
void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
@@ -685,12 +685,15 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
}
else {
// Invalidate deferred flags early
// OF, SF, AF, PF all undefined
InvalidateDeferredFlags();
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
}
if constexpr (poptwice) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, X87Tag::Empty);
@@ -14,11 +14,11 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeH0F38Tables() {
#define OPD(prefix, opcode) ((prefix << 8) | opcode)
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
static constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -74,6 +74,7 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x37), 1, X86InstInfo{"PCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -98,8 +99,10 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_66, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66, 0xF6), 1, X86InstInfo{"ADCX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F3, 0xF6), 1, X86InstInfo{"ADOX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
+23 -2
View File
@@ -306,7 +306,9 @@
},
"EntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"Desc": ["Returns the <entrypoint> + Offset address",
"When the size is 4 bytes then 32-bit overflow and underflow needs to work"
],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
@@ -320,7 +322,9 @@
},
"InlineEntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"Desc": ["Returns the <entrypoint> + Offset address",
"When the size is 4 bytes then 32-bit overflow and underflow needs to work"
],
"OpClass": "ALU",
"DestSize": "RegisterSize",
"Args": [
@@ -3551,6 +3555,23 @@
]
},
"CRC32": {
"Desc": ["CRC32 using polynomial 0x1EDC6F41"
],
"OpClass": "Crypto",
"HasDest": true,
"DestSize": "std::max<uint8_t>(4, GetOpSize(ssa0))",
"DestClass": "GPR",
"SSAArgs": "2",
"SSANames": [
"Src1",
"Src2"
],
"Args": [
"uint8_t", "SrcSize"
]
},
"GetHostFlag": {
"OpClass": "Flags",
"HasDest": true,
+2 -2
View File
@@ -225,7 +225,7 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
NumElements /= ElementSize;
}
*out << "%ssa" << ID;
*out << "%ssa" << std::dec << ID;
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(ID);
@@ -265,7 +265,7 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
NumElements = IROp->Size / ElementSize;
}
*out << "(%ssa" << ID << ' ';
*out << "(%ssa" << std::dec << ID << ' ';
*out << 'i' << std::dec << (ElementSize * 8);
if (NumElements > 1) {
*out << 'v' << std::dec << NumElements;
@@ -47,10 +47,12 @@ namespace {
return (Type & ACCESS_TYPE_MASK) == ACCESS_READ;
}
[[maybe_unused]]
static bool IsInvalidAccess(LastAccessType Type) {
return (Type & ACCESS_TYPE_MASK) == ACCESS_INVALID;
}
[[maybe_unused]]
static bool IsPartialAccess(LastAccessType Type) {
return (Type & ACCESS_PARTIAL) == ACCESS_PARTIAL;
}
@@ -254,7 +256,7 @@ namespace {
});
size_t ClassifiedStructSize{};
[[maybe_unused]] size_t ClassifiedStructSize{};
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
for (auto &it : *ContextClassification) {
LOGMAN_THROW_A_FMT(it.Class.Offset == ContextClassificationInfo->Lookup.size(), "Offset mismatch (offset={})", it.Class.Offset);
@@ -77,7 +77,7 @@ struct FPRInfo {
bool IsFPR(uint32_t Offset) {
auto begin = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto end = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[17][0]);
auto end = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[16][0]);
if (Offset < begin || Offset >= end)
return false;
@@ -570,10 +570,10 @@ namespace {
// Get SRA Reg and Class from a Context offset
auto GetRegAndClassFromOffset = [](uint32_t Offset) {
auto beginGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0]);
auto endGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[17]);
auto endGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[16]);
auto beginFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[17][0]);
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[16][0]);
if (Offset >= beginGpr && Offset < endGpr) {
auto reg = (Offset - beginGpr) / 8;
@@ -594,10 +594,10 @@ namespace {
// Get a StaticMap entry from context offset
const auto GetStaticMapFromOffset = [&](uint32_t Offset) -> LiveRange** {
auto beginGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0]);
auto endGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[17]);
auto endGpr = offsetof(FEXCore::Core::CpuStateFrame, State.gregs[16]);
auto beginFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[17][0]);
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[16][0]);
if (Offset >= beginGpr && Offset < endGpr) {
auto reg = (Offset - beginGpr) / 8;
@@ -1115,8 +1115,6 @@ namespace {
auto InterferenceNodeOpBeginIter = IR.at(InterferenceLiveRange->Begin);
auto InterferenceNodeOpEndIter = IR.at(InterferenceLiveRange->End);
bool Found{};
// If the nodes live range is entirely encompassed by the interference node's range
// then spilling that range will /potentially/ lower RA
// Will only lower register pressure if the interference node does NOT have a use inside of
@@ -1153,7 +1151,6 @@ namespace {
const auto NextUseDistance = InterferenceNodeNextUse.ID().Value - CurrentLocation.Value;
if (NextUseDistance >= InterferenceFarthestNextUse) {
Found = true;
InterferenceIdToSpill = InterferenceNode;
InterferenceFarthestNextUse = NextUseDistance;
}
@@ -25,7 +25,7 @@ public:
bool IsStaticAllocGpr(uint32_t Offset, RegisterClassType Class) {
const auto begin = offsetof(FEXCore::Core::CPUState, gregs[0]);
const auto end = offsetof(FEXCore::Core::CPUState, gregs[17]);
const auto end = offsetof(FEXCore::Core::CPUState, gregs[16]);
if (Offset >= begin && Offset < end) {
const auto reg = (Offset - begin) / 8;
@@ -40,7 +40,7 @@ bool IsStaticAllocGpr(uint32_t Offset, RegisterClassType Class) {
bool IsStaticAllocFpr(uint32_t Offset, RegisterClassType Class, bool AllowGpr) {
const auto begin = offsetof(FEXCore::Core::CPUState, xmm[0][0]);
const auto end = offsetof(FEXCore::Core::CPUState, xmm[17][0]);
const auto end = offsetof(FEXCore::Core::CPUState, xmm[16][0]);
if (Offset >= begin && Offset < end) {
const auto reg = (Offset - begin) / 16;
+5 -4
View File
@@ -4,6 +4,7 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
@@ -183,7 +184,7 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
size_t NumberOfPages = length / PAGE_SIZE;
// This needs a mutex to be thread safe
std::scoped_lock<std::mutex> lk{AllocationMutex};
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
uint64_t AllocatedOffset{};
LiveVMARegion *LiveRegion{};
@@ -442,7 +443,7 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
}
// This needs a mutex to be thread safe
std::scoped_lock<std::mutex> lk{AllocationMutex};
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
length = FEXCore::AlignUp(length, PAGE_SIZE);
@@ -621,8 +622,8 @@ OSAllocator_64Bit::OSAllocator_64Bit() {
}
OSAllocator_64Bit::~OSAllocator_64Bit() {
// For consistency, pull the mutex
std::scoped_lock<std::mutex> lk{AllocationMutex};
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
// Walk the pages and deallocate
// First walk the live regions
+10
View File
@@ -172,6 +172,16 @@ namespace FEXCore {
return val;
}
siginfo_t(::siginfo_t val) {
si_signo = val.si_signo;
si_errno = val.si_errno;
si_code = val.si_code;
// Copy over the union
// The union is different sizes on 64-bit versus 32-bit
memcpy(val._sifields._pad, _sifields.pad, std::min(sizeof(val._sifields._pad), sizeof(_sifields.pad)));
}
};
static_assert(sizeof(FEXCore::x86::siginfo_t) == 128, "This needs to be the right size");
+11 -1
View File
@@ -7,6 +7,8 @@
#include <cstring>
#include <type_traits>
#include <fmt/format.h>
namespace FEXCore::IR {
///< Forward declaration of OpDispatchBuilder
class OpDispatchBuilder;
@@ -466,7 +468,7 @@ constexpr size_t MAX_INST_GROUP_TABLE_SIZE = 512;
constexpr size_t MAX_INST_SECOND_GROUP_TABLE_SIZE = 512;
constexpr size_t MAX_X87_TABLE_SIZE = 1 << 11;
constexpr size_t MAX_SECOND_MODRM_TABLE_SIZE = 32;
// 3 prefixes | 8 bit opcode
// (3 bit prefixes) | 8 bit opcode
constexpr size_t MAX_0F_38_TABLE_SIZE = (1 << 11);
// 1 REX | 1 prefixes | 8 bit opcode
constexpr size_t MAX_0F_3A_TABLE_SIZE = (1 << 11);
@@ -513,3 +515,11 @@ extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE>
// EVEX
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> EVEXTableOps;
}
template <>
struct fmt::formatter<FEXCore::X86Tables::DecodedOperand::OpType> : formatter<uint32_t> {
template <typename FormatContext>
auto format(FEXCore::X86Tables::DecodedOperand::OpType type, FormatContext& ctx) const {
return fmt::formatter<uint32_t>::format(static_cast<uint32_t>(type), ctx);
}
};
+1 -1
+1 -1
@@ -0,0 +1,46 @@
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <mutex>
#include <signal.h>
#include <sys/syscall.h>
#include <unistd.h>
namespace FHU {
/**
* @brief A class that masks signals and locks a mutex until it goes out of scope
*
* Constructor order:
* 1) Mask signals
* 2) Lock Mutex
*
* Destructor Order:
* 1) Unlock Mutex
* 2) Unmask signals
*/
class ScopedSignalMaskWithMutex final {
public:
ScopedSignalMaskWithMutex(std::mutex &_Mutex, uint64_t Mask = ~0ULL)
: Mutex {_Mutex} {
// Mask all signals, storing the original incoming mask
::syscall(SYS_rt_sigprocmask, SIG_SETMASK, &Mask, &OriginalMask, sizeof(OriginalMask));
// Lock the mutex
Mutex.lock();
}
~ScopedSignalMaskWithMutex() {
// Unlock the mutex
Mutex.unlock();
// Unmask back to the original signal mask
::syscall(SYS_rt_sigprocmask, SIG_SETMASK, &OriginalMask, nullptr, sizeof(OriginalMask));
}
private:
uint64_t OriginalMask{};
std::mutex &Mutex;
};
}
+11 -2
View File
@@ -1,4 +1,5 @@
#!/usr/bin/python3
import os
import sys
import subprocess
@@ -13,6 +14,7 @@ known_failures_file = sys.argv[1]
expected_output_file = sys.argv[2]
disabled_tests_file = sys.argv[3]
test_name = sys.argv[4]
fexecutable = sys.argv[5]
known_failures = { }
expected_output = { }
@@ -42,9 +44,16 @@ with open(disabled_tests_file) as dtf:
# run with timeout to avoid locking up
RunnerArgs = []
RunnerArgs.append(fexecutable)
ROOTFS_ENV = os.getenv("ROOTFS")
if ROOTFS_ENV != None:
RunnerArgs.append("-R")
RunnerArgs.append(ROOTFS_ENV)
# Add the rest of the arguments
for i in range(len(sys.argv) - 5):
RunnerArgs.append(sys.argv[5 + i])
for i in range(len(sys.argv) - 6):
RunnerArgs.append(sys.argv[6 + i])
#print(RunnerArgs)
+2 -2
View File
@@ -108,7 +108,7 @@ namespace FEX::SocketLogging {
static std::unique_ptr<ClientConnector> Client{};
void MsgHandler(LogMan::DebugLevels Level, char const *Message) {
Client->MsgHandler(Level, false, Message);
Client->MsgHandler(Level, Level == LogMan::DebugLevels::ASSERT, Message);
}
void AssertHandler(char const *Message) {
@@ -333,7 +333,7 @@ namespace FEX::SocketLogging {
while (CurrentOffset < CurrentRead) {
Common::PacketHeader *Header = reinterpret_cast<Common::PacketHeader*>(&Data[CurrentOffset]);
if (Header->PacketType == Common::PacketTypes::TYPE_MSG) {
Common::PacketMsg *Msg = reinterpret_cast<Common::PacketMsg*>(&Data[0]);
Common::PacketMsg *Msg = reinterpret_cast<Common::PacketMsg*>(&Data[CurrentOffset]);
MsgHandler(Socket, Msg->Header.Timestamp, Msg->Header.PID, Msg->Header.TID, Msg->Level, Msg->Msg);
CurrentOffset += sizeof(Common::PacketMsg) + strlen(Msg->Msg) + 1;
+5
View File
@@ -1,7 +1,12 @@
#pragma once
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <functional>
#include <memory>
#include <string>
namespace FEX::SocketLogging {
// Client side
namespace Client {
-4
View File
@@ -877,7 +877,6 @@ void ELFContainer::PrintRelocationTable() const {
}
else {
Elf64_Shdr const *RelaHeader{nullptr};
Elf64_Shdr const *GOTHeader {nullptr};
Elf64_Shdr const *DynSymHeader {nullptr};
Elf64_Shdr const *StrHeader = SectionHeaders.at(Header._64.e_shstrndx)._64;
@@ -897,7 +896,6 @@ void ELFContainer::PrintRelocationTable() const {
if (RelaHeader->sh_info != 0) {
LOGMAN_THROW_A_FMT(RelaHeader->sh_info < SectionHeaders.size(), "Rela header pointers to invalid GOT header");
GOTHeader = SectionHeaders.at(RelaHeader->sh_info)._64;
}
if (RelaHeader->sh_link != 0) {
@@ -966,7 +964,6 @@ void ELFContainer::FixupRelocations(void *ELFBase, uint64_t GuestELFBase, Symbol
}
else {
Elf64_Shdr const *RelaHeader{nullptr};
Elf64_Shdr const *GOTHeader {nullptr};
Elf64_Shdr const *DynSymHeader {nullptr};
Elf64_Shdr const *StringTableHeader{nullptr};
@@ -982,7 +979,6 @@ void ELFContainer::FixupRelocations(void *ELFBase, uint64_t GuestELFBase, Symbol
if (RelaHeader->sh_info != 0) {
LOGMAN_THROW_A_FMT(RelaHeader->sh_info < SectionHeaders.size(), "Rela header pointers to invalid GOT header");
GOTHeader = SectionHeaders.at(RelaHeader->sh_info)._64;
}
if (RelaHeader->sh_link != 0) {
+3 -2
View File
@@ -177,7 +177,9 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
return LoadBase;
}
std::string ResolveRootfsFile(std::string File, std::string RootFS) {
public:
static std::string ResolveRootfsFile(std::string const &File, std::string RootFS) {
// If the path is relative then just run that
if (std::filesystem::path(File).is_relative()) {
return File;
@@ -201,7 +203,6 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
return RootFSLink;
}
public:
struct LoadedSection {
uintptr_t ElfBase;
uintptr_t Base;
+11 -1
View File
@@ -203,6 +203,15 @@ void InterpreterHandler(std::string *Filename, std::string const &RootFS, std::v
}
}
void RootFSRedirect(std::string *Filename, std::string const &RootFS) {
auto RootFSLink = ELFCodeLoader2::ResolveRootfsFile(*Filename, RootFS);
std::error_code ec{};
if (std::filesystem::exists(RootFSLink, ec)) {
*Filename = RootFSLink;
}
}
bool RanAsInterpreter(const char *Program) {
return ExecutedWithFD || strstr(Program, "FEXInterpreter") != nullptr;
}
@@ -317,6 +326,8 @@ int main(int argc, char **argv, char **const envp) {
}
FEXCore::Telemetry::Initialize();
RootFSRedirect(&Program, LDPath());
InterpreterHandler(&Program, LDPath(), &Args);
std::error_code ec{};
@@ -327,7 +338,6 @@ int main(int argc, char **argv, char **const envp) {
return -ENOEXEC;
}
uint32_t KernelVersion = FEX::HLE::SyscallHandler::CalculateHostKernelVersion();
if (KernelVersion < FEX::HLE::SyscallHandler::KernelVersion(4, 17)) {
// We require 4.17 minimum for MAP_FIXED_NOREPLACE
+2 -2
View File
@@ -61,13 +61,13 @@ void MsgHandler(LogMan::DebugLevels Level, char const *Message)
break;
}
fmt::format("[{}] {}\n", CharLevel, Message);
fmt::print("[{}] {}\n", CharLevel, Message);
fflush(stdout);
}
void AssertHandler(char const *Message)
{
fmt::format("[ASSERT] {}\n", Message);
fmt::print("[ASSERT] {}\n", Message);
fflush(stdout);
}
+10 -9
View File
@@ -11,6 +11,7 @@ $end_info$
#include "Tests/LinuxSyscalls/x64/Syscalls.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
@@ -336,7 +337,7 @@ uint64_t FileManager::Open(const char *pathname, [[maybe_unused]] int flags, [[m
uint64_t FileManager::Close(int fd) {
{
std::lock_guard<std::mutex> lk(FDLock);
FHU::ScopedSignalMaskWithMutex lk(FDLock);
FDToNameMap.erase(fd);
}
return ::close(fd);
@@ -350,7 +351,7 @@ uint64_t FileManager::CloseRange(unsigned int first, unsigned int last, unsigned
if (!(flags & CLOSE_RANGE_CLOEXEC)) {
// If the flag was set then it doesn't actually close the FDs
// Just sets the flag on a range
std::lock_guard<std::mutex> lk(FDLock);
FHU::ScopedSignalMaskWithMutex lk(FDLock);
for (unsigned int i = first; i <= last; ++i) {
// We remove from first to last inclusive
FDToNameMap.erase(i);
@@ -421,7 +422,7 @@ uint64_t FileManager::FAccessat2(int dirfd, const char *pathname, int mode, int
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
auto Path = GetEmulatedPath(SelfPath, (flags & AT_SYMLINK_NOFOLLOW) == 0);
if (!Path.empty()) {
uint64_t Result = ::syscall(SYSCALL_DEF(faccessat2), dirfd, Path.c_str(), mode, flags);
if (Result != -1)
@@ -557,7 +558,7 @@ uint64_t FileManager::Openat([[maybe_unused]] int dirfs, const char *pathname, i
}
if (fd != -1) {
std::lock_guard lk(FDLock);
FHU::ScopedSignalMaskWithMutex lk(FDLock);
FDToNameMap.insert_or_assign(fd, SelfPath);
}
@@ -582,7 +583,7 @@ uint64_t FileManager::Openat2(int dirfs, const char *pathname, FEX::HLE::open_ho
}
if (fd != -1) {
std::lock_guard lk(FDLock);
FHU::ScopedSignalMaskWithMutex lk(FDLock);
FDToNameMap.insert_or_assign(fd, SelfPath);
}
@@ -594,7 +595,7 @@ uint64_t FileManager::Statx(int dirfd, const char *pathname, int flags, uint32_t
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
auto Path = GetEmulatedPath(SelfPath, (flags & AT_SYMLINK_NOFOLLOW) == 0);
if (!Path.empty()) {
uint64_t Result = FHU::Syscalls::statx(dirfd, Path.c_str(), flags, mask, statxbuf);
if (Result != -1)
@@ -630,7 +631,7 @@ uint64_t FileManager::NewFSStatAt(int dirfd, const char *pathname, struct stat *
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
auto Path = GetEmulatedPath(SelfPath, (flag & AT_SYMLINK_NOFOLLOW) == 0);
if (!Path.empty()) {
uint64_t Result = ::fstatat(dirfd, Path.c_str(), buf, flag);
if (Result != -1) {
@@ -644,7 +645,7 @@ uint64_t FileManager::NewFSStatAt64(int dirfd, const char *pathname, struct stat
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
auto Path = GetEmulatedPath(SelfPath, (flag & AT_SYMLINK_NOFOLLOW) == 0);
if (!Path.empty()) {
uint64_t Result = ::fstatat64(dirfd, Path.c_str(), buf, flag);
if (Result != -1) {
@@ -655,7 +656,7 @@ uint64_t FileManager::NewFSStatAt64(int dirfd, const char *pathname, struct stat
}
std::string *FileManager::FindFDName(int fd) {
std::lock_guard<std::mutex> lk(FDLock);
FHU::ScopedSignalMaskWithMutex lk(FDLock);
auto it = FDToNameMap.find(fd);
if (it == FDToNameMap.end()) {
return nullptr;
@@ -65,6 +65,8 @@ public:
std::string GetEmulatedPath(const char *pathname, bool FollowSymlink = false);
std::mutex *GetFDLock() { return &FDLock; }
private:
FEX::EmulatedFile::EmulatedFDManager EmuFD;
@@ -50,17 +50,5 @@ namespace FEX::HLE {
REGISTER_SYSCALL_IMPL(signalfd4, [](FEXCore::Core::CpuStateFrame *Frame, int fd, const uint64_t *mask, size_t sigsetsize, int flags) -> uint64_t {
return FEX::HLE::_SyscallHandler->GetSignalDelegator()->GuestSignalFD(fd, mask, sigsetsize, flags);
});
REGISTER_SYSCALL_IMPL_PASS(rt_sigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t pid, int sig, siginfo_t *info) -> uint64_t {
uint64_t Result = ::syscall(SYSCALL_DEF(rt_sigqueueinfo), pid, sig, info);
SYSCALL_ERRNO();
});
// XXX: siginfo_t definitely isn't correct for 32-bit
REGISTER_SYSCALL_IMPL_PASS(rt_tgsigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t tgid, pid_t tid, int sig, siginfo_t *info) -> uint64_t {
uint64_t Result = ::syscall(SYSCALL_DEF(rt_tgsigqueueinfo), tgid, tid, sig, info);
SYSCALL_ERRNO();
});
}
}
@@ -69,15 +69,5 @@ namespace FEX::HLE {
uint64_t Result = ::socketpair(domain, type, protocol, sv);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_PASS(setsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, const void *optval, socklen_t optlen) -> uint64_t {
uint64_t Result = ::setsockopt(sockfd, level, optname, optval, optlen);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_PASS(getsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, void *optval, socklen_t *optlen) -> uint64_t {
uint64_t Result = ::getsockopt(sockfd, level, optname, optval, optlen);
SYSCALL_ERRNO();
});
}
}
@@ -168,6 +168,11 @@ namespace FEX::HLE {
}
uint64_t ForkGuest(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::CpuStateFrame *Frame, uint32_t flags, void *stack, pid_t *parent_tid, pid_t *child_tid, void *tls) {
// Just before we fork, pull the Filemanagement mutex so it is in a safe state
// Locking the mutex will mean that both processes will end up with a locked mutex
auto Mutex = FEX::HLE::_SyscallHandler->FM.GetFDLock();
Mutex->lock();
pid_t Result{};
if (flags & CLONE_VFORK) {
// XXX: We don't currently support a vfork as it causes problems.
@@ -178,6 +183,9 @@ namespace FEX::HLE {
Result = fork();
}
// Unlock the mutex on both sides of the fork
Mutex->unlock();
if (Result == 0) {
// Child
// update the internal TID
@@ -212,5 +212,29 @@ namespace FEX::HLE::x32 {
else {
REGISTER_SYSCALL_IMPL_X32(pidfd_send_signal, UnimplementedSyscallSafe);
}
REGISTER_SYSCALL_IMPL_X32_PASS(rt_sigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t pid, int sig, compat_ptr<FEXCore::x86::siginfo_t> info) -> uint64_t {
siginfo_t info64{};
siginfo_t *info64_p{};
if (info) {
info64_p = &info64;
}
uint64_t Result = ::syscall(SYSCALL_DEF(rt_sigqueueinfo), pid, sig, info64_p);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X32_PASS(rt_tgsigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t tgid, pid_t tid, int sig, compat_ptr<FEXCore::x86::siginfo_t> info) -> uint64_t {
siginfo_t info64{};
siginfo_t *info64_p{};
if (info) {
info64_p = &info64;
}
uint64_t Result = ::syscall(SYSCALL_DEF(rt_tgsigqueueinfo), tgid, tid, sig, info64_p);
SYSCALL_ERRNO();
});
}
}
+343 -3
View File
@@ -7,6 +7,7 @@ $end_info$
#include "Tests/LinuxSyscalls/Syscalls.h"
#include "Tests/LinuxSyscalls/x32/Syscalls.h"
#include "Tests/LinuxSyscalls/x32/Types.h"
#include "Tests/LinuxSyscalls/x64/Syscalls.h"
#include <FEXCore/Utils/LogManager.h>
@@ -17,6 +18,7 @@ $end_info$
#include <memory>
#include <stddef.h>
#include <sys/socket.h>
#include <unistd.h>
#include <vector>
ARG_TO_STR(FEX::HLE::x32::compat_ptr<FEX::HLE::x32::mmsghdr_32>, "%lx")
@@ -26,6 +28,63 @@ namespace FEXCore::Core {
}
namespace FEX::HLE::x32 {
// Some sockopt defines for older build environments
#ifndef SO_RCVTIMEO_OLD
#define SO_RCVTIMEO_OLD 20
#endif
#ifndef SO_SNDTIMEO_OLD
#define SO_SNDTIMEO_OLD 21
#endif
#ifndef SO_TIMESTAMP_OLD
#define SO_TIMESTAMP_OLD 29
#endif
#ifndef SO_TIMESTAMPNS_OLD
#define SO_TIMESTAMPNS_OLD 35
#endif
#ifndef SO_TIMESTAMPING_OLD
#define SO_TIMESTAMPING_OLD 37
#endif
#ifndef SO_TXTIME
#define SO_TXTIME 61
#endif
#ifndef SO_BINDTOIFINDEX
#define SO_BINDTOIFINDEX 62
#endif
#ifndef SO_TIMESTAMP_NEW
#define SO_TIMESTAMP_NEW 63
#endif
#ifndef SO_TIMESTAMPNS_NEW
#define SO_TIMESTAMPNS_NEW 64
#endif
#ifndef SO_TIMESTAMPING_NEW
#define SO_TIMESTAMPING_NEW 65
#endif
#ifndef SO_RCVTIMEO_NEW
#define SO_RCVTIMEO_NEW 66
#endif
#ifndef SO_SNDTIMEO_NEW
#define SO_SNDTIMEO_NEW 67
#endif
#ifndef SO_DETACH_REUSEPORT_BPF
#define SO_DETACH_REUSEPORT_BPF 68
#endif
#ifndef SO_PREFER_BUSY_POLL
#define SO_PREFER_BUSY_POLL 69
#endif
#ifndef SO_BUSY_POLL_BUDGET
#define SO_BUSY_POLL_BUDGET 70
#endif
#ifndef SO_NETNS_COOKIE
#define SO_NETNS_COOKIE 71
#endif
#ifndef SO_BUF_LOCK
#define SO_BUF_LOCK 72
#endif
#ifndef SO_RESERVE_MEM
#define SO_RESERVE_MEM 73
#endif
enum SockOp {
OP_SOCKET = 1,
OP_BIND = 2,
@@ -243,6 +302,279 @@ namespace FEX::HLE::x32 {
SYSCALL_ERRNO();
}
static uint64_t SetSockOpt(int sockfd, int level, int optname, compat_ptr<void> optval, int optlen) {
uint64_t Result{};
if (level == SOL_SOCKET) {
switch (optname) {
case SO_ATTACH_FILTER:
case SO_ATTACH_REUSEPORT_CBPF: {
struct sock_fprog32 {
uint16_t len;
uint32_t filter;
};
struct sock_fprog64 {
uint16_t len;
uint64_t filter;
};
if (optlen != sizeof(sock_fprog32)) {
return -EINVAL;
}
sock_fprog32 *prog = reinterpret_cast<sock_fprog32*>(optval.Ptr);
sock_fprog64 prog64{};
prog64.len = prog->len;
prog64.filter = prog->filter;
Result = ::syscall(SYSCALL_DEF(setsockopt),
sockfd,
level,
optname,
&prog64,
sizeof(sock_fprog64)
);
break;
}
case SO_RCVTIMEO_OLD: {
// _OLD uses old_timeval32. Needs to be converted
struct timeval tv64 = *reinterpret_cast<timeval32*>(optval.Ptr);
Result = ::syscall(SYSCALL_DEF(setsockopt),
sockfd,
level,
SO_RCVTIMEO_NEW,
&tv64,
sizeof(tv64)
);
break;
}
case SO_SNDTIMEO_OLD: {
// _OLD uses old_timeval32. Needs to be converted
struct timeval tv64 = *reinterpret_cast<timeval32*>(optval.Ptr);
Result = ::syscall(SYSCALL_DEF(setsockopt),
sockfd,
level,
SO_SNDTIMEO_NEW,
&tv64,
sizeof(tv64)
);
break;
}
// Each optname as a reminder which setting has been manually checked
case SO_DEBUG:
case SO_REUSEADDR:
case SO_TYPE:
case SO_ERROR:
case SO_DONTROUTE:
case SO_BROADCAST:
case SO_SNDBUF:
case SO_RCVBUF:
case SO_SNDBUFFORCE:
case SO_RCVBUFFORCE:
case SO_KEEPALIVE:
case SO_OOBINLINE:
case SO_NO_CHECK:
case SO_PRIORITY:
case SO_LINGER:
case SO_BSDCOMPAT:
case SO_REUSEPORT:
/**
* @name These end up differing between {x86,arm} and {powerpc, alpha, sparc, mips, parisc}
* @{ */
case SO_PASSCRED:
case SO_PEERCRED:
case SO_RCVLOWAT:
case SO_SNDLOWAT:
/** @} */
case SO_SECURITY_AUTHENTICATION:
case SO_SECURITY_ENCRYPTION_TRANSPORT:
case SO_SECURITY_ENCRYPTION_NETWORK:
case SO_DETACH_FILTER:
case SO_PEERNAME:
case SO_TIMESTAMP_OLD: // Returns int32_t boolean
case SO_ACCEPTCONN:
case SO_PEERSEC:
// Gap 32, 33
case SO_PASSSEC:
case SO_TIMESTAMPNS_OLD: // Returns int32_t boolean
case SO_MARK:
case SO_TIMESTAMPING_OLD: // Returns so_timestamping
case SO_PROTOCOL:
case SO_DOMAIN:
case SO_RXQ_OVFL:
case SO_WIFI_STATUS:
case SO_PEEK_OFF:
case SO_NOFCS:
case SO_LOCK_FILTER:
case SO_SELECT_ERR_QUEUE:
case SO_BUSY_POLL:
case SO_MAX_PACING_RATE:
case SO_BPF_EXTENSIONS:
case SO_INCOMING_CPU:
case SO_ATTACH_BPF:
case SO_ATTACH_REUSEPORT_EBPF:
case SO_CNX_ADVICE:
// Gap 54 (SCM_TIMESTAMPING_OPT_STATS)
case SO_MEMINFO:
case SO_INCOMING_NAPI_ID:
case SO_COOKIE: // Cookie always returns 64-bit even on 32-bit
// Gap 58 (SCM_TIMESTAMPING_PKTINFO)
case SO_PEERGROUPS:
case SO_ZEROCOPY:
case SO_TXTIME:
case SO_BINDTOIFINDEX:
case SO_TIMESTAMP_NEW:
case SO_TIMESTAMPNS_NEW:
case SO_TIMESTAMPING_NEW:
case SO_RCVTIMEO_NEW:
case SO_SNDTIMEO_NEW:
case SO_DETACH_REUSEPORT_BPF:
case SO_PREFER_BUSY_POLL:
case SO_BUSY_POLL_BUDGET:
case SO_NETNS_COOKIE: // Cookie always returns 64-bit even on 32-bit
case SO_BUF_LOCK:
case SO_RESERVE_MEM:
default:
Result = ::syscall(SYSCALL_DEF(setsockopt),
sockfd,
level,
optname,
reinterpret_cast<const void*>(optval.Ptr),
optlen
);
break;
}
}
else {
Result = ::syscall(SYSCALL_DEF(setsockopt),
sockfd,
level,
optname,
reinterpret_cast<const void*>(optval.Ptr),
optlen
);
}
SYSCALL_ERRNO();
}
static uint64_t GetSockOpt(int sockfd, int level, int optname, compat_ptr<void> optval, compat_ptr<socklen_t> optlen) {
uint64_t Result{};
if (level == SOL_SOCKET) {
switch (optname) {
case SO_RCVTIMEO_OLD: {
// _OLD uses old_timeval32. Needs to be converted
struct timeval tv64{};
Result = ::syscall(SYSCALL_DEF(getsockopt),
sockfd,
level,
SO_RCVTIMEO_NEW,
&tv64,
sizeof(tv64)
);
*reinterpret_cast<timeval32*>(optval.Ptr) = tv64;
break;
}
case SO_SNDTIMEO_OLD: {
// _OLD uses old_timeval32. Needs to be converted
struct timeval tv64{};
Result = ::syscall(SYSCALL_DEF(getsockopt),
sockfd,
level,
SO_SNDTIMEO_NEW,
&tv64,
sizeof(tv64)
);
*reinterpret_cast<timeval32*>(optval.Ptr) = tv64;
break;
}
// Each optname as a reminder which setting has been manually checked
case SO_DEBUG:
case SO_REUSEADDR:
case SO_TYPE:
case SO_ERROR:
case SO_DONTROUTE:
case SO_BROADCAST:
case SO_SNDBUF:
case SO_RCVBUF:
case SO_SNDBUFFORCE:
case SO_RCVBUFFORCE:
case SO_KEEPALIVE:
case SO_OOBINLINE:
case SO_NO_CHECK:
case SO_PRIORITY:
case SO_LINGER:
case SO_BSDCOMPAT:
case SO_REUSEPORT:
/**
* @name These end up differing between {x86,arm} and {powerpc, alpha, sparc, mips, parisc}
* @{ */
case SO_PASSCRED:
case SO_PEERCRED:
case SO_RCVLOWAT:
case SO_SNDLOWAT:
/** @} */
case SO_SECURITY_AUTHENTICATION:
case SO_SECURITY_ENCRYPTION_TRANSPORT:
case SO_SECURITY_ENCRYPTION_NETWORK:
case SO_ATTACH_FILTER: // Renamed to SO_GET_FILTER on get. Same between 32-bit and 64-bit
case SO_DETACH_FILTER:
case SO_PEERNAME:
case SO_TIMESTAMP_OLD: // Returns int32_t boolean
case SO_ACCEPTCONN:
case SO_PEERSEC:
// Gap 32, 33
case SO_PASSSEC:
case SO_TIMESTAMPNS_OLD: // Returns int32_t boolean
case SO_MARK:
case SO_TIMESTAMPING_OLD: // Returns so_timestamping
case SO_PROTOCOL:
case SO_DOMAIN:
case SO_RXQ_OVFL:
case SO_WIFI_STATUS:
case SO_PEEK_OFF:
case SO_NOFCS:
case SO_LOCK_FILTER:
case SO_SELECT_ERR_QUEUE:
case SO_BUSY_POLL:
case SO_MAX_PACING_RATE:
case SO_BPF_EXTENSIONS:
case SO_INCOMING_CPU:
case SO_ATTACH_BPF:
case SO_ATTACH_REUSEPORT_CBPF: // Doesn't do anything in get
case SO_ATTACH_REUSEPORT_EBPF:
case SO_CNX_ADVICE:
// Gap 54 (SCM_TIMESTAMPING_OPT_STATS)
case SO_MEMINFO:
case SO_INCOMING_NAPI_ID:
case SO_COOKIE: // Cookie always returns 64-bit even on 32-bit
// Gap 58 (SCM_TIMESTAMPING_PKTINFO)
case SO_PEERGROUPS:
case SO_ZEROCOPY:
case SO_TXTIME:
case SO_BINDTOIFINDEX:
case SO_TIMESTAMP_NEW:
case SO_TIMESTAMPNS_NEW:
case SO_TIMESTAMPING_NEW:
case SO_RCVTIMEO_NEW:
case SO_SNDTIMEO_NEW:
case SO_DETACH_REUSEPORT_BPF:
case SO_PREFER_BUSY_POLL:
case SO_BUSY_POLL_BUDGET:
case SO_NETNS_COOKIE: // Cookie always returns 64-bit even on 32-bit
case SO_BUF_LOCK:
case SO_RESERVE_MEM:
default:
Result = ::syscall(SYSCALL_DEF(getsockopt), sockfd, level, optname, optval, optlen);
break;
}
}
else {
Result = ::syscall(SYSCALL_DEF(getsockopt), sockfd, level, optname, optval, optlen);
}
SYSCALL_ERRNO();
}
void RegisterSocket() {
REGISTER_SYSCALL_IMPL_X32(socketcall, [](FEXCore::Core::CpuStateFrame *Frame, uint32_t call, uint32_t *Arguments) -> uint64_t {
uint64_t Result{};
@@ -313,17 +645,17 @@ namespace FEX::HLE::x32 {
break;
}
case OP_SETSOCKOPT: {
Result = ::setsockopt(
return SetSockOpt(
Arguments[0],
Arguments[1],
Arguments[2],
reinterpret_cast<const void*>(Arguments[3]),
Arguments[3],
reinterpret_cast<socklen_t>(Arguments[4])
);
break;
}
case OP_GETSOCKOPT: {
Result = ::getsockopt(
return GetSockOpt(
Arguments[0],
Arguments[1],
Arguments[2],
@@ -463,5 +795,13 @@ namespace FEX::HLE::x32 {
REGISTER_SYSCALL_IMPL_X32(recvmsg, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, struct msghdr32 *msg, int flags) -> uint64_t {
return RecvMsg(sockfd, msg, flags);
});
REGISTER_SYSCALL_IMPL_X32(setsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, compat_ptr<void> optval, socklen_t optlen) -> uint64_t {
return SetSockOpt(sockfd, level, optname, optval, optlen);
});
REGISTER_SYSCALL_IMPL_X32(getsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, compat_ptr<void> optval, compat_ptr<socklen_t> optlen) -> uint64_t {
return GetSockOpt(sockfd, level, optname, optval, optlen);
});
}
}
+19 -4
View File
@@ -13,6 +13,7 @@ $end_info$
#include "Tests/LinuxSyscalls/x64/Syscalls.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
@@ -31,6 +32,7 @@ $end_info$
#include <vector>
ARG_TO_STR(FEX::HLE::x32::compat_ptr<FEX::HLE::x32::stack_t32>, "%x")
ARG_TO_STR(FEX::HLE::x32::compat_ptr<FEXCore::x86::siginfo_t>, "%x")
namespace FEX::HLE::x32 {
// The kernel only gives 32-bit userspace 3 TLS segments
@@ -353,19 +355,32 @@ namespace FEX::HLE::x32 {
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X32(waitid, [](FEXCore::Core::CpuStateFrame *Frame, int which, pid_t upid, siginfo_t *infop, int options, struct rusage_32 *rusage) -> uint64_t {
REGISTER_SYSCALL_IMPL_X32(waitid, [](FEXCore::Core::CpuStateFrame *Frame, int which, pid_t upid, compat_ptr<FEXCore::x86::siginfo_t> info, int options, struct rusage_32 *rusage) -> uint64_t {
struct rusage usage64{};
struct rusage *usage64_p{};
siginfo_t info64{};
siginfo_t *info64_p{};
if (rusage) {
usage64 = *rusage;
usage64_p = &usage64;
}
uint64_t Result = ::syscall(SYSCALL_DEF(waitid), which, upid, infop, options, usage64_p);
if (info) {
info64_p = &info64;
}
if (rusage) {
*rusage = usage64;
uint64_t Result = ::syscall(SYSCALL_DEF(waitid), which, upid, info64_p, options, usage64_p);
if (Result != -1) {
if (rusage) {
*rusage = usage64;
}
if (info) {
*info = info64;
}
}
SYSCALL_ERRNO();
@@ -40,6 +40,16 @@ namespace FEX::HLE::x64 {
else {
REGISTER_SYSCALL_IMPL_X64(pidfd_send_signal, UnimplementedSyscallSafe);
}
REGISTER_SYSCALL_IMPL_X64_PASS(rt_sigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t pid, int sig, siginfo_t *info) -> uint64_t {
uint64_t Result = ::syscall(SYSCALL_DEF(rt_sigqueueinfo), pid, sig, info);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X64_PASS(rt_tgsigqueueinfo, [](FEXCore::Core::CpuStateFrame *Frame, pid_t tgid, pid_t tid, int sig, siginfo_t *info) -> uint64_t {
uint64_t Result = ::syscall(SYSCALL_DEF(rt_tgsigqueueinfo), tgid, tid, sig, info);
SYSCALL_ERRNO();
});
}
}
+10
View File
@@ -40,5 +40,15 @@ namespace FEX::HLE::x64 {
uint64_t Result = ::recvmsg(sockfd, msg, flags);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X64_PASS(setsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, const void *optval, socklen_t optlen) -> uint64_t {
uint64_t Result = ::setsockopt(sockfd, level, optname, optval, optlen);
SYSCALL_ERRNO();
});
REGISTER_SYSCALL_IMPL_X64_PASS(getsockopt, [](FEXCore::Core::CpuStateFrame *Frame, int sockfd, int level, int optname, void *optval, socklen_t *optlen) -> uint64_t {
uint64_t Result = ::getsockopt(sockfd, level, optname, optval, optlen);
SYSCALL_ERRNO();
});
}
}
+4 -2
View File
@@ -201,8 +201,10 @@ namespace {
char EmulatedCPUCores[32]{};
if (ImGui::BeginTabItem("CPU")) {
std::optional<std::string*> Value{};
#ifdef INTERPRETER_ENABLED
ImGui::Text("Core:");
auto Value = LoadedConfig->Get(FEXCore::Config::ConfigOption::CONFIG_CORE);
Value = LoadedConfig->Get(FEXCore::Config::ConfigOption::CONFIG_CORE);
ImGui::SameLine();
if (ImGui::RadioButton("Int", Value.has_value() && **Value == "0")) {
@@ -214,7 +216,7 @@ namespace {
LoadedConfig->EraseSet(FEXCore::Config::ConfigOption::CONFIG_CORE, "1");
ConfigChanged = true;
}
#endif
Value = LoadedConfig->Get(FEXCore::Config::ConfigOption::CONFIG_MAXINST);
if (Value.has_value() && !(*Value)->empty()) {
strncpy(BlockSize, &(*Value)->at(0), 32);
+222 -22
View File
@@ -7,6 +7,7 @@
#include <fstream>
#include <iostream>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/wait.h>
namespace Exec {
@@ -20,7 +21,48 @@ namespace Exec {
int32_t Status{};
waitpid(pid, &Status, 0);
if (WIFEXITED(Status)) {
return WEXITSTATUS(Status);
return (int8_t)WEXITSTATUS(Status);
}
}
return -1;
}
int32_t ExecAndWaitForResponseRedirect(const char *path, char* const* args, int stdoutRedirect = -2, int stderrRedirect = -2) {
pid_t pid = fork();
if (pid == 0) {
if (stdoutRedirect == -1) {
close(STDOUT_FILENO);
}
else if (stdoutRedirect == -2) {
// Do nothing
}
else {
if (stdoutRedirect != STDOUT_FILENO) {
close(STDOUT_FILENO);
}
dup2(stdoutRedirect, STDOUT_FILENO);
}
if (stderrRedirect == -1) {
close(STDERR_FILENO);
}
else if (stderrRedirect == -2) {
// Do nothing
}
else {
if (stderrRedirect != STDOUT_FILENO) {
close(STDERR_FILENO);
}
dup2(stderrRedirect, STDERR_FILENO);
}
execvp(path, args);
_exit(-1);
}
else {
int32_t Status{};
waitpid(pid, &Status, 0);
if (WIFEXITED(Status)) {
return (int8_t)WEXITSTATUS(Status);
}
}
@@ -70,6 +112,88 @@ namespace Exec {
}
}
namespace WorkingAppsTester {
static bool Has_Curl {false};
static bool Has_Squashfuse {false};
static bool Has_Unsquashfs {false};
void CheckCurl() {
// Check if curl exists on the host
std::vector<const char*> ExecveArgs = {
"curl",
"-V",
nullptr,
};
int32_t Result = Exec::ExecAndWaitForResponseRedirect(ExecveArgs[0], const_cast<char* const*>(ExecveArgs.data()), -1, -1);
Has_Curl = Result != -1;
}
void CheckSquashfuse() {
std::vector<const char*> ExecveArgs = {
"squashfuse",
"--help",
nullptr,
};
int32_t Result = Exec::ExecAndWaitForResponseRedirect(ExecveArgs[0], const_cast<char* const*>(ExecveArgs.data()), -1, -1);
Has_Squashfuse = Result != -1;
}
void CheckUnsquashfs() {
std::vector<const char*> ExecveArgs = {
"unsquashfs",
"--help",
nullptr,
};
int fd = memfd_create("stdout", 0);
int32_t Result = Exec::ExecAndWaitForResponseRedirect(ExecveArgs[0], const_cast<char* const*>(ExecveArgs.data()), fd, fd);
Has_Unsquashfs = Result != -1;
if (Has_Unsquashfs) {
// Seek back to the start
lseek(fd, 0, SEEK_SET);
// Unsquashfs needs to support zstd
// Scan its output to find the zstd compressor
FILE *fp = fdopen(fd, "r");
char *Line {nullptr};
ssize_t NumRead;
size_t Len;
bool ReadingDecompressors = false;
bool SupportsZSTD = false;
while ((NumRead = getline(&Line, &Len, fp)) != -1) {
if (!ReadingDecompressors) {
if (strstr(Line, "Decompressors available")) {
ReadingDecompressors = true;
}
}
else {
if (strstr(Line, "zstd")) {
SupportsZSTD = true;
}
}
}
free(Line);
fclose(fp);
// Disable unsquashfs if it doesn't support ZSTD
if (!SupportsZSTD) {
Has_Unsquashfs = false;
}
}
close(fd);
}
void Init() {
CheckCurl();
CheckSquashfuse();
CheckUnsquashfs();
}
}
namespace DistroQuery {
struct DistroInfo {
std::string DistroName;
@@ -239,7 +363,7 @@ namespace WebFileFetcher {
auto PathName = Path + filename;
std::string BigArgs =
fmt::format("curl {} -o {}", URL, PathName);
fmt::format("curl -C - {} -o {}", URL, PathName);
std::vector<const char*> ExecveArgs = {
"/bin/sh",
"-c",
@@ -257,10 +381,12 @@ namespace WebFileFetcher {
// -# for progress bar
// -o for output file
// -f for silent fail
std::string CurlPipe = fmt::format("curl -#f {} -o {} 2>&1", URL, PathName);
std::string CurlPipe = fmt::format("curl -C - -#f {} -o {} 2>&1", URL, PathName);
const std::string StdBuf = "stdbuf -oL tr '\\r' '\\n'";
const std::string SedBuf = "sed -u 's/[^0-9]*\\([0-9]*\\).*/\\1/'";
const std::string ZenityBuf = "zenity --time-remaining --progress --auto-close --no-cancel --title 'Downloading'";
// zenity --auto-close can't be used since `curl -C` for whatever reason prints 100% at the start.
// Making zenity vanish immediately
const std::string ZenityBuf = "zenity --time-remaining --progress --no-cancel --title 'Downloading'";
std::string BigArgs =
fmt::format("{} | {} | {} | {}", CurlPipe, StdBuf, SedBuf, ZenityBuf);
std::vector<const char*> ExecveArgs = {
@@ -482,7 +608,6 @@ namespace Zenity {
}
if (!WebFileFetcher::DownloadToPathWithZenityProgress(Target.URL, RootFS)) {
ExecWithInfo("Couldn't download RootFS");
return false;
}
@@ -642,12 +767,25 @@ namespace TTY {
return false;
}
}
auto DoDownload = [&Target, &RootFS]() -> bool {
if (!WebFileFetcher::DownloadToPath(Target.URL, RootFS)) {
fmt::print("Couldn't download RootFS\n");
return false;
}
if (!WebFileFetcher::DownloadToPath(Target.URL, RootFS)) {
fmt::print("Couldn't download RootFS\n");
return false;
return true;
};
while (DoDownload() == false) {
if (AskForConfirmation("Curl RootFS download failed. Do you want to retry?")) {
// Loop to retry
}
else {
return false;
}
}
// Got here then we passed
return true;
}
return false;
@@ -778,6 +916,13 @@ int main(int argc, char **argv, char **const envp) {
return 0;
}
WorkingAppsTester::Init();
// Check if curl exists on the host
if (!WorkingAppsTester::Has_Curl) {
ExecWithInfo("curl is required to use this tool. Please install curl before using.");
return -1;
}
FEX_CONFIG_OPT(LDPath, ROOTFS);
std::error_code ec;
@@ -802,22 +947,52 @@ int main(int argc, char **argv, char **const envp) {
if (!ValidateCheckExists(Target)) {
// Keep going
}
else if (ValidateDownloadSelection(Target)) {
uint64_t ExpectedHash = std::stoul(Target.Hash, nullptr, 16);
else {
auto ValidateDownload = [&Target, &PathName]() -> std::pair<int32_t, bool> {
std::error_code ec;
if (ValidateDownloadSelection(Target)) {
uint64_t ExpectedHash = std::stoul(Target.Hash, nullptr, 16);
if (std::filesystem::exists(PathName, ec)) {
auto Res = XXFileHash::HashFile(PathName);
if (Res.first == false ||
Res.second != ExpectedHash) {
std::string Text = fmt::format("Couldn't hash the rootfs or hash didn't match\n");
Text += fmt::format("Hash {:x} != Expected Hash {:x}\n", Res.second, ExpectedHash);
ExecWithInfo(Text);
return -1;
if (std::filesystem::exists(PathName, ec)) {
auto Res = XXFileHash::HashFile(PathName);
if (Res.first == false ||
Res.second != ExpectedHash) {
std::string Text = fmt::format("Couldn't hash the rootfs or hash didn't match\n");
Text += fmt::format("Hash {:x} != Expected Hash {:x}\n", Res.second, ExpectedHash);
ExecWithInfo(Text);
return std::make_pair(-1, true);
}
}
else {
ExecWithInfo("Correctly downloaded RootFS but doesn't exist?");
return std::make_pair(-1, false);
}
}
else {
ExecWithInfo("Couldn't download rootfs for some reason.");
return std::make_pair(-1, false);
}
return std::make_pair(0, false);
};
std::pair<int32_t, bool> Result{};
while ((Result = ValidateDownload()).second == true &&
Result.first == -1) {
if (AskForConfirmation("Do you want to try downloading the RootFS again?")) {
// Continue the loop
}
else {
// Didn't want to retry, just exit now
return Result.first;
}
}
else {
ExecWithInfo("Correctly downloaded RootFS but doesn't exist?");
return -1;
// Early exit on other errors
if (Result.first == -1 &&
Result.second == false) {
return Result.first;
}
}
@@ -825,7 +1000,32 @@ int main(int argc, char **argv, char **const envp) {
"Extract",
"As-Is",
};
auto Result = AskForConfirmationList("Do you wish to extract the squashfs file or use it as-is?", Args);
int32_t Result{};
if (WorkingAppsTester::Has_Unsquashfs) {
if (WorkingAppsTester::Has_Squashfuse) {
Result = AskForConfirmationList("Do you wish to extract the squashfs file or use it as-is?", Args);
}
else {
Args.pop_back();
Result = AskForConfirmationList("Squashfuse doesn't work. Do you wish to extract the squashfs file?", Args);
}
}
else {
if (WorkingAppsTester::Has_Squashfuse) {
Args.erase(Args.begin());
Result = AskForConfirmationList("Unsquashfs doesn't work. Do you want to use the squashfs file as-is?", Args);
if (Result == 0) {
// We removed an argument, Just change "As-Is" from 0 to 1 for later logic to work
Result = 1;
}
}
else {
Args.erase(Args.begin());
ExecWithInfo("Unsquashfs and squashfuse isn't working. Leaving rootfs as-is");
Result = -1;
}
}
if (Result == -1 ||
Result == 1) {
// Nothing
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: thunklibs|Vulkan
$end_info$
*/
#define VK_USE_PLATFORM_XLIB_XRANDR_EXT
#define VK_USE_PLATFORM_XLIB_KHR
#define VK_USE_PLATFORM_XCB_KHR
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: thunklibs|Vulkan
$end_info$
*/
#define VK_USE_PLATFORM_XLIB_XRANDR_EXT
#define VK_USE_PLATFORM_XLIB_KHR
#define VK_USE_PLATFORM_XCB_KHR
+5 -1
View File
@@ -1,4 +1,4 @@
# FEX-2201
# FEX-2202
## External/FEXCore
See [FEXCore/Readme.md](../External/FEXCore/Readme.md) for more details
@@ -185,6 +185,10 @@ These are generated + glue logic 1:1 thunks unless noted otherwise
- [libSDL2_Guest.cpp](../ThunkLibs/libSDL2/libSDL2_Guest.cpp): Handles sdlglproc, dload, stubs a few log fns
- [libSDL2_Host.cpp](../ThunkLibs/libSDL2/libSDL2_Host.cpp)
#### Vulkan
- [Guest.cpp](../ThunkLibs/libvulkan_device/Guest.cpp)
- [Host.cpp](../ThunkLibs/libvulkan_device/Host.cpp)
#### X11
- [libX11_Guest.cpp](../ThunkLibs/libX11/libX11_Guest.cpp): Handles callbacks and varargs
- [libX11_Host.cpp](../ThunkLibs/libX11/libX11_Host.cpp): Handles callbacks and varargs
+3
View File
@@ -5,3 +5,6 @@ Test_32Bit_X87/D9_FD.asm
Test_32Bit_X87/D9_F9.asm
Test_32Bit_X87/D9_F2.asm
# Relies on rounding correctness
Test_32Bit_X87/D9_F8.asm
+1
View File
@@ -0,0 +1 @@
Test_32Bit_X87/D9_F8.asm
@@ -0,0 +1,44 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for underflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a mov eax + hlt over there first
; 0xb8'44'43'42'41: mov eax, 0x41424344
; 0xf4: hlt
mov ebx, 0xe0000000
mov al, 0xb8
mov byte [ebx], al
mov eax, 0x41424344
mov dword [ebx + 1], eax
mov al, 0xf4
mov byte [ebx + 5], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; Setup esp
mov esp, 0xe0001000
; This is dependent on where it is in the code!
call -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
@@ -0,0 +1,50 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for overflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a call 0x11000 over there
; 0xe8'fb'0f'01'20 : call 0x11000
; 0xf4: hlt - Just in case
mov ebx, 0xe0000000
mov al, 0xe8
mov byte [ebx], al
mov eax, 0x20010ffb
mov dword [ebx + 1], eax
mov al, 0xf4
mov byte [ebx + 5], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; Setup esp
mov esp, 0xe0001000
; This is dependent on where it is in the code!
call -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
; This is where the JIT code will land
align 0x1000
mov eax, 0x41424344
hlt
@@ -0,0 +1,41 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for underflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a mov eax + hlt over there first
; 0xb8'44'43'42'41: mov eax, 0x41424344
; 0xf4: hlt
mov ebx, 0xe0000000
mov al, 0xb8
mov byte [ebx], al
mov eax, 0x41424344
mov dword [ebx + 1], eax
mov al, 0xf4
mov byte [ebx + 5], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; This is dependent on where it is in the code!
jmp -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
@@ -0,0 +1,48 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for overflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a jmp 0x11000 over there
; 0xe9'fb'0f'01'20 : jmp 0x11000
; 0xf4: hlt - Just in case
mov ebx, 0xe0000000
mov al, 0xe9
mov byte [ebx], al
mov eax, 0x20010ffb
mov dword [ebx + 1], eax
mov al, 0xf4
mov byte [ebx + 5], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; This is dependent on where it is in the code!
jmp -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
; This is where the JIT code will land
align 0x1000
mov eax, 0x41424344
hlt
+44
View File
@@ -0,0 +1,44 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for underflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a mov eax + hlt over there first
; 0xb8'44'43'42'41: mov eax, 0x41424344
; 0xf4: hlt
mov ebx, 0xe0000000
mov al, 0xb8
mov byte [ebx], al
mov eax, 0x41424344
mov dword [ebx + 1], eax
mov al, 0xf4
mov byte [ebx + 5], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; Clear the lower flags so the branch gets taken
sahf
; This is dependent on where it is in the code!
jnb -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
+50
View File
@@ -0,0 +1,50 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x41424344"
},
"Mode": "32BIT"
}
%endif
; Tests for 32-bit signed displacement wrapping
; Testing for overflow specifically
; Will crash or hit the code we emit to memory
; We map ten pages to 0xe000'0000
; Generate a call 0x11000 over there
; 0x0f'83'fa'0f'01'20 : jnb 0x11000
; 0xf4: hlt - Just in case
mov ebx, 0xe0000000
mov ax, 0x830f
mov word [ebx], ax
mov eax, 0x20010ffa
mov dword [ebx + 2], eax
mov al, 0xf4
mov byte [ebx + 6], al
; Do a jump dance to stop multiblock from trying to optimize
; Otherwise it will JIT code from 0xe000'0000 before written
lea ebx, [rel next]
jmp ebx
next:
; Move temp to eax to overwrite
mov eax, 0
; Clear the lower flags so the branch gets taken
sahf
; This is dependent on where it is in the code!
jnb -0x20000000
; Definitely wrong if we hit here
mov eax, -1
hlt
; This is where the JIT code will land
align 0x1000
mov eax, 0x41424344
hlt
@@ -2,6 +2,8 @@
#include <chrono>
#include <csetjmp>
#include <FEXCore/Utils/InterruptableConditionVariable.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <thread>
#include <signal.h>
@@ -62,7 +64,7 @@ void WaitThreadLongJump(
int32_t *TID) {
// Store the TID
*TID = ::gettid();
*TID = FHU::Syscalls::gettid();
// Setup a long jump signal handler
struct sigaction sa{};
@@ -106,7 +108,7 @@ TEST_CASE("SignaledWaitLongJumpNoSignal") {
// Wait for the thread to signal that it is ready to receive signal
WaitMutex.WaitFor(std::chrono::milliseconds(500));
// Send the signal to the thread
tgkill(::getpid(), TID, SIGUSR1);
FHU::Syscalls::tgkill(::getpid(), TID, SIGUSR1);
}
// Wait for thread join
@@ -138,7 +140,7 @@ TEST_CASE("SignaledWaitLongJumpSignal") {
WaitMutex.WaitFor(std::chrono::milliseconds(500));
// Send the signal to the thread
tgkill(::getpid(), TID, SIGUSR1);
FHU::Syscalls::tgkill(::getpid(), TID, SIGUSR1);
}
// Notify the thread's mutex now
+3
View File
@@ -23,3 +23,6 @@ Test_Primary/Primary_E9.asm
Test_X87/D9_F9.asm
Test_X87/D9_F2.asm
# Relies on rounding correctness
Test_X87/D9_F8.asm
+32
View File
@@ -0,0 +1,32 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0xFFFFFFFFFFFF43FF",
"RBX": "0"
}
}
%endif
; Set EFLAGS to known value with sahf
mov rax, -1
sahf
movups xmm0, [rel .data]
; Tests a bug that FEX had where ptest would not set OF, SF, AF, PF to zero
ptest xmm0, xmm0
; Now load back
; ZF = 1
; CF = 1
; OF, SF, AF, PF should be zero
lahf
; lahf doesn't get OF, get it with seto
mov rbx, 0
seto bl
hlt
.data:
dq 0
dq 0
+63
View File
@@ -0,0 +1,63 @@
%ifdef CONFIG
{
"RegData": {
"XMM0": ["0x0000000000000000", "0xFFFFFFFFFFFFFFFF"],
"XMM1": ["0x0000000000000000", "0x0000000000000000"],
"XMM2": ["0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0x0000000000000000", "0xFFFFFFFFFFFFFFFF"],
"XMM4": ["0xFFFFFFFFFFFFFFFF", "0x0000000000000000"],
"XMM5": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM6": ["0x0000000000000000", "0x0000000000000000"],
"XMM7": ["0xFFFFFFFFFFFFFFFF", "0x0000000000000000"],
"XMM8": ["0x0000000000000000", "0xFFFFFFFFFFFFFFFF"],
"XMM9": ["0xFFFFFFFFFFFFFFFF", "0x0000000000000000"]
}
}
%endif
movaps xmm0, [rel .data0]
movaps xmm1, [rel .data1]
movaps xmm2, [rel .data2]
movaps xmm3, [rel .data3]
movaps xmm4, [rel .data4]
movaps xmm5, [rel .data0]
movaps xmm6, [rel .data1]
movaps xmm7, [rel .data2]
movaps xmm8, [rel .data3]
movaps xmm9, [rel .data4]
pcmpgtq xmm0, [rel .data4]
pcmpgtq xmm1, [rel .data3]
pcmpgtq xmm2, [rel .data2]
pcmpgtq xmm3, [rel .data1]
pcmpgtq xmm4, [rel .data0]
pcmpgtq xmm5, [rel .data1]
pcmpgtq xmm6, [rel .data2]
pcmpgtq xmm7, [rel .data3]
pcmpgtq xmm8, [rel .data4]
pcmpgtq xmm9, [rel .data0]
hlt
align 16
.data0:
dq 0
dq 0
.data1:
dq -1
dq -1
.data2:
dq 1
dq 1
.data3:
dq -1
dq 1
.data4:
dq 1
dq -1
+63
View File
@@ -0,0 +1,63 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000E330A81A",
"RBX": "0x00000000BE2DA0A5",
"RCX": "0x00000000ADBE9F64"
}
}
%endif
; This is a clone of the F2_F0 crc32 test with manually coded crc32 with prefix 66 instead of F2
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
lea rsi, [rel .data]
; crc32 rax, byte [rel .data]
db 0xf2
db 0x66
db 0x48, 0x0f, 0x38, 0xf0, 0x05
dd 0x0000002A
; crc32 ebx, byte [rel .data + 1]
db 0xf2
db 0x66
db 0x0f, 0x38, 0xf0, 0x1d,
dd 0x00000021
.again:
cmp rdx, 256
je .done
; crc32 rcx, byte [rsi + rdx]
db 0xf2
db 0x66
db 0x48, 0x0f, 0x38, 0xf0, 0x0c, 0x16
add rdx, 1
jmp .again
.done:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+63
View File
@@ -0,0 +1,63 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000E330A81A",
"RBX": "0x00000000BE2DA0A5",
"RCX": "0x00000000ADBE9F64"
}
}
%endif
; This is a clone of the F2_F0 crc32 test with manually coded crc32 with prefix 66 instead of F2
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
lea rsi, [rel .data]
; crc32 rax, byte [rel .data]
db 0x66
db 0xf2
db 0x48, 0x0f, 0x38, 0xf0, 0x05
dd 0x0000002A
; crc32 ebx, byte [rel .data + 1]
db 0x66
db 0xf2
db 0x0f, 0x38, 0xf0, 0x1d,
dd 0x00000021
.again:
cmp rdx, 256
je .done
; crc32 rcx, byte [rsi + rdx]
db 0x66
db 0xf2
db 0x48, 0x0f, 0x38, 0xf0, 0x0c, 0x16
add rdx, 1
jmp .again
.done:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+103
View File
@@ -0,0 +1,103 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000C5727F5A",
"RBX": "0x00000000FAC690D7",
"RCX": "0x000000002AAF1F77",
"RDX": "0x00000000ADBE9F64",
"RSI": "0x00000000ADBE9F64",
"RDI": "0x00000000ADBE9F64"
}
}
%endif
; This is a clone of the F2_F1 crc32 test with manually coded crc32 with prefix 66 instead of F2
; This can't user operand size override on 32-bit instructions for testing since it WILL override to 16-bit
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
mov rsi, 0
mov rdi, 0
mov rbp, 0
lea rsp, [rel .data]
; crc32 eax, word [rel .data]
db 0x66 ; Operand size override
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x05
dd 0x0000006c
; crc32 ebx, dword [rel .data + 2]
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x1d,
dd 0x00000065
; crc32 rcx, qword [rel .data + 8]
db 0x66 ; Operand size override
db 0xf2 ; Prefix
db 0x48, 0x0f, 0x38, 0xf1, 0x0d,
dd 0x00000060
mov rbp, 0
.again16:
cmp rbp, 128
je .done16
; crc32 edx, word [rsp + rbp * 2]
db 0x66 ; Operand size override
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x14, 0x6c
add rbp, 1
jmp .again16
.done16:
mov rbp, 0
.again32:
cmp rbp, 64
je .done32
; crc32 esi, dword [rsp + rbp * 4]
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x34, 0xac
add rbp, 1
jmp .again32
.done32:
mov rbp, 0
.again64:
cmp rbp, 32
je .done64
; crc32 rdi, qword [rsp + rbp * 8]
db 0x66 ; Operand size override
db 0xf2 ; Prefix
db 0x48, 0x0f, 0x38, 0xf1, 0x3c, 0xec
add rbp, 1
jmp .again64
.done64:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+103
View File
@@ -0,0 +1,103 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000C5727F5A",
"RBX": "0x00000000FAC690D7",
"RCX": "0x000000002AAF1F77",
"RDX": "0x00000000ADBE9F64",
"RSI": "0x00000000ADBE9F64",
"RDI": "0x00000000ADBE9F64"
}
}
%endif
; This is a clone of the F2_F1 crc32 test with manually coded crc32 with prefix 66 instead of F2
; This can't user operand size override on 32-bit instructions for testing since it WILL override to 16-bit
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
mov rsi, 0
mov rdi, 0
mov rbp, 0
lea rsp, [rel .data]
; crc32 eax, word [rel .data]
db 0xf2 ; Prefix
db 0x66 ; Operand size override
db 0x0f, 0x38, 0xf1, 0x05
dd 0x0000006c
; crc32 ebx, dword [rel .data + 2]
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x1d,
dd 0x00000065
; crc32 rcx, qword [rel .data + 8]
db 0xf2 ; Prefix
db 0x66 ; Operand size override
db 0x48, 0x0f, 0x38, 0xf1, 0x0d,
dd 0x00000060
mov rbp, 0
.again16:
cmp rbp, 128
je .done16
; crc32 edx, word [rsp + rbp * 2]
db 0xf2 ; Prefix
db 0x66 ; Operand size override
db 0x0f, 0x38, 0xf1, 0x14, 0x6c
add rbp, 1
jmp .again16
.done16:
mov rbp, 0
.again32:
cmp rbp, 64
je .done32
; crc32 esi, dword [rsp + rbp * 4]
db 0xf2 ; Prefix
db 0x0f, 0x38, 0xf1, 0x34, 0xac
add rbp, 1
jmp .again32
.done32:
mov rbp, 0
.again64:
cmp rbp, 32
je .done64
; crc32 rdi, qword [rsp + rbp * 8]
db 0xf2 ; Prefix
db 0x66 ; Operand size override
db 0x48, 0x0f, 0x38, 0xf1, 0x3c, 0xec
add rbp, 1
jmp .again64
.done64:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+36
View File
@@ -0,0 +1,36 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x000000005c3bc5b0",
"RBX": "0x000000001dd5b1e5",
"RCX": "0x0000000015d1c92d"
}
}
%endif
mov rax, 0x41424344454647
mov rbx, 0x51525354555657
mov rcx, 0x61626364656667
mov rdx, 0x71727374757677
; crc32 rax, rbx
db 0x66 ; Override, Should be ignored
db 0xf2 ; Prefix
db 0x48 ; REX.W
db 0x0f, 0x38, 0xf1, 0xc3
; crc32 rbx, rcx
db 0xf2 ; Prefix
db 0x66 ; Override, Should be ignored
db 0x48 ; REX.W
db 0x0f, 0x38, 0xf1, 0xd9
; crc32 rcx, rdx
db 0x66 ; Override, Should be ignored
db 0xf2 ; Prefix
db 0x66 ; Override, Should be ignored
db 0x48 ; REX.W
db 0x0f, 0x38, 0xf1, 0xca
hlt
+49
View File
@@ -0,0 +1,49 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000E330A81A",
"RBX": "0x00000000BE2DA0A5",
"RCX": "0x00000000ADBE9F64"
}
}
%endif
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
lea rsi, [rel .data]
crc32 rax, byte [rel .data]
crc32 ebx, byte [rel .data + 1]
.again:
cmp rdx, 256
je .done
crc32 rcx, byte [rsi + rdx]
add rdx, 1
jmp .again
.done:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+78
View File
@@ -0,0 +1,78 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x00000000C5727F5A",
"RBX": "0x00000000FAC690D7",
"RCX": "0x000000002AAF1F77",
"RDX": "0x00000000ADBE9F64",
"RSI": "0x00000000ADBE9F64",
"RDI": "0x00000000ADBE9F64"
}
}
%endif
mov rax, 0
mov rbx, 0
mov rcx, 0
mov rdx, 0
mov rsi, 0
mov rdi, 0
mov rbp, 0
lea rsp, [rel .data]
crc32 eax, word [rel .data]
crc32 ebx, dword [rel .data + 2]
crc32 rcx, qword [rel .data + 8]
mov rbp, 0
.again16:
cmp rbp, 128
je .done16
crc32 edx, word [rsp + rbp * 2]
add rbp, 1
jmp .again16
.done16:
mov rbp, 0
.again32:
cmp rbp, 64
je .done32
crc32 esi, dword [rsp + rbp * 4]
add rbp, 1
jmp .again32
.done32:
mov rbp, 0
.again64:
cmp rbp, 32
je .done64
crc32 rdi, qword [rsp + rbp * 8]
add rbp, 1
jmp .again64
.done64:
hlt
align 16
; 256bytes of random data
.data:
db 0xe0, 0xfc, 0x2b, 0xa1, 0x06, 0x4f, 0x6c, 0xa7, 0x0f, 0x06, 0x6a, 0x1e, 0x7f, 0x76, 0x80, 0x9b
db 0xe0, 0x56, 0xed, 0xaa, 0xf3, 0xc3, 0x68, 0x68, 0xde, 0xe6, 0xe6, 0x94, 0xe2, 0xe9, 0xfc, 0xf0
db 0x6e, 0x35, 0xa8, 0x54, 0xd7, 0xab, 0x8b, 0x6c, 0x77, 0x5f, 0x92, 0xca, 0x25, 0xa6, 0x7e, 0x27
db 0xc7, 0xcd, 0x73, 0xec, 0x95, 0xd6, 0x6f, 0x6a, 0xbb, 0xae, 0xf2, 0xbb, 0x27, 0xb9, 0xa1, 0xdd
db 0x73, 0x4d, 0xd1, 0xc7, 0xd5, 0x2c, 0x31, 0x88, 0xfe, 0xe7, 0xdb, 0xfd, 0x1e, 0x1e, 0x09, 0x7f
db 0x14, 0xfa, 0x4e, 0x95, 0xef, 0xe6, 0x9a, 0xf2, 0xa0, 0x42, 0x62, 0x9a, 0xa4, 0xa8, 0x73, 0x82
db 0x0e, 0x0f, 0x16, 0x82, 0x38, 0x07, 0x12, 0x32, 0x07, 0x35, 0x92, 0xc1, 0x63, 0x07, 0x78, 0xb3
db 0xcb, 0x46, 0x19, 0x57, 0x2b, 0x37, 0x2a, 0x46, 0x1f, 0x04, 0x0e, 0x79, 0x3d, 0xcd, 0x8d, 0xa3
db 0x2b, 0xf3, 0x86, 0x2f, 0xab, 0xba, 0x57, 0x30, 0x2e, 0xd6, 0x2c, 0xf0, 0x46, 0x4f, 0x3f, 0xef
db 0xef, 0xd1, 0xbb, 0x85, 0x34, 0x4b, 0x3c, 0xde, 0x9e, 0x48, 0xa3, 0xb9, 0x8d, 0x71, 0xe3, 0x9d
db 0x09, 0x72, 0xfb, 0xde, 0x8a, 0x32, 0x50, 0x9d, 0x69, 0x98, 0xf1, 0xf6, 0x52, 0xeb, 0xf7, 0xee
db 0xd6, 0x99, 0xc2, 0xff, 0x30, 0x1c, 0x02, 0xce, 0x70, 0x05, 0xb2, 0xf1, 0x56, 0x9c, 0x0e, 0xa6
db 0x18, 0x62, 0xc4, 0xe2, 0x86, 0x38, 0x76, 0x30, 0x2f, 0xa1, 0xe4, 0xa7, 0x0e, 0x5d, 0x53, 0xeb
db 0x14, 0x45, 0xe0, 0xb7, 0xe1, 0xe8, 0x02, 0x68, 0x1a, 0xfe, 0x8e, 0xc1, 0x8f, 0xf2, 0xeb, 0x46
db 0x7f, 0x5d, 0x6a, 0x23, 0x46, 0x97, 0x2e, 0x03, 0x98, 0x12, 0x32, 0x8f, 0x54, 0x76, 0x59, 0xac
db 0xc8, 0x76, 0x5f, 0xc8, 0x71, 0x0c, 0xd3, 0xb6, 0xc5, 0x19, 0xea, 0xab, 0xa6, 0x2c, 0x1d, 0x88
+1 -1
View File
@@ -1 +1 @@
# No known failures
Test_X87/D9_F8.asm
+180
View File
@@ -0,0 +1,180 @@
%ifdef CONFIG
{
"RegData": {
"XMM0": ["0x0506070801020304", "0x0000000000000012"],
"XMM1": ["0x6576879821324354", "0x0000000000000000"],
"XMM2": ["0xB90984060D355548", "0x000000000000C03B"],
"XMM3": ["0xA83732340C01F070", "0x000000000000C03B"],
"XMM4": ["0xFFAA6DA43613FED0", "0x000000000000C03A"],
"XMM5": ["0x0000000000000001", "0x0000000000000000"],
"XMM6": ["0x0000000000000001", "0x0000000000008000"],
"XMM7": ["0x0000000000000000", "0x0000000000008000"],
"XMM8": ["0x0000000000000000", "0x0000000000008000"],
"XMM9": ["0x0000000000000000", "0x0000000000000000"],
"XMM10": ["0x0000000000000001", "0x0000000000008000"],
"XMM11": ["0x0000000000000001", "0x0000000000000000"]
}
}
%endif
fbld [rel .data_0]
fbstp [rel .res_data_0]
movups xmm0, [rel .res_data_0]
fbld [rel .data_1]
fbstp [rel .res_data_1]
movups xmm1, [rel .res_data_1]
; Check encoding of invalid BCD
fbld [rel .data_2]
fstp tword [rel .res_data_2]
movups xmm2, [rel .res_data_2]
fbld [rel .data_3]
fstp tword [rel .res_data_3]
movups xmm3, [rel .res_data_3]
fbld [rel .data_4]
fstp tword [rel .res_data_4]
movups xmm4, [rel .res_data_4]
; Some special values
fld tword [rel .data_5]
fbstp [rel .res_data_5]
movups xmm5, [rel .res_data_5]
fld tword [rel .data_6]
fbstp [rel .res_data_6]
movups xmm6, [rel .res_data_6]
fld tword [rel .data_7]
fbstp [rel .res_data_7]
movups xmm7, [rel .res_data_7]
; Values that choose +- 0 or +-1 depending on rounding mode
; -1 < F < -0
; +0 < F < +1
fld tword [rel .data_8]
fbstp [rel .res_data_8]
movups xmm8, [rel .res_data_8]
fld tword [rel .data_9]
fbstp [rel .res_data_9]
movups xmm9, [rel .res_data_9]
; Swap control word
fnstcw [rel .cw]
mov ax, [rel .cw]
and ax, ~(3 << 10)
or eax, 1 << 10 ; Round down
mov [rel .cw], ax
fldcw [rel .cw]
fld tword [rel .data_10]
fbstp [rel .res_data_10]
movups xmm10, [rel .res_data_10]
; Swap control word
fnstcw [rel .cw]
mov ax, [rel .cw]
and ax, ~(3 << 10)
or eax, 2 << 10 ; Round up
mov [rel .cw], ax
fldcw [rel .cw]
fld tword [rel .data_11]
fbstp [rel .res_data_11]
movups xmm11, [rel .res_data_11]
; Values that generate Invalicating floating point operation exception
; -inf
; +inf
; Negative value too large for destination format
; Positive value too large for destination format
; NaN
; On IA the indefinite BCD result is still stored to memory
; XXX: We don't support IA on this
hlt
.cw:
dw 0
.data_0:
dd 0x01020304
dd 0x05060708
dd 0x09101112
dd 0x13141516
.data_1:
dd 0x21324354
dd 0x65768798
dd 0x00000000
dd 0x00000000
.data_2:
dd 0xFFFFFFFF
dd 0xFFFFFFFF
dd 0xFFFFFFFF
dd 0xFFFFFFFF
.data_3:
dd 0xF0F0F0F0
dd 0xF0F0F0F0
dd 0xF0F0F0F0
dd 0xF0F0F0F0
.data_4:
dd 0x0A0B0C0D
dd 0x0E0FAAAB
dd 0xACADAEAF
dd 0xBABBBCBD
.data_5:
dt 1.0
.data_6:
dt -1.0
.data_7:
dt -0.0
.data_8:
dt -0.5
.data_9:
dt 0.5
.data_10:
dt -0.5
.data_11:
dt 0.5
.res_data_0:
dq 0
dq 0
.res_data_1:
dq 0
dq 0
.res_data_2:
dq 0
dq 0
.res_data_3:
dq 0
dq 0
.res_data_4:
dq 0
dq 0
.res_data_5:
dq 0
dq 0
.res_data_6:
dq 0
dq 0
.res_data_7:
dq 0
dq 0
.res_data_8:
dq 0
dq 0
.res_data_9:
dq 0
dq 0
.res_data_10:
dq 0
dq 0
.res_data_11:
dq 0
dq 0
+2 -2
View File
@@ -16,7 +16,7 @@ foreach(POSIX_TEST ${POSIX_TESTS})
"${CMAKE_SOURCE_DIR}/unittests/POSIX/Disabled_Tests"
"${TEST_NAME}"
"${CMAKE_BINARY_DIR}/Bin/FEXLoader"
"--no-silent" "-c" "irint" "-n" "500" "-R" $ENV{ROOTFS} "--"
"--no-silent" "-c" "irint" "-n" "500" "--"
"${POSIX_TEST}")
add_test(NAME "${TEST_NAME}.jit.posix"
@@ -26,7 +26,7 @@ foreach(POSIX_TEST ${POSIX_TESTS})
"${CMAKE_SOURCE_DIR}/unittests/POSIX/Disabled_Tests"
"${TEST_NAME}"
"${CMAKE_BINARY_DIR}/Bin/FEXLoader"
"--no-silent" "-c" "irjit" "-n" "500" "-R" $ENV{ROOTFS} "--"
"--no-silent" "-c" "irjit" "-n" "500" "--"
"${POSIX_TEST}")
endforeach()
+38 -264
View File
@@ -12,6 +12,7 @@ conformance-interfaces-mmap-6-3.test
# These tests take too long to run, and might timeout
conformance-interfaces-clock-1-1.test
conformance-interfaces-clock_gettime-4-1.test
conformance-interfaces-clock_getcpuclockid-1-1.test
# These tests fail when run natively on x86-64 host
conformance-interfaces-clock_getcpuclockid-2-1.test
@@ -193,262 +194,6 @@ src-conformance-interfaces-shm_open-24-1.test
src-conformance-interfaces-shm_open-29-1.test
src-conformance-interfaces-shm_open-6-1.test
# These test use signals
conformance-interfaces-sigsuspend-6-1.test
conformance-interfaces-sigsuspend-4-1.test
conformance-interfaces-sigsuspend-3-1.test
conformance-interfaces-sigsuspend-1-1.test
conformance-interfaces-sigtimedwait-5-1.test
conformance-interfaces-sigwaitinfo-3-1.test
conformance-interfaces-strftime-1-1.test
conformance-interfaces-sigqueue-3-1.test
conformance-interfaces-sigqueue-12-1.test
conformance-interfaces-sigqueue-1-1.test
conformance-interfaces-sigqueue-9-1.test
conformance-interfaces-aio_read-2-1.test
conformance-interfaces-raise-1-2.test
conformance-interfaces-sigaction-17-11.test
conformance-interfaces-sigaction-4-49.test
conformance-interfaces-sigaction-21-1.test
conformance-interfaces-sigaction-4-16.test
conformance-interfaces-sigaction-22-7.test
conformance-interfaces-sigaction-22-23.test
conformance-interfaces-sigaction-22-6.test
conformance-interfaces-sigaction-4-26.test
conformance-interfaces-sigaction-22-16.test
conformance-interfaces-sigaction-17-25.test
conformance-interfaces-sigaction-4-5.test
conformance-interfaces-sigaction-25-17.test
conformance-interfaces-sigaction-4-52.test
conformance-interfaces-sigaction-4-32.test
conformance-interfaces-sigaction-17-15.test
conformance-interfaces-sigaction-17-6.test
conformance-interfaces-sigaction-4-6.test
conformance-interfaces-sigaction-17-13.test
conformance-interfaces-sigaction-4-23.test
conformance-interfaces-sigaction-4-13.test
conformance-interfaces-sigaction-17-7.test
conformance-interfaces-sigaction-22-11.test
conformance-interfaces-sigaction-25-6.test
conformance-interfaces-sigaction-17-19.test
conformance-interfaces-sigaction-25-11.test
conformance-interfaces-sigaction-17-14.test
conformance-interfaces-sigaction-17-3.test
conformance-interfaces-sigaction-22-25.test
conformance-interfaces-sigaction-22-5.test
conformance-interfaces-sigaction-17-26.test
conformance-interfaces-sigaction-4-17.test
conformance-interfaces-sigaction-25-3.test
conformance-interfaces-sigaction-17-16.test
conformance-interfaces-sigaction-22-4.test
conformance-interfaces-sigaction-4-40.test
conformance-interfaces-sigaction-4-27.test
conformance-interfaces-sigaction-17-8.test
conformance-interfaces-sigaction-17-24.test
conformance-interfaces-sigaction-4-22.test
conformance-interfaces-sigaction-4-21.test
conformance-interfaces-sigaction-25-26.test
conformance-interfaces-sigaction-22-26.test
conformance-interfaces-sigaction-4-25.test
conformance-interfaces-sigaction-4-1.test
conformance-interfaces-sigaction-25-7.test
conformance-interfaces-sigaction-4-14.test
conformance-interfaces-sigaction-17-4.test
conformance-interfaces-sigaction-4-34.test
conformance-interfaces-sigaction-17-17.test
conformance-interfaces-sigaction-4-47.test
conformance-interfaces-sigaction-22-3.test
conformance-interfaces-sigaction-25-9.test
conformance-interfaces-sigaction-25-8.test
conformance-interfaces-sigaction-25-22.test
conformance-interfaces-sigaction-25-19.test
conformance-interfaces-sigaction-25-21.test
conformance-interfaces-sigaction-22-1.test
conformance-interfaces-sigaction-4-29.test
conformance-interfaces-sigaction-22-10.test
conformance-interfaces-sigaction-22-17.test
conformance-interfaces-sigaction-4-15.test
conformance-interfaces-sigaction-4-12.test
conformance-interfaces-sigaction-4-18.test
conformance-interfaces-sigaction-22-22.test
conformance-interfaces-sigaction-22-18.test
conformance-interfaces-sigaction-17-1.test
conformance-interfaces-sigaction-4-36.test
conformance-interfaces-sigaction-4-51.test
conformance-interfaces-sigaction-22-12.test
conformance-interfaces-sigaction-4-50.test
conformance-interfaces-sigaction-25-10.test
conformance-interfaces-sigaction-22-13.test
conformance-interfaces-sigaction-4-24.test
conformance-interfaces-sigaction-25-2.test
conformance-interfaces-sigaction-25-14.test
conformance-interfaces-sigaction-22-2.test
conformance-interfaces-sigaction-4-8.test
conformance-interfaces-sigaction-22-14.test
conformance-interfaces-sigaction-17-18.test
conformance-interfaces-sigaction-22-21.test
conformance-interfaces-sigaction-25-16.test
conformance-interfaces-sigaction-4-28.test
conformance-interfaces-sigaction-25-12.test
conformance-interfaces-sigaction-4-39.test
conformance-interfaces-sigaction-4-33.test
conformance-interfaces-sigaction-4-42.test
conformance-interfaces-sigaction-17-12.test
conformance-interfaces-sigaction-22-9.test
conformance-interfaces-sigaction-17-9.test
conformance-interfaces-sigaction-4-19.test
conformance-interfaces-sigaction-25-20.test
conformance-interfaces-sigaction-17-20.test
conformance-interfaces-sigaction-4-41.test
conformance-interfaces-sigaction-22-8.test
conformance-interfaces-sigaction-25-1.test
conformance-interfaces-sigaction-25-15.test
conformance-interfaces-sigaction-25-5.test
conformance-interfaces-sigaction-17-10.test
conformance-interfaces-sigaction-4-9.test
conformance-interfaces-sigaction-25-4.test
conformance-interfaces-sigaction-25-13.test
conformance-interfaces-sigaction-22-24.test
conformance-interfaces-sigaction-4-4.test
conformance-interfaces-sigaction-17-23.test
conformance-interfaces-sigaction-22-19.test
conformance-interfaces-sigaction-4-10.test
conformance-interfaces-sigaction-4-48.test
conformance-interfaces-sigaction-25-24.test
conformance-interfaces-sigaction-4-46.test
conformance-interfaces-sigaction-17-22.test
conformance-interfaces-sigaction-4-3.test
conformance-interfaces-sigaction-4-7.test
conformance-interfaces-sigaction-25-18.test
conformance-interfaces-sigaction-17-2.test
conformance-interfaces-sigaction-17-21.test
conformance-interfaces-sigaction-4-35.test
conformance-interfaces-sigaction-4-37.test
conformance-interfaces-sigaction-17-5.test
conformance-interfaces-sigaction-4-30.test
conformance-interfaces-sigaction-4-2.test
conformance-interfaces-sigaction-4-45.test
conformance-interfaces-sigaction-4-38.test
conformance-interfaces-sigaction-4-31.test
conformance-interfaces-sigaction-4-20.test
conformance-interfaces-sigaction-4-43.test
conformance-interfaces-sigaction-25-25.test
conformance-interfaces-sigaction-22-15.test
conformance-interfaces-sigaction-4-11.test
conformance-interfaces-sigaction-22-20.test
conformance-interfaces-sigaction-25-23.test
conformance-interfaces-sigaction-4-44.test
conformance-interfaces-sigaction-1-10.test
conformance-interfaces-sigaction-1-11.test
conformance-interfaces-sigaction-1-12.test
conformance-interfaces-sigaction-1-13.test
conformance-interfaces-sigaction-1-14.test
conformance-interfaces-sigaction-1-15.test
conformance-interfaces-sigaction-1-16.test
conformance-interfaces-sigaction-1-17.test
conformance-interfaces-sigaction-1-18.test
conformance-interfaces-sigaction-1-19.test
conformance-interfaces-sigaction-1-1.test
conformance-interfaces-sigaction-1-20.test
conformance-interfaces-sigaction-12-10.test
conformance-interfaces-sigaction-12-11.test
conformance-interfaces-sigaction-12-12.test
conformance-interfaces-sigaction-12-13.test
conformance-interfaces-sigaction-12-14.test
conformance-interfaces-sigaction-12-15.test
conformance-interfaces-sigaction-12-16.test
conformance-interfaces-sigaction-12-17.test
conformance-interfaces-sigaction-12-18.test
conformance-interfaces-sigaction-12-19.test
conformance-interfaces-sigaction-1-21.test
conformance-interfaces-sigaction-12-1.test
conformance-interfaces-sigaction-12-20.test
conformance-interfaces-sigaction-12-21.test
conformance-interfaces-sigaction-12-22.test
conformance-interfaces-sigaction-12-23.test
conformance-interfaces-sigaction-12-24.test
conformance-interfaces-sigaction-12-25.test
conformance-interfaces-sigaction-12-26.test
conformance-interfaces-sigaction-1-22.test
conformance-interfaces-sigaction-12-2.test
conformance-interfaces-sigaction-1-23.test
conformance-interfaces-sigaction-12-3.test
conformance-interfaces-sigaction-1-24.test
conformance-interfaces-sigaction-12-4.test
conformance-interfaces-sigaction-1-25.test
conformance-interfaces-sigaction-12-5.test
conformance-interfaces-sigaction-1-26.test
conformance-interfaces-sigaction-12-6.test
conformance-interfaces-sigaction-12-7.test
conformance-interfaces-sigaction-12-8.test
conformance-interfaces-sigaction-12-9.test
conformance-interfaces-sigaction-1-2.test
conformance-interfaces-sigaction-1-3.test
conformance-interfaces-sigaction-1-4.test
conformance-interfaces-sigaction-1-5.test
conformance-interfaces-sigaction-1-6.test
conformance-interfaces-sigaction-1-7.test
conformance-interfaces-sigaction-1-8.test
conformance-interfaces-sigaction-1-9.test
conformance-interfaces-sigaction-3-10.test
conformance-interfaces-sigaction-3-11.test
conformance-interfaces-sigaction-3-12.test
conformance-interfaces-sigaction-3-13.test
conformance-interfaces-sigaction-3-14.test
conformance-interfaces-sigaction-3-15.test
conformance-interfaces-sigaction-3-16.test
conformance-interfaces-sigaction-3-17.test
conformance-interfaces-sigaction-3-18.test
conformance-interfaces-sigaction-3-19.test
conformance-interfaces-sigaction-3-1.test
conformance-interfaces-sigaction-3-20.test
conformance-interfaces-sigaction-3-21.test
conformance-interfaces-sigaction-3-22.test
conformance-interfaces-sigaction-3-23.test
conformance-interfaces-sigaction-3-24.test
conformance-interfaces-sigaction-3-25.test
conformance-interfaces-sigaction-3-26.test
conformance-interfaces-sigaction-3-2.test
conformance-interfaces-sigaction-3-3.test
conformance-interfaces-sigaction-3-4.test
conformance-interfaces-sigaction-3-5.test
conformance-interfaces-sigaction-3-6.test
conformance-interfaces-sigaction-3-7.test
conformance-interfaces-sigaction-3-8.test
conformance-interfaces-sigaction-3-9.test
conformance-interfaces-sigaction-6-10.test
conformance-interfaces-sigaction-6-11.test
conformance-interfaces-sigaction-6-12.test
conformance-interfaces-sigaction-6-13.test
conformance-interfaces-sigaction-6-14.test
conformance-interfaces-sigaction-6-15.test
conformance-interfaces-sigaction-6-16.test
conformance-interfaces-sigaction-6-17.test
conformance-interfaces-sigaction-6-18.test
conformance-interfaces-sigaction-6-19.test
conformance-interfaces-sigaction-6-1.test
conformance-interfaces-sigaction-6-20.test
conformance-interfaces-sigaction-6-21.test
conformance-interfaces-sigaction-6-22.test
conformance-interfaces-sigaction-6-23.test
conformance-interfaces-sigaction-6-24.test
conformance-interfaces-sigaction-6-25.test
conformance-interfaces-sigaction-6-26.test
conformance-interfaces-sigaction-6-2.test
conformance-interfaces-sigaction-6-3.test
conformance-interfaces-sigaction-6-4.test
conformance-interfaces-sigaction-6-5.test
conformance-interfaces-sigaction-6-6.test
conformance-interfaces-sigaction-6-7.test
conformance-interfaces-sigaction-6-8.test
conformance-interfaces-sigaction-6-9.test
conformance-interfaces-sigaltstack-2-1.test
conformance-interfaces-sigaltstack-3-1.test
conformance-interfaces-sigaltstack-8-1.test
conformance-interfaces-sigprocmask-9-1.test
conformance-interfaces-sigrelse-1-1.test
# These use signals and will fail on x86-64
conformance-interfaces-sigaction-8-1.test
conformance-interfaces-sigaction-8-10.test
@@ -542,14 +287,12 @@ conformance-interfaces-raise-2-1.test
# Signals change this behaviour
conformance-interfaces-sigpending-1-1.test
conformance-interfaces-sigpending-2-1.test
conformance-interfaces-sigwait-1-1.test
conformance-interfaces-sigwait-2-1.test
conformance-interfaces-sigwait-3-1.test
conformance-interfaces-sigwait-8-1.test
conformance-interfaces-sigwaitinfo-1-1.test
conformance-interfaces-sigwaitinfo-5-1.test
conformance-interfaces-sigwaitinfo-6-1.test
conformance-interfaces-sigwaitinfo-9-1.test
# Both of these pass signals in their handler
# Since we exit the signal handler our sa_mask changes
# and we don't currently resolve this
conformance-interfaces-sigpending-1-2.test
conformance-interfaces-sigpending-1-3.test
# Causes long timeout with signal change
conformance-interfaces-mmap-11-2.test
@@ -601,3 +344,34 @@ conformance-interfaces-killpg-8-1.test
# It puts the thread to sleep for 1 second and expects to wake up within 10ms of the timer
# If the kernel is doing other things then this 10ms time is too strict and fails periodically
conformance-interfaces-sigtimedwait-1-1.test
# We accidentally pass through signals to the guest that shouldn't be
conformance-interfaces-sigsuspend-1-1.test
# Received signal in handler even though it should be masked
conformance-interfaces-sigaction-25-1.test
conformance-interfaces-sigaction-25-2.test
conformance-interfaces-sigaction-25-3.test
conformance-interfaces-sigaction-25-4.test
conformance-interfaces-sigaction-25-5.test
conformance-interfaces-sigaction-25-6.test
conformance-interfaces-sigaction-25-7.test
conformance-interfaces-sigaction-25-8.test
conformance-interfaces-sigaction-25-9.test
conformance-interfaces-sigaction-25-10.test
conformance-interfaces-sigaction-25-11.test
conformance-interfaces-sigaction-25-12.test
conformance-interfaces-sigaction-25-13.test
conformance-interfaces-sigaction-25-14.test
conformance-interfaces-sigaction-25-15.test
conformance-interfaces-sigaction-25-16.test
conformance-interfaces-sigaction-25-17.test
conformance-interfaces-sigaction-25-18.test
conformance-interfaces-sigaction-25-19.test
conformance-interfaces-sigaction-25-20.test
conformance-interfaces-sigaction-25-21.test
conformance-interfaces-sigaction-25-22.test
conformance-interfaces-sigaction-25-23.test
conformance-interfaces-sigaction-25-24.test
conformance-interfaces-sigaction-25-25.test
conformance-interfaces-sigaction-25-26.test
+35 -273
View File
@@ -1,20 +1,3 @@
# these fail
conformance-interfaces-clock-1-1.test
conformance-interfaces-clock_getcpuclockid-2-1.testconformance-interfaces-clock-1-1.test
conformance-interfaces-munmap-2-1.test
conformance-interfaces-sigqueue-12-1.test
conformance-interfaces-sigqueue-3-1.test
conformance-interfaces-sigqueue-9-1.test
conformance-interfaces-sigtimedwait-5-1.test
conformance-interfaces-sigwait-1-1.test
conformance-interfaces-sigwait-2-1.test
conformance-interfaces-sigwait-3-1.test
conformance-interfaces-sigwait-8-1.test
conformance-interfaces-sigwaitinfo-1-1.test
conformance-interfaces-sigwaitinfo-5-1.test
conformance-interfaces-sigwaitinfo-6-1.test
conformance-interfaces-sigwaitinfo-9-1.test
# these are disabled
# These tests are inconsistent
conformance-interfaces-mmap-12-1.test
@@ -25,6 +8,7 @@ conformance-interfaces-mmap-6-3.test
# These tests take too long to run, and might timeout
conformance-interfaces-clock-1-1.test
conformance-interfaces-clock_gettime-4-1.test
conformance-interfaces-clock_getcpuclockid-1-1.test
# These tests fail when run natively on x86-64 host
conformance-interfaces-clock_getcpuclockid-2-1.test
@@ -206,262 +190,6 @@ src-conformance-interfaces-shm_open-24-1.test
src-conformance-interfaces-shm_open-29-1.test
src-conformance-interfaces-shm_open-6-1.test
# These test use signals
conformance-interfaces-sigsuspend-6-1.test
conformance-interfaces-sigsuspend-4-1.test
conformance-interfaces-sigsuspend-3-1.test
conformance-interfaces-sigsuspend-1-1.test
conformance-interfaces-sigtimedwait-5-1.test
conformance-interfaces-sigwaitinfo-3-1.test
conformance-interfaces-strftime-1-1.test
conformance-interfaces-sigqueue-3-1.test
conformance-interfaces-sigqueue-12-1.test
conformance-interfaces-sigqueue-1-1.test
conformance-interfaces-sigqueue-9-1.test
conformance-interfaces-aio_read-2-1.test
conformance-interfaces-raise-1-2.test
conformance-interfaces-sigaction-17-11.test
conformance-interfaces-sigaction-4-49.test
conformance-interfaces-sigaction-21-1.test
conformance-interfaces-sigaction-4-16.test
conformance-interfaces-sigaction-22-7.test
conformance-interfaces-sigaction-22-23.test
conformance-interfaces-sigaction-22-6.test
conformance-interfaces-sigaction-4-26.test
conformance-interfaces-sigaction-22-16.test
conformance-interfaces-sigaction-17-25.test
conformance-interfaces-sigaction-4-5.test
conformance-interfaces-sigaction-25-17.test
conformance-interfaces-sigaction-4-52.test
conformance-interfaces-sigaction-4-32.test
conformance-interfaces-sigaction-17-15.test
conformance-interfaces-sigaction-17-6.test
conformance-interfaces-sigaction-4-6.test
conformance-interfaces-sigaction-17-13.test
conformance-interfaces-sigaction-4-23.test
conformance-interfaces-sigaction-4-13.test
conformance-interfaces-sigaction-17-7.test
conformance-interfaces-sigaction-22-11.test
conformance-interfaces-sigaction-25-6.test
conformance-interfaces-sigaction-17-19.test
conformance-interfaces-sigaction-25-11.test
conformance-interfaces-sigaction-17-14.test
conformance-interfaces-sigaction-17-3.test
conformance-interfaces-sigaction-22-25.test
conformance-interfaces-sigaction-22-5.test
conformance-interfaces-sigaction-17-26.test
conformance-interfaces-sigaction-4-17.test
conformance-interfaces-sigaction-25-3.test
conformance-interfaces-sigaction-17-16.test
conformance-interfaces-sigaction-22-4.test
conformance-interfaces-sigaction-4-40.test
conformance-interfaces-sigaction-4-27.test
conformance-interfaces-sigaction-17-8.test
conformance-interfaces-sigaction-17-24.test
conformance-interfaces-sigaction-4-22.test
conformance-interfaces-sigaction-4-21.test
conformance-interfaces-sigaction-25-26.test
conformance-interfaces-sigaction-22-26.test
conformance-interfaces-sigaction-4-25.test
conformance-interfaces-sigaction-4-1.test
conformance-interfaces-sigaction-25-7.test
conformance-interfaces-sigaction-4-14.test
conformance-interfaces-sigaction-17-4.test
conformance-interfaces-sigaction-4-34.test
conformance-interfaces-sigaction-17-17.test
conformance-interfaces-sigaction-4-47.test
conformance-interfaces-sigaction-22-3.test
conformance-interfaces-sigaction-25-9.test
conformance-interfaces-sigaction-25-8.test
conformance-interfaces-sigaction-25-22.test
conformance-interfaces-sigaction-25-19.test
conformance-interfaces-sigaction-25-21.test
conformance-interfaces-sigaction-22-1.test
conformance-interfaces-sigaction-4-29.test
conformance-interfaces-sigaction-22-10.test
conformance-interfaces-sigaction-22-17.test
conformance-interfaces-sigaction-4-15.test
conformance-interfaces-sigaction-4-12.test
conformance-interfaces-sigaction-4-18.test
conformance-interfaces-sigaction-22-22.test
conformance-interfaces-sigaction-22-18.test
conformance-interfaces-sigaction-17-1.test
conformance-interfaces-sigaction-4-36.test
conformance-interfaces-sigaction-4-51.test
conformance-interfaces-sigaction-22-12.test
conformance-interfaces-sigaction-4-50.test
conformance-interfaces-sigaction-25-10.test
conformance-interfaces-sigaction-22-13.test
conformance-interfaces-sigaction-4-24.test
conformance-interfaces-sigaction-25-2.test
conformance-interfaces-sigaction-25-14.test
conformance-interfaces-sigaction-22-2.test
conformance-interfaces-sigaction-4-8.test
conformance-interfaces-sigaction-22-14.test
conformance-interfaces-sigaction-17-18.test
conformance-interfaces-sigaction-22-21.test
conformance-interfaces-sigaction-25-16.test
conformance-interfaces-sigaction-4-28.test
conformance-interfaces-sigaction-25-12.test
conformance-interfaces-sigaction-4-39.test
conformance-interfaces-sigaction-4-33.test
conformance-interfaces-sigaction-4-42.test
conformance-interfaces-sigaction-17-12.test
conformance-interfaces-sigaction-22-9.test
conformance-interfaces-sigaction-17-9.test
conformance-interfaces-sigaction-4-19.test
conformance-interfaces-sigaction-25-20.test
conformance-interfaces-sigaction-17-20.test
conformance-interfaces-sigaction-4-41.test
conformance-interfaces-sigaction-22-8.test
conformance-interfaces-sigaction-25-1.test
conformance-interfaces-sigaction-25-15.test
conformance-interfaces-sigaction-25-5.test
conformance-interfaces-sigaction-17-10.test
conformance-interfaces-sigaction-4-9.test
conformance-interfaces-sigaction-25-4.test
conformance-interfaces-sigaction-25-13.test
conformance-interfaces-sigaction-22-24.test
conformance-interfaces-sigaction-4-4.test
conformance-interfaces-sigaction-17-23.test
conformance-interfaces-sigaction-22-19.test
conformance-interfaces-sigaction-4-10.test
conformance-interfaces-sigaction-4-48.test
conformance-interfaces-sigaction-25-24.test
conformance-interfaces-sigaction-4-46.test
conformance-interfaces-sigaction-17-22.test
conformance-interfaces-sigaction-4-3.test
conformance-interfaces-sigaction-4-7.test
conformance-interfaces-sigaction-25-18.test
conformance-interfaces-sigaction-17-2.test
conformance-interfaces-sigaction-17-21.test
conformance-interfaces-sigaction-4-35.test
conformance-interfaces-sigaction-4-37.test
conformance-interfaces-sigaction-17-5.test
conformance-interfaces-sigaction-4-30.test
conformance-interfaces-sigaction-4-2.test
conformance-interfaces-sigaction-4-45.test
conformance-interfaces-sigaction-4-38.test
conformance-interfaces-sigaction-4-31.test
conformance-interfaces-sigaction-4-20.test
conformance-interfaces-sigaction-4-43.test
conformance-interfaces-sigaction-25-25.test
conformance-interfaces-sigaction-22-15.test
conformance-interfaces-sigaction-4-11.test
conformance-interfaces-sigaction-22-20.test
conformance-interfaces-sigaction-25-23.test
conformance-interfaces-sigaction-4-44.test
conformance-interfaces-sigaction-1-10.test
conformance-interfaces-sigaction-1-11.test
conformance-interfaces-sigaction-1-12.test
conformance-interfaces-sigaction-1-13.test
conformance-interfaces-sigaction-1-14.test
conformance-interfaces-sigaction-1-15.test
conformance-interfaces-sigaction-1-16.test
conformance-interfaces-sigaction-1-17.test
conformance-interfaces-sigaction-1-18.test
conformance-interfaces-sigaction-1-19.test
conformance-interfaces-sigaction-1-1.test
conformance-interfaces-sigaction-1-20.test
conformance-interfaces-sigaction-12-10.test
conformance-interfaces-sigaction-12-11.test
conformance-interfaces-sigaction-12-12.test
conformance-interfaces-sigaction-12-13.test
conformance-interfaces-sigaction-12-14.test
conformance-interfaces-sigaction-12-15.test
conformance-interfaces-sigaction-12-16.test
conformance-interfaces-sigaction-12-17.test
conformance-interfaces-sigaction-12-18.test
conformance-interfaces-sigaction-12-19.test
conformance-interfaces-sigaction-1-21.test
conformance-interfaces-sigaction-12-1.test
conformance-interfaces-sigaction-12-20.test
conformance-interfaces-sigaction-12-21.test
conformance-interfaces-sigaction-12-22.test
conformance-interfaces-sigaction-12-23.test
conformance-interfaces-sigaction-12-24.test
conformance-interfaces-sigaction-12-25.test
conformance-interfaces-sigaction-12-26.test
conformance-interfaces-sigaction-1-22.test
conformance-interfaces-sigaction-12-2.test
conformance-interfaces-sigaction-1-23.test
conformance-interfaces-sigaction-12-3.test
conformance-interfaces-sigaction-1-24.test
conformance-interfaces-sigaction-12-4.test
conformance-interfaces-sigaction-1-25.test
conformance-interfaces-sigaction-12-5.test
conformance-interfaces-sigaction-1-26.test
conformance-interfaces-sigaction-12-6.test
conformance-interfaces-sigaction-12-7.test
conformance-interfaces-sigaction-12-8.test
conformance-interfaces-sigaction-12-9.test
conformance-interfaces-sigaction-1-2.test
conformance-interfaces-sigaction-1-3.test
conformance-interfaces-sigaction-1-4.test
conformance-interfaces-sigaction-1-5.test
conformance-interfaces-sigaction-1-6.test
conformance-interfaces-sigaction-1-7.test
conformance-interfaces-sigaction-1-8.test
conformance-interfaces-sigaction-1-9.test
conformance-interfaces-sigaction-3-10.test
conformance-interfaces-sigaction-3-11.test
conformance-interfaces-sigaction-3-12.test
conformance-interfaces-sigaction-3-13.test
conformance-interfaces-sigaction-3-14.test
conformance-interfaces-sigaction-3-15.test
conformance-interfaces-sigaction-3-16.test
conformance-interfaces-sigaction-3-17.test
conformance-interfaces-sigaction-3-18.test
conformance-interfaces-sigaction-3-19.test
conformance-interfaces-sigaction-3-1.test
conformance-interfaces-sigaction-3-20.test
conformance-interfaces-sigaction-3-21.test
conformance-interfaces-sigaction-3-22.test
conformance-interfaces-sigaction-3-23.test
conformance-interfaces-sigaction-3-24.test
conformance-interfaces-sigaction-3-25.test
conformance-interfaces-sigaction-3-26.test
conformance-interfaces-sigaction-3-2.test
conformance-interfaces-sigaction-3-3.test
conformance-interfaces-sigaction-3-4.test
conformance-interfaces-sigaction-3-5.test
conformance-interfaces-sigaction-3-6.test
conformance-interfaces-sigaction-3-7.test
conformance-interfaces-sigaction-3-8.test
conformance-interfaces-sigaction-3-9.test
conformance-interfaces-sigaction-6-10.test
conformance-interfaces-sigaction-6-11.test
conformance-interfaces-sigaction-6-12.test
conformance-interfaces-sigaction-6-13.test
conformance-interfaces-sigaction-6-14.test
conformance-interfaces-sigaction-6-15.test
conformance-interfaces-sigaction-6-16.test
conformance-interfaces-sigaction-6-17.test
conformance-interfaces-sigaction-6-18.test
conformance-interfaces-sigaction-6-19.test
conformance-interfaces-sigaction-6-1.test
conformance-interfaces-sigaction-6-20.test
conformance-interfaces-sigaction-6-21.test
conformance-interfaces-sigaction-6-22.test
conformance-interfaces-sigaction-6-23.test
conformance-interfaces-sigaction-6-24.test
conformance-interfaces-sigaction-6-25.test
conformance-interfaces-sigaction-6-26.test
conformance-interfaces-sigaction-6-2.test
conformance-interfaces-sigaction-6-3.test
conformance-interfaces-sigaction-6-4.test
conformance-interfaces-sigaction-6-5.test
conformance-interfaces-sigaction-6-6.test
conformance-interfaces-sigaction-6-7.test
conformance-interfaces-sigaction-6-8.test
conformance-interfaces-sigaction-6-9.test
conformance-interfaces-sigaltstack-2-1.test
conformance-interfaces-sigaltstack-3-1.test
conformance-interfaces-sigaltstack-8-1.test
conformance-interfaces-sigprocmask-9-1.test
conformance-interfaces-sigrelse-1-1.test
# These use signals and will fail on x86-64
conformance-interfaces-sigaction-8-1.test
conformance-interfaces-sigaction-8-10.test
@@ -625,3 +353,37 @@ conformance-interfaces-sigtimedwait-1-1.test
# Sadly since this test is so old it happens to set SS_ONSTACK which is a no-op on
# newer kernels. So this is expected to fail.
conformance-interfaces-sigaltstack-11-1.test
# This test fails since FEX doesn't fully support SS_DISABLE
conformance-interfaces-sigaltstack-2-1.test
# We accidentally pass through signals to the guest that shouldn't be
conformance-interfaces-sigsuspend-1-1.test
# Received signal in handler even though it should be masked
conformance-interfaces-sigaction-25-1.test
conformance-interfaces-sigaction-25-2.test
conformance-interfaces-sigaction-25-3.test
conformance-interfaces-sigaction-25-4.test
conformance-interfaces-sigaction-25-5.test
conformance-interfaces-sigaction-25-6.test
conformance-interfaces-sigaction-25-7.test
conformance-interfaces-sigaction-25-8.test
conformance-interfaces-sigaction-25-9.test
conformance-interfaces-sigaction-25-10.test
conformance-interfaces-sigaction-25-11.test
conformance-interfaces-sigaction-25-12.test
conformance-interfaces-sigaction-25-13.test
conformance-interfaces-sigaction-25-14.test
conformance-interfaces-sigaction-25-15.test
conformance-interfaces-sigaction-25-16.test
conformance-interfaces-sigaction-25-17.test
conformance-interfaces-sigaction-25-18.test
conformance-interfaces-sigaction-25-19.test
conformance-interfaces-sigaction-25-20.test
conformance-interfaces-sigaction-25-21.test
conformance-interfaces-sigaction-25-22.test
conformance-interfaces-sigaction-25-23.test
conformance-interfaces-sigaction-25-24.test
conformance-interfaces-sigaction-25-25.test
conformance-interfaces-sigaction-25-26.test
+1 -1
View File
@@ -18,7 +18,7 @@ foreach(TEST ${TESTS})
"${CMAKE_SOURCE_DIR}/unittests/gcc-target-tests-32/Disabled_Tests"
"${TEST_NAME}"
"${CMAKE_BINARY_DIR}/Bin/FEXLoader"
"--no-silent" "-c" "irjit" "-n" "500" "-R" $ENV{ROOTFS} "--"
"--no-silent" "-c" "irjit" "-n" "500" "--"
"${TEST}")
endforeach()
+1 -1
View File
@@ -18,7 +18,7 @@ foreach(TEST ${TESTS})
"${CMAKE_SOURCE_DIR}/unittests/gcc-target-tests-64/Disabled_Tests"
"${TEST_NAME}"
"${CMAKE_BINARY_DIR}/Bin/FEXLoader"
"--no-silent" "-c" "irjit" "-n" "500" "-R" $ENV{ROOTFS} "--"
"--no-silent" "-c" "irjit" "-n" "500" "--"
"${TEST}")
endforeach()
+1 -1
View File
@@ -18,7 +18,7 @@ foreach(TEST ${TESTS})
"${CMAKE_SOURCE_DIR}/unittests/gvisor-tests/Disabled_Tests"
"${TEST_NAME}"
"${CMAKE_BINARY_DIR}/Bin/FEXLoader"
"--no-silent" "-c" "irjit" "-n" "500" "-R" $ENV{ROOTFS} "--"
"--no-silent" "-c" "irjit" "-n" "500" "--"
"${TEST}")
endforeach()