Compare commits

..
433 Commits
Author SHA1 Message Date
Ryan Houdek a9d00b3f8d Docs: Update for release FEX-2207 2022-07-07 10:06:29 -07:00
Ryan Houdek fb41ba172d Merge pull request #1835 from wannacu/main
AOTIR: Fix IRList delete
2022-07-07 10:03:33 -07:00
Ryan Houdek aec5b21d2a Merge pull request #1833 from neobrain/feature_generic_callbacks
Thunks: Implement generic callback support
2022-07-07 10:02:02 -07:00
Ryan Houdek 124097d563 Merge pull request #1839 from neobrain/fix_value_dom_validation
ValueDominanceValidation: Avoid stack exhaustion when aggregating predecessors
2022-07-07 08:51:54 -07:00
Ryan Houdek e137c2edff Merge pull request #1837 from neobrain/fix_unknown_vulkan_functions
Vulkan: Handle queries for unknown functions more gracefully
2022-07-07 08:45:48 -07:00
Tony Wasserka 2b35dd829c Thunks: Mark inline assembly in CallHostThunkFromRuntimePointer as volatile 2022-07-07 17:33:31 +02:00
Tony Wasserka 686895802a Thunks/GL: Drop unused manual function implementations
These functions are annotated with fexgen::callback_stub nowadays, so
they don't need custom host endpoints anymore.
2022-07-07 17:33:31 +02:00
Tony Wasserka 72d0228cd7 Thunks/gen: Remove now unneeded callback_structs and callback_typedefs files 2022-07-07 17:33:31 +02:00
Tony Wasserka 298e6cad0c Thunks/Xext: Enable automatic handling of callbacks 2022-07-07 17:33:31 +02:00
Tony Wasserka ad34c228e3 unittests/ThunkLibs: Extend FunctionPointerParameter test 2022-07-07 17:33:31 +02:00
Tony Wasserka 66f74a431f Thunks/gen: Clean up implementation of generic callbacks 2022-07-07 17:33:31 +02:00
Tony Wasserka 57bae90c69 Thunks: Fix generic callbacks on ARM hosts 2022-07-07 17:20:28 +02:00
Tony Wasserka 83c2e52ba5 Vulkan: Handle queries for unknown functions more gracefully 2022-07-07 17:00:15 +02:00
Tony Wasserka 4f8525da4d ValueDominanceValidation: Avoid stack exhaustion when aggregating predecessors
The recursive algorithm used here previously led to deeply nested function
calls, which eventually exhausted the available stack space. The simple
non-recursive algorithm used now avoids this problem at the expense of
small overhead.
2022-07-07 16:46:38 +02:00
Ryan Houdek 3b8491b558 Merge pull request #1836 from neobrain/fix_duplicate_gch_links
Thunks: Soften error condition to be non-fatal
2022-07-07 03:00:41 -07:00
Tony Wasserka f69c53d294 Thunks: Soften error condition to be non-fatal
Dota Underlords hit this when querying Vulkan two different symbols
that resolve to the same function. Ignoring the error in that specific
case is safe, since the linked guest functions have the same implementation.
2022-07-07 11:33:01 +02:00
Ryan Houdek 6b226dd6af Merge pull request #1834 from Sonicadvance1/fix_steam_vulkan_thunks
Thunks: Adds libvulkan steam pinned library thunking support
2022-07-06 23:57:04 -07:00
wannacu 227462e4d8 AOTIR: Fix IRList delete 2022-07-07 14:39:00 +08:00
Ryan Houdek 5f00ba5cae Thunks: Adds libvulkan steam pinned library thunking support
Since we have switched over to thunking the vulkan loader, behaviour has
changed here and we need to thunk this library.

While a bit unsafe to thunk arbitrary libraries, we know this one is
safe to thunk.

Fixes thunking Vulkan on steam games which have been broken since
switching over to vulkan loader thunking.

In the future this may become unnecessary but it is required for now.
2022-07-06 20:45:14 -07:00
Ryan Houdek 0b8a6d9599 Thunks: Add support for a @HOME@ prefix
Currently thunk prefixes are mutally exclusive. No thunks database path
can currently have more than one prefix. If this is necessary then we
can add it in the future.

Most minor of optimizations here, we only scan the string once to find a
prefix, instead of searching for it on all four prefix replacements.
2022-07-06 20:42:00 -07:00
Ryan Houdek 13bf04a81e FEXCore: Expose the Paths namespace publically
We will need the ability to get the home directory from the frontend.
Get it the same way as FEXCore.
2022-07-06 20:40:49 -07:00
Tony Wasserka 7eb8409f71 Thunks: Move definitions for LoadLib and IsLibLoaded together 2022-07-06 18:41:53 +02:00
Tony Wasserka e821072cb9 unittests/ThunkLibs: Fix FunctionPointerParameter test 2022-07-06 18:41:53 +02:00
Stefanos Kornilios Misis Poiitidis 9c01dd9d9f Thunks: PoC Callbacks using sha256 exports from host 2022-07-06 18:41:53 +02:00
Ryan Houdek 6a43db8c8f Merge pull request #1832 from Sonicadvance1/fix_thunk_crash
Thunks: Fix std::set crash
2022-07-06 02:00:04 -07:00
Ryan Houdek 3ffc301dd0 Thunks: Fix std::set crash
std::string_view was sticking around for longer than libraries being
loaded.
This was causing a crash.
Change this to a std::string directly until we support cleanly removing
this data on library unload.
2022-07-05 21:25:40 -07:00
Ryan Houdek 46fcbe2fc0 Merge pull request #1830 from Sonicadvance1/support_erofs
Support EroFS
2022-07-05 07:21:34 -07:00
Ryan Houdek b7806e47e9 Merge pull request #1831 from Sonicadvance1/minor_fexserver_fixes_pt2
FEXServer: Stop leaking FDs to subprocesses
2022-07-05 06:16:34 -07:00
Ryan Houdek a1ed54adc7 FEXRootFSFetcher: Add support for EroFS
Fixes a bug in `ExecAndWaitForResponse` where results > 1024 bytes would
overwrite data.

Switches from a custom format txt file to a json file.
JSON file now has a "Type" field to specify squashfs versus erofs.

JSON is now versioned so we don't need to move the file around, just
append to a new versioned segment.

Only shows erofs files if you have the bleeding edge `erofsfuse`
application.
This application was available starting with erofs-utils v1.5 which was
released on 2022-06-13, so it isn't available pretty much everywhere.
2022-07-05 04:40:13 -07:00
Ryan Houdek 250a4ea4e3 FEXServer: Stop leaking FDs to subprocesses
Attach the CLOEXEC flag on each of the FDs/Sockets we open.
2022-07-05 03:53:39 -07:00
Stefanos Kornilios Mitsis Poiitidis b3e090c8ff Merge pull request #1826 from Sonicadvance1/fix_gdb_install_library
FEXGDBReader: Fix install path
2022-07-05 08:47:42 +00:00
Stefanos Kornilios Mitsis Poiitidis 982518d3a4 Merge pull request #1829 from Sonicadvance1/fix_ioctl32_vblank
Ioctl32: Fix DRM_IOCTL_WAIT_VBLANK
2022-07-05 08:45:46 +00:00
Ryan Houdek 9ad1d5548d FEXServer: Add support for EroFS 2022-07-04 20:51:59 -07:00
Ryan Houdek 0de2558ef7 FEXConfig: Add support for erofs 2022-07-04 20:51:34 -07:00
Ryan Houdek 3e0e601616 FileFormatCheck: Add support for EroFS 2022-07-04 20:51:13 -07:00
Ryan Houdek 8b202b0ebf Ioctl32: Fix DRM_IOCTL_WAIT_VBLANK
Oops. Used the incorrect ioctl for this one. Copy and paste fail.
2022-07-04 18:14:07 -07:00
Ryan Houdek 6d2f98a379 Merge pull request #1822 from Sonicadvance1/check_binfmt_misc_conflict
CMake: Check for binfmt_misc conflicts before install
2022-07-02 01:38:37 -07:00
Ryan Houdek 84379b5fdf CMake: Check for binfmt_misc conflicts before install
Check for qemu and box binfmt_misc file conflicts before the
`binfmt_misc` install command.

This ensures if you're building from source that you won't inadvertently
install conflicting binfmt_misc files, breaking program execution.
2022-07-01 13:42:45 -07:00
Ryan Houdek d005fdcd03 Merge pull request #1823 from Sonicadvance1/classify_arm
unittests: Classify CPU based on CPU features
2022-06-30 23:34:36 -07:00
Ryan Houdek a97fb2f34f Merge pull request #1825 from Sonicadvance1/disable_posix_flake
unittests: Disable known flake in posix tests
2022-06-30 23:34:22 -07:00
Ryan Houdek d48981b6b0 FEXGDBReader: Fix install path 2022-06-30 22:34:25 -07:00
Ryan Houdek af32228e38 unittests: Disable known flake in posix tests
Interpreter is flakey here, likely due to some race, but is periodically
fails and makes CI red.
2022-06-30 14:41:14 -07:00
Ryan Houdek 8b35275ec1 unittests: Classify CPU based on CPU features
Instead of relying on runner features, classify based on CPU features.

This fixes an annoying issue where if running unit tests locally without
it set then you get an unexpected failure.

Fixes #1807
2022-06-30 13:55:38 -07:00
Ryan Houdek 302a6c96ff Merge pull request #1818 from FEX-Emu/skmp/fix-guest-h
ThunkLibs: Fix Guest.h
2022-06-30 10:02:48 -07:00
Stefanos Kornilios Misis Poiitidis f4a4b2c14b ThunkLibs: Fix Guest.h 2022-06-30 14:08:01 +03:00
Stefanos Kornilios Mitsis Poiitidis 88b94bef54 Merge pull request #1812 from FEX-Emu/skmp/add-thunks-islibloaded
Thunks: Add fex:is_lib_loaded
2022-06-30 04:45:37 +00:00
Stefanos Kornilios Mitsis Poiitidis 751b66d45d Merge pull request #1816 from Sonicadvance1/fix_vulkan_debug_report
Thunks/vulkan: Disable debug report callback
2022-06-30 04:43:43 +00:00
Stefanos Kornilios Mitsis Poiitidis ae6a57e667 Merge pull request #1815 from Sonicadvance1/fix_wine_preloader
Config: Fixes AppConfig for wine-preloader
2022-06-30 04:43:37 +00:00
Stefanos Kornilios Mitsis Poiitidis 9110546d34 Merge pull request #1814 from Sonicadvance1/pressure_vessel_fexserver_fix
FEXServerClient: When running under pressure-vessel don't use FEXServer rootfs
2022-06-30 04:43:26 +00:00
Stefanos Kornilios Mitsis Poiitidis 1f1d0706dd Merge pull request #1813 from Sonicadvance1/fexserver_changes
FEXServer: Minor changes
2022-06-30 04:43:16 +00:00
Ryan Houdek a82d41ee22 Thunks/vulkan: Disable debug report callback
This was checking for the wrong debug structure and it wasn't actually
unlinking it correctly from the linked list.
Since it was talking directly to a_0 instead of the current modifying
struct.

Instead of having a stubbed debug report struct, we can just remove the
structure from the linked list. Since we are already abusing a const
cast there anyway.

Fixes Vulkan thunks on Snapdragon.
2022-06-29 20:02:52 -07:00
Ryan Houdek e8e70828d1 Config: Fixes AppConfig for wine-preloader
When wine-preloader is executed it doesn't do an execve to passed in
wine program. It will instead map the executable directly in to memory
and start executing it.

This way we end up with a program executing like `wine-preloader
<absolute wine path> Game.exe`

This now handles the wine-preloader case so we can get the correct
application profile here.
2022-06-29 20:00:32 -07:00
Ryan Houdek b7a58fdb8a FEXServerClient: When running under pressure-vessel don't use FEXServer rootfs
pressure-vessel overrides our rootfs when it does a pivot_root.
Since we are still communicating to the FEXServer we were pulling the
configured rootfs.

Instead check if we are in pressure vessel and avoid doing that.

This fixes FEX running under pressure-vessel.

(There may be some implications to this down the road with code caching
but let's worry about that later)
2022-06-29 19:57:07 -07:00
Ryan Houdek dc9fde8fae FEXServer: Change server lock fifo to regular file
This doesn't need to be a FIFO.
Resolves an issue of trying to run a rootfs from a fex config folder mapped over nfs/sshfs.
2022-06-29 13:46:49 -07:00
Ryan Houdek c0cf4f6a61 FEXServer: Change socket pathname to include euid
Just to ensure the socket path is unique per user.

Noticed this while running independent FEXServers with multiple users.
2022-06-29 13:44:48 -07:00
Stefanos Kornilios Misis Poiitidis 21fc6bedcb Thunks: Add fex:is_lib_loaded 2022-06-29 19:20:08 +03:00
Ryan Houdek ad6fd5ab72 Merge pull request #1804 from neobrain/fix_cmake_thunks_portability
Allow building thunks on a wider range of platforms
2022-06-29 02:08:55 -07:00
Ryan Houdek 0f696c6092 Merge pull request #1811 from FEX-Emu/skmp/fix-vixl-assert
Dispatcher/Arm64: Fix vixl assert
2022-06-29 02:08:08 -07:00
Stefanos Kornilios Misis Poiitidis 9af3bc1864 Dispatcher/Arm64: Fix vixl assert 2022-06-29 10:41:24 +03:00
Ryan Houdek b020e593a5 Merge pull request #1802 from Sonicadvance1/fexserver_wait
FEXServer: Adds -w option for waiting on current FEXServer
2022-06-28 09:33:31 -07:00
Tony Wasserka 4771a340f5 Merge pull request #1803 from neobrain/fix_vulkan_debug_report
ThunkLibs/vulkan: Work around lack of generic callback support in VK_EXT_debug_report
2022-06-28 16:39:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 4449b60459 Merge pull request #1787 from FEX-Emu/skmp/gdb-jit-integration
gdb: jit integration
2022-06-27 11:57:40 +00:00
Stefanos Kornilios Misis Poiitidis 4139332ad9 GDBSymbols: Cleanups 2022-06-27 14:44:07 +03:00
Stefanos Kornilios Mitsis Poiitidis 30a28ff1ad Merge pull request #1801 from FEX-Emu/skmp/remove-thunk-warnings
ThunkLibs: silence warnings
2022-06-27 10:35:43 +00:00
Tony Wasserka 498d0fc145 CMake: Use toolchain files to set up x86 cross compilation 2022-06-25 14:00:36 +02:00
Tony Wasserka 4a547fe95f Thunks/CMake: Link against clang-cpp instead of clangTooling 2022-06-25 13:54:12 +02:00
Tony Wasserka ca906589d4 Thunks/CMake: Automatically discover the clang resource directory 2022-06-25 13:53:48 +02:00
Tony Wasserka 8bafae2262 ThunkLibs/vulkan: Work around lack of generic callback support in VK_EXT_debug_report 2022-06-25 12:44:58 +02:00
Tony Wasserka 4ef82c82b5 Revert "Thunks/vulkan: Disable support for debug extensions as they require callback support"
This reverts commit c04d2409da.
2022-06-25 11:48:39 +02:00
Ryan Houdek e03d253310 FEXServer: Adds -w option for waiting on current FEXServer
It can be useful to know in tooling when the current active FEXServer
has exited.

Two things can happen when this command is run.
No FEXServer is active, returns immediately.
A FEXServer is active, we query for a pidfd from the active server, then
we wait until it exits.

Both instances of this is valid to use.
2022-06-24 22:56:13 -07:00
Stefanos Kornilios Misis Poiitidis 48c45da2de ThunkLibs: silence warnings 2022-06-25 08:27:57 +03:00
Ryan Houdek 04a1ac967c Merge pull request #1760 from neobrain/feature_guest_callable_hostptrs
Thunks: Support returning host function pointers to the guest
2022-06-24 18:22:24 -07:00
Ryan Houdek aa17f64593 Merge pull request #1799 from FEX-Emu/skmp/fix-get_fdpath
FDUtils: Fix get_fdpath
2022-06-24 16:42:30 -07:00
Stefanos Kornilios Misis Poiitidis ae5cfcc249 Symbols: Add SymName function 2022-06-24 19:38:40 +03:00
Stefanos Kornilios Misis Poiitidis 8f578b57f2 GDBSymbols: Cleanups 2022-06-24 17:02:39 +03:00
Stefanos Kornilios Misis Poiitidis 3871646611 GDBSymbols: Add gdb reader for fex, GDBSymbols option to enable 2022-06-24 16:27:16 +03:00
Stefanos Kornilios Mitsis Poiitidis 48a574dfd3 FDUtils: Fix get_fdpath 2022-06-24 16:11:10 +03:00
Tony Wasserka ec49100c63 Thunks/CMake: Remove now unneeded helper functionality
This was needed for libvulkan_device. With libvulkan thunked directly now,
there is no further use of this code.
2022-06-24 11:56:36 +02:00
Tony Wasserka c04d2409da Thunks/vulkan: Disable support for debug extensions as they require callback support 2022-06-24 11:56:36 +02:00
Tony Wasserka b93b713179 Thunks/vulkan: Thunk libvulkan directly instead of libvulkan_device 2022-06-24 11:56:36 +02:00
Tony Wasserka fe2f54fc3d Thunks/GL: Use guest-callable host function pointers to implement glXGetProcAddress 2022-06-24 11:56:36 +02:00
Tony Wasserka 599f8f99ed Thunks/gen: Add support for thunking APIs that return host function pointers
Interface definitions must enable this functionality by enclosing functions
that may be called through function pointers in a namespace annotated with
`fexgen::indirect_guest_calls`. The guest thunk must further link any host
function pointers to a guest-side instance of CallHostThunkFromRuntimePointer.
2022-06-24 11:56:36 +02:00
Tony Wasserka b1b338ed5a Thunks: Simplify guest helper macros 2022-06-24 11:56:36 +02:00
Ryan Houdek e4d659a619 Merge pull request #1797 from FEX-Emu/skmp/remove-used-irs
IR: Remove GuestCallDirect, GuestCallIndirect
2022-06-23 10:45:01 -07:00
Stefanos Kornilios Misis Poiitidis 9ce94266e1 IR: Remove GuestCallDirect, GuestCallIndirect 2022-06-23 19:46:25 +03:00
Ryan Houdek 5a19425b28 Merge pull request #1792 from Sonicadvance1/fexserver
FEXServer: Adds new FEXServer service
2022-06-23 09:25:29 -07:00
Ryan Houdek 1494aac861 Merge pull request #1796 from Sonicadvance1/fix_clone3_stack_again
Linux: Fixes for clone3 stack size
2022-06-23 09:21:39 -07:00
Ryan Houdek 8a21ecabee Merge pull request #1795 from neobrain/fix_radata_asan
Fix inconsistent allocation schemes used for RegisterAllocationData
2022-06-23 09:03:05 -07:00
Ryan Houdek 005b2bc3db Linux: Fixes for clone3 stack size
We weren't adjusting the guest stack size when using clone3.
If the clone comes from CLONE3 then we need to offset the RSP by the
provided stack size.

This also translates to fork/vfork through clone, so make sure to adjust
stack in that case as well.

Fixes Ender Lilies crashing with Ubuntu 22.04 rootfs
2022-06-23 08:41:15 -07:00
Tony Wasserka f27830bf41 Fix inconsistent allocation schemes used for RegisterAllocationData 2022-06-23 17:16:44 +02:00
Ryan Houdek 5fe6afc270 FEXServer: Adds new FEXServer service
This is a relatively invasive change since multiple things needed to
happen at once.

* Socket based logging is removed
  * Logging has been replaced to only support stdout, stderr, and server
  * Server is now default and replaces what FEXLogServer did
  * Server logging now uses a pipe instead of a socket
  * Can be faster than stderr and stdout since the application doesn't
    need to wait on terminal output

* FEXMountDaemon has been removed
  * Functionality has been merged in to FEXServer

* FEXServer is always executed on FEX initialization time
  * Similar in behaviour to Wine's wineserver
  * Can explicitly start this before using FEX for logging purposes
  * Stays around until all instances of FEX exit
  * Will stick around for a short amount of time in case of spurious
    execution

* FEXServer will soon be extended to do more than logging and squashfs
  mounting

* FEX rootfs scripts will need to be updated to support this path
  * Just means rbinding the /tmp folder and forcing a FEXServer instance
    to be alive
* Pressure-vessel works fine in this case since FEXServer will already
  be running
  * It already rbinds the host /tmp folder which is why this works
2022-06-23 07:51:56 -07:00
Stefanos Kornilios Mitsis Poiitidis 4f9bc70562 Merge pull request #1794 from FEX-Emu/skmp/fix-unidispatch-multithread
Dispatchers: Use thread local emitters for backend callbacks
2022-06-23 10:45:12 +00:00
Stefanos Kornilios Misis Poiitidis 835155021d Dispatchers: Use thread local emitters for backend callbacks 2022-06-23 13:33:56 +03:00
Stefanos Kornilios Misis Poiitidis e2c5c2292d Dispatchers: Use thread local emitters for backend callbacks 2022-06-23 13:20:03 +03:00
Stefanos Kornilios Mitsis Poiitidis 072690a191 Merge pull request #1782 from FEX-Emu/skmp/ipr-unified-dispatch
Backends: Unified dispatch, interface rework, cleanups
2022-06-23 07:03:14 +00:00
Ryan Houdek a2c9d5a398 Merge pull request #1790 from neobrain/fix_missing_packfn_define
unittests/ThunkLibs: Fix test failures due to missing FEX_PACKFN_LINKAGE define
2022-06-22 19:08:55 -07:00
Ryan Houdek 5944342604 Externals: update cpp-optparse 2022-06-22 17:39:59 -07:00
Stefanos Kornilios Misis Poiitidis 4dbcd34548 Backends: Rework some of the interfaces, unified backend dispatch 2022-06-22 22:24:42 +03:00
Tony Wasserka f02c73a2c2 unittests/ThunkLibs: Fix test failures due to missing FEX_PACKFN_LINKAGE define 2022-06-22 17:03:01 +02:00
Ryan Houdek 158ba1ae3b Merge pull request #1788 from Sonicadvance1/safe_v3d_csd
Ioctl: Safely access v3d csd ioctl structure
2022-06-21 15:33:50 -07:00
Ryan Houdek a85a77e088 Ioctl: Safely access v3d csd ioctl structure
DRM ioctls are taking advantage of the fact that any ioctl structure
that is passed in to the kernel smaller than the expected value will
zero out the remaining members of the structure.

Some more ioctls coming down the pipe will also abuse this, so we might
as well as get started with the v3d ioctl that requires it.

Due to how this works, the kernel knows how big an ioctl structure
should be and if the userspace passes in a smaller struct, it will
memset the remaining size to zero.
2022-06-19 09:53:50 -07:00
Ryan Houdek 542ab04671 Merge pull request #1786 from Sonicadvance1/optimize_fd_path
Linux: Make `get_fdpath` more optimal
2022-06-18 01:57:55 -07:00
Ryan Houdek 28ee2ca5a2 Linux: Make get_fdpath more optimal
std::filesystem::canonical is very heavyweight and walks the full path
to ensure that each folder in the path is not a symlink.

eg:
```
readlink("/proc", 0x7ffd5646e210, 1023)                                         = -1 EINVAL (Invalid argument)
readlink("/proc/self", "880556", 1023)                                          = 6
readlink("/proc/880556", 0x7ffd5646e210, 1023)                                  = -1 EINVAL (Invalid argument)
readlink("/proc/880556/fd", 0x7ffd5646e210, 1023)                               = -1 EINVAL (Invalid argument)
readlink("/proc/880556/fd/5", "/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-6"..., 1023) = 86
readlink("/home", 0x7ffd5646e210, 1023)                                         = -1 EINVAL (Invalid argument)
readlink("/home/ryanh", 0x7ffd5646e210, 1023)                                   = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu", 0x7ffd5646e210, 1023)                          = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS", 0x7ffd5646e210, 1023)                   = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04", 0x7ffd5646e210, 1023)      = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr", 0x7ffd5646e210, 1023)  = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
```

This is what was occuring for every single mmap that occurs. /really/
adding to the time for the syscall to take.
This also happens on a couple of other syscalls which are using this new
path now.

The primary reason why this works is that we know that every entry in
`/proc/self/fd/` is a symlink. So instead of asking for canonical, we
can just read the symlink and this will redirect us to the canonical
path.
So this previous example goes from 14 syscalls down to 1.

eg:
```
readlinkat(AT_FDCWD, "/proc/self/fd/5", "/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-6"..., 4096) = 86
```

While this is only a minor improvement in the "typical" operating environment,
this significantly improves performance of FEX under proot or if the
rootfs lives on a network share.
2022-06-18 01:45:37 -07:00
Stefanos Kornilios Mitsis Poiitidis 3913dd6c8b Merge pull request #1785 from Sonicadvance1/wine_app_config
Common: Support application profiles for games launched through wine
2022-06-18 07:42:16 +00:00
Ryan Houdek 1ecf147e3e Common: Support application profiles for games launched through wine
Wine will set the application name later in the boot process but we
can't defer application loading that late.

Once an application is loaded with wine or wine64, then check the next
argument for the application name instead.

This will allow us to have wine application application profiles.

eg: FEXInterpreter `which wine` $HOME/.wine/drive_c/GOG\ Games/Oblivion/Oblivion.exe
This will give us the application name of `Oblivion.exe`

Same with: FEXInterpreter `which wine` C:\\GOG\ Games\\Oblivion\\Oblivion.exe
2022-06-17 23:24:42 -07:00
Ryan Houdek d8fa53a445 Merge pull request #1783 from FEX-Emu/skmp/fix-thunksdb-prefixing
ThunksDB: Fix String.find error check
2022-06-17 15:32:08 -07:00
Ryan Houdek 1c1ad876ca Merge pull request #1784 from lioncash/regsize
CoreState: Add register size constants
2022-06-17 15:31:59 -07:00
lioncash 58ae49372c CoreState: Add register size constants
Gets rid of some magic numbers and reduces the number of things that
need to manually change (e.g. when supporting AVX and needing to
increase the xmm size).
2022-06-17 10:55:07 -04:00
Stefanos Kornilios Misis Poiitidis 917e69b021 ThunksDB: Fix String.find error check 2022-06-17 15:04:17 +03:00
Mai 9242e59841 Merge pull request #1781 from Sonicadvance1/fix_free
AOTIR: Fix RAData free
2022-06-16 20:11:55 -04:00
Ryan Houdek 05d15fa052 AOTIR: Fix RAData free
This thing is allocated with malloc so it needs to be free'd with free.
This was poisoning asan runs.
2022-06-16 17:01:06 -07:00
Stefanos Kornilios Mitsis Poiitidis 0a62a4c571 Merge pull request #1775 from FEX-Emu/skmp/dispatcher-per-context
Make Dispatcher per Context from per Thread, Simplify TestHarnessRunner
2022-06-16 21:08:32 +00:00
Stefanos Kornilios Misis Poiitidis 525266e482 Core: Move common signal handling setup to context 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 503881f5df Jit/Arm64: Fix ubuntu 20.04 include 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 04fd4f2bb4 Arm64Dispatcher: Add missing ExitFunctionLinker Pointer 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 8be1e5e260 TestHarnessRunner: Fix formatting 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis e84e084d56 FEXLogServer: Work around linking issues 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 12ad05e089 HostRunner: Fix typo for arm64 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 472ce7c1e4 Context: Add DipsatcherConfig to class, minor header massaging 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 19fbc10c16 TestHarnessRunner: Don't setup a FEX Context when running host tests 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 39f3406fba Backends+Context: Make Dispatcher per Context from per Thread 2022-06-16 19:41:42 +00:00
Ryan Houdek 19b0a9cd7d Merge pull request #1779 from lioncash/irname
Arm64/JIT: Use IR names in opcode implementations
2022-06-16 12:37:18 -07:00
Stefanos Kornilios Mitsis Poiitidis eac579f714 Merge pull request #1778 from FEX-Emu/skmp/fix-thread-creation-race
Context: Fix CreateThread partial initialization issue
2022-06-16 18:17:34 +00:00
Stefanos Kornilios Mitsis Poiitidis b0a31f7105 Merge pull request #1770 from FEX-Emu/skmp/custom-ir-handlers-2
Context: Decouple from CodeLoader, introduce generic CustomIREntrypoints
2022-06-16 18:04:18 +00:00
Stefanos Kornilios Misis Poiitidis 29163cfd07 Context: Fix CreateThread partial initialization issue 2022-06-16 17:51:49 +00:00
Stefanos Kornilios Mitsis Poiitidis f05b24636f OpDisp: Re add ShouldDump 2022-06-16 17:38:44 +00:00
lioncash 39466a3a8f Arm64/VectorOps: Move arguments over to IR names 2022-06-16 12:03:33 -04:00
lioncash bccf18f46c Arm64/MoveOps: Move arguments over to IR names 2022-06-16 12:03:26 -04:00
lioncash 1109126e35 Arm64/MiscOps: Move arguments over to IR names 2022-06-16 12:03:18 -04:00
lioncash 98dd40c3a8 Arm64/MemoryOps: Move arguments over to IR names 2022-06-16 12:03:12 -04:00
lioncash b1cfb104ff Arm64/FlagOps: Move arguments over to IR names 2022-06-16 12:03:02 -04:00
lioncash c339cfe184 Arm64/EncryptionOps: Move arguments over to IR names 2022-06-16 12:02:53 -04:00
lioncash ca8e6a83e1 Arm64/ConversionOps: Move arguments over to IR names 2022-06-16 12:02:45 -04:00
lioncash 760217ab29 Arm64/BranchOps: Move arguments over to IR names 2022-06-16 12:02:38 -04:00
lioncash ee9361bf41 Arm64/ALUOps: Move arguments over to IR names
Makes it a little easier to read.
2022-06-16 12:02:28 -04:00
Stefanos Kornilios Mitsis Poiitidis e62bc24b3f Merge pull request #1777 from FEX-Emu/skmp/create-directories-during-configuration
CMAKE: Create directories during configuration, fixes endless generation of unittests
2022-06-15 01:58:57 +03:00
Stefanos Kornilios Misis Poiitidis 18074307f6 Core: Add Creator, Data to CustomIRHandlers, return them + lock on Add 2022-06-15 01:48:32 +03:00
Ryan Houdek 63b70ff3d4 Merge pull request #1776 from lioncash/vixl-pcl
Arm64/EncryptionOps: Fix register specifiers in PCLMUL movs
2022-06-14 15:46:40 -07:00
Stefanos Kornilios Misis Poiitidis dacdfd5c02 CMAKE: Create directories during configuration, fixes endless generation of unittests 2022-06-15 01:10:33 +03:00
lioncash d5c039cdc7 Arm64/EncryptionOps: Fix register specifiers in PCLMUL movs
Prevents an assertion from firing in vixl, since the destination needs
to be a scalar
2022-06-14 10:15:27 -04:00
Ryan Houdek e2e6f2a92b Merge pull request #1767 from Sonicadvance1/pressure_vessel_thunks
Thunks: Support pressure-vessel prefixes
2022-06-13 10:26:00 -07:00
Ryan Houdek d7d8244593 Thunks: Support pressure-vessel prefixes
Instead of duplicating prefixes in the ThunksDB json file even more,
just do a prefix replacement when the thunks database is being parsed.

Changes the JSON overlay arrays over to @PREFIX@ and reduce the
duplication.

Adds support for the pressure-vessel prefix `/usr/lib/pressure-vessel/overrides/lib`

This gets thunks ready for running under pressure-vessel, once vulkan
thunking switches from device libraries to the vulkan loader it should
just work.
2022-06-13 09:48:17 -07:00
Ryan Houdek b05e5ce14e Merge pull request #1773 from lioncash/header
JITs: Qualify external includes consistently
2022-06-13 09:46:54 -07:00
lioncash 08730c817a X86Dispatcher: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:55:44 -04:00
lioncash 34e086eec2 Arm64Dispatcher: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:54:49 -04:00
lioncash 83c8a9b675 ArchHelpers/Arm64: Qualify external includes
Qualifies external includes with <> instead of quotes.
2022-06-13 10:52:39 -04:00
lioncash a3db629fd6 Arm64Emitter: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:50:03 -04:00
lioncash 64b3cef126 ARM64/JITClass: include dependencies
Includes dependencies directly used by the class. Also qualifies the
vixl includes with <> instead of quotes.
2022-06-13 10:43:13 -04:00
Stefanos Kornilios Misis Poiitidis 7100a2eaee Context: Decouple from CodeLoader, introduce generic CustomIRHandlers 2022-06-13 11:07:16 +03:00
Stefanos Kornilios Mitsis Poiitidis ffcde1823b Merge pull request #1769 from FEX-Emu/skmp/interlocked-invalidate-2
Invalidations: Move invalidation locks to Context
2022-06-13 11:06:34 +03:00
Stefanos Kornilios Mitsis Poiitidis 4c73c715ad Merge pull request #1771 from Sonicadvance1/fix_32bit_allocator
Linux: Fixes 32-bit allocator range scanning
2022-06-13 11:06:14 +03:00
Ryan Houdek 26c3ebd6a7 Linux: Fixes 32-bit allocator range scanning
32-bit allocations were scanning for available pages by shifting by the
length of allocation. This is incorrect since large allocations would
then only scan through the region in very large chunks. Especially if
MAP_32BIT was present. Instead scan by page size to ensure better fit.

This fixes X-Plane 11.

Also ensure we don't try to munmap a range unless the result is
MAP_FAILED, noticed we were trying to munmap ~0ULL
2022-06-12 17:53:12 -07:00
Ryan Houdek 790bd9747f Merge pull request #1756 from FEX-Emu/skmp/ipr-own-irlists
IPR: Store copy of IRLists, Dispatcher cleanups
2022-06-11 12:15:18 -07:00
Stefanos Kornilios Misis Poiitidis 83aa8731d1 Invalidations: Move invalidation locks to Context, extend InvalidateGuestCodeRange with optional callback 2022-06-11 17:15:41 +03:00
Stefanos Kornilios Mitsis Poiitidis bbd9eb5b9a Merge pull request #1766 from FEX-Emu/skmp/fix-smc-mt-2
SMC: Track code pages before frontend decode
2022-06-11 16:05:02 +03:00
Stefanos Kornilios Misis Poiitidis cc90fa5773 LookupCache: Review feedback & cleanups 2022-06-11 13:17:49 +03:00
Stefanos Kornilios Mitsis PoiitidisandTony Wasserka 30ebc6c938 Update External/FEXCore/Source/Interface/Core/LookupCache.h
Co-authored-by: Tony Wasserka <4840017+neobrain@users.noreply.github.com>
2022-06-11 12:52:37 +03:00
Ryan Houdek c0a8984799 Merge pull request #1764 from Sonicadvance1/fix_xxhash
FEXRootFSFetcher: Update and fix xxhash file hashing
2022-06-10 04:57:09 -07:00
Stefanos Kornilios Misis Poiitidis a6e34d301e Backends: Move and use INITIAL_CODE_SIZE, MAX_CODE_SIZE consts 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 82c88168cc JITPointers: Should be a struct now 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis d169c3ed4a Context: Remove unused param from ClearCodeCache 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 563d11a702 IPR: Actually works now 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 31e2ba4696 IPR: Fix the build 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 2c3baaad1c Context: Split LocalIR to PrecompiledIR and DebugStore 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 01bd43aa2b IPR: Store IR in the code bufffer, use executer function, cleanups 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis f4d095ce0f Jit/dispatch: ExecuteBlocksWithCall -> ExecuteBlocksWithCall 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 8c5bd9a058 IPR: Move dispatcher creation to InterpreterCore ctor 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 335ce1d056 JIT: Merge common CPUFrame::Pointers 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 0afb3caaed Refactors: Move CodeBuffers to CPUBackend, SignalHandlerRefCounter to CpuStateFrame, Arm64 Relocation to ARM64Jit 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 8521aaf65c Core, Frontend, LookupCache: Track code pages before frontend decode 2022-06-10 10:13:47 +03:00
Stefanos Kornilios Misis Poiitidis 0a451cfe0d FEXLT: Run 100 iterations of smc-mt-2 to catch races 2022-06-10 10:01:37 +03:00
Stefanos Kornilios Mitsis Poiitidis cb0935c96b Merge pull request #1737 from FEX-Emu/skmp/fex-linux-tests
unittests: Add FEXLinuxTests with a few tests
2022-06-10 08:51:15 +03:00
Stefanos Kornilios Misis Poiitidis 40ec910108 unittests: Add FEXLinuxTests, with a few signal and smc tests 2022-06-10 08:34:40 +03:00
Stefanos Kornilios Mitsis Poiitidis 2d3c6efae2 Merge pull request #1761 from FEX-Emu/skmp/ir-meta-fns
IR: add IsFragmentExit, IsBlockExit
2022-06-10 08:15:03 +03:00
Stefanos Kornilios Mitsis Poiitidis c027acecf8 Merge pull request #1763 from Sonicadvance1/ci_auto_rootfs
CI: Auto rootfs fetching
2022-06-10 07:34:56 +03:00
Ryan Houdek bc2840e4a7 FEXRootFSFetcher: Update and fix xxhash file hashing
The final tail of the file reading was incorrect, so our hashing was
"correct" but it was using stale data from the previous block size read.

Noticed this while wiring up the CI rootfs fetching since the hashing is
a lot simpler there.

Now instead of reading a tail, just attempt to read the full block size
and use the resulting data size instead. Confirmed it matches expected
results now.

In the process we are going to need to update hyperlinks and hashes
anyway, change the hash to XXH3 so it is faster to run.
2022-06-09 20:20:43 -07:00
Ryan Houdek 869b472d91 xxhash: Update xxhash to 0.8.1 2022-06-09 20:09:49 -07:00
Ryan Houdek 81dfc21700 unittests/gvisor-test: Disable semaphore test for now
New rootfs images cause different data to be returned in permissions
than this test was expecting. Disable this while we are upgrading CI.
2022-06-09 20:08:24 -07:00
Ryan Houdek a3d8fe2362 github: Allow CI to fetch its own rootfs
Just need to set some environment variables and execute our new script
2022-06-09 20:08:24 -07:00
Ryan Houdek d2a57f6231 Scripts: Add CI rootfs fetch script
This will allow our CI system to automatically pull their rootfs.
2022-06-09 20:08:24 -07:00
Ryan Houdek 9e9ceb3894 Merge pull request #1762 from lioncash/pcl
OpcodeDispatcher: Handle CLMUL opcode extension
2022-06-09 12:55:54 -07:00
Stefanos Kornilios Misis Poiitidis bd70088877 IR: add IsFragmentExit, IsBlockExit 2022-06-09 20:44:24 +03:00
lioncash 63c9d99b33 CPUID: Enable PCLMULQDQ CPUID bit
Now that we handle all PCLMULQDQ facilities, we can enable this bit.
2022-06-09 12:42:34 -04:00
lioncash a2469f48e7 OpcodeDispatcher: Handle VPCLMULQDQ 2022-06-09 12:42:34 -04:00
lioncash 62f0f421a9 OpcodeDispatcher: Handle PCLMULQDQ 2022-06-09 11:22:32 -04:00
lioncash 0b24758f41 IR: Add PCLMUL IR opcode 2022-06-09 11:22:29 -04:00
lioncash 147897e13d External: Update vixl submodule
Allows use of the 128-bit variant of PMULL/PMULL2
2022-06-09 11:06:47 -04:00
Ryan Houdek ade3a5275d Merge pull request #1758 from lioncash/warn
Tests/IRLoader: Silence missing override warning
2022-06-08 12:35:29 -07:00
Ryan Houdek 51c5f945b4 Merge pull request #1759 from lioncash/ir
IR.json: Correct 'Dest' key to 'Desc'
2022-06-08 12:35:04 -07:00
lioncash a4e2e36243 IR.json: Correct 'Dest' key to 'Desc'
In a few places the description key was accidentally written as Dest.
2022-06-08 10:58:49 -04:00
lioncash 2c5dc13201 Tests/IRLoader: Silence missing override warning 2022-06-08 10:41:28 -04:00
Ryan Houdek 43234939ca Merge pull request #1757 from Sonicadvance1/fix_open_wrapping
Linux: Fixes `open` syscall emulated path handling
2022-06-08 03:52:16 -07:00
Ryan Houdek 87b3d50899 Linux: Fixes open syscall emulated path handling
open is rarely used compared to openat and openat2, so this has been
just missed but nothing really got upset about it.

Fixes Proton Experimental while running under pressure-vessel.
2022-06-07 17:27:25 -07:00
Ryan Houdek da8dbf1777 Merge pull request #1755 from Sonicadvance1/hypervisor_bit
CPUID: Enable the hypervisor bit
2022-06-06 11:52:21 -07:00
Ryan Houdek 34e1fcccf8 CPUID: Enable the hypervisor bit
This was originally set to zero out of concern for any application that
is doing anti-emulation, anti-cheat, anti-VM checks.

This concern is likely unwarranted, and if any application/game starts
hitting this as a problem then we can throw an application profile at it
instead.

This makes pressure-vessel emulation checking more optimal by it not
having to do a uname dance.
2022-06-06 11:36:42 -07:00
Stefanos Kornilios Mitsis Poiitidis c99d1e48bd Merge pull request #1752 from FEX-Emu/skmp/tso-auto-migration
TSO: Add auto migration optimisation for applications that don't need TSO
2022-06-06 19:42:39 +03:00
Stefanos Kornilios Misis Poiitidis 6a428043f6 TSO: Add auto migration optimisation for applications that don't need TSO 2022-06-06 16:51:32 +03:00
Stefanos Kornilios Mitsis Poiitidis aafe7ff10f Merge pull request #1751 from Sonicadvance1/allow_override
Scripts: Allow user override on tagged version
2022-06-05 12:32:03 +03:00
Ryan Houdek 952e157770 Scripts: Allow user override on tagged version
In the case of a missing month or a minor version, need to allow user
defined overrides.
2022-06-04 18:06:28 -07:00
Ryan Houdek cae4f2f873 Docs: Update for release FEX-2206 2022-06-04 12:55:42 -07:00
Stefanos Kornilios Mitsis Poiitidis 0fc6d6b6b5 Merge pull request #1749 from Sonicadvance1/atomic_tests
unittests: Reenable atomic tests on ARMv8.0
2022-06-04 14:02:25 +03:00
Stefanos Kornilios Mitsis Poiitidis 7227ee9b2e Merge pull request #1748 from Sonicadvance1/gvisor_investigations
unittests: Investigate failing CI changes
2022-06-04 13:49:48 +03:00
Stefanos Kornilios Mitsis Poiitidis c6153d6a52 Merge pull request #1747 from Sonicadvance1/struct_verifier_fixes
Struct verifier fixes and reenable
2022-06-04 13:46:30 +03:00
Ryan Houdek d82d2944a9 Merge pull request #1745 from FEX-Emu/skmp/mtrack-fixes
mtrack: Fixes 32-bit shmat, shmdt tracking, guaranteed invalidation atomicity
2022-06-04 00:36:58 -07:00
Ryan Houdek 58ad400519 unittests: Reenable atomic tests on ARMv8.0 2022-06-04 00:30:02 -07:00
Ryan Houdek d23c76d0a9 Arm64: Work with more unaligned atomic operations
The latest ARMv8.0 toolchain is implementing fetch_add with a bic rather
than an and. Not sure why they started doing this but support the
remaining logical operations in our unaligned atomics handler.

Fixes the Interpreter ARMv8.0 atomic ops.

Fixes #1742
2022-06-04 00:30:02 -07:00
Ryan Houdek 6a5b9e2a93 unittests: Investigate failing CI changes
One gvisor test didn't expect a file header to change layout.
Another one was testing behaviour that was removed from upstream Linux

Fixes #1741
2022-06-03 23:03:59 -07:00
Ryan Houdek f33a93b0a1 github: Reenable struct verifier tests 2022-06-03 19:15:32 -07:00
Ryan Houdek 015200f511 StructVerifier: Reenable DRM testing 2022-06-03 19:12:33 -07:00
Ryan Houdek 92f48819b6 IoctlEmulation: Fix DRM includes
These were being overridden by system includes.
2022-06-03 19:12:33 -07:00
Ryan Houdek 0c6483cad5 External: Update drm-headers 2022-06-03 18:50:33 -07:00
Stefanos Kornilios Misis Poiitidis ee02b1ca51 Review feedback 2022-06-03 14:15:22 +03:00
Stefanos Kornilios Misis Poiitidis 33845a3112 Mtrack: Remove race conditions around concurrent invalidation and compilation 2022-06-03 12:41:38 +03:00
Stefanos Kornilios Misis Poiitidis 096ed29b5e Mtrack/x86: Track shmat, shmdt via ipc syscall as well 2022-06-03 12:41:32 +03:00
Ryan Houdek 3bbff8a948 Merge pull request #1744 from lioncash/sha256
OpcodeDispatcher: Implement SHA256 instructions
2022-06-02 13:53:52 -07:00
lioncash 726918b82c CPUID: Enable SHA extension bit
Now that all SHA instructions have an implementation, we can enable the
CPUID bit for it.
2022-06-02 15:59:35 -04:00
lioncash 0f59a18223 unittests: Disable SHA256 tests
Currently our x86 CI doesn't have SHA instruction extensions.
2022-06-02 15:58:17 -04:00
lioncash 3402cde334 OpcodeDispatcher: Implement SHA256RNDS2 2022-06-02 15:56:22 -04:00
lioncash 8f53c6bb96 OpcodeDispatcher: Implement SHA256MSG2 2022-06-02 15:12:46 -04:00
lioncash 0d6e4631a3 OpcodeDispatcher: Implement SHA256MSG1 2022-06-02 15:03:33 -04:00
Ryan Houdek 8dd9a5bd38 Merge pull request #1739 from lioncash/sha1
OpcodeDispatcher: Handle SHA-1 instructions
2022-06-02 11:28:54 -07:00
lioncash 903cf84874 unittests: Disable SHA-1 tests for now
Currently the x86 CI doesn't support the SHA instruction extension set
2022-06-02 14:02:30 -04:00
lioncash e997da48c7 OpcodeDispatcher: Implement SHA1RNDS4 2022-06-02 13:37:42 -04:00
lioncash 5ff89fd171 OpcodeDispatcher: Implement SHA1MSG2 2022-06-02 13:37:42 -04:00
lioncash fad4254c0e OpcodeDispatcher: Implement SHA1MSG1 2022-06-02 13:37:42 -04:00
lioncash 2fb3c4f11c OpcodeDispatcher: Implement SHA1NEXTE 2022-06-02 13:37:42 -04:00
Stefanos Kornilios Mitsis Poiitidis ce5297b75f Merge pull request #1738 from Sonicadvance1/workaround_tests
unittests: Workaround runner issues
2022-06-02 15:48:51 +03:00
Ryan Houdek b2b0c277f6 github: Disable struct verifier
Needs to be validated again. Xavier is really hating it.
2022-06-02 05:00:10 -07:00
Ryan Houdek 46919979ce StructVerifier: Ensure drm include is in place
drm testing disabled while investigations occur
2022-06-02 04:46:00 -07:00
Ryan Houdek 8b716c6a22 unittests ASM: Disable failing ARMv8 tests 2022-06-02 04:40:13 -07:00
Ryan Houdek 75090f8f6c GVisor: Disable failing unit tests 2022-06-02 04:40:13 -07:00
Stefanos Kornilios Mitsis Poiitidis c14c0c2e3b Merge pull request #1736 from Sonicadvance1/argument_injector
AppConfig: Inject --no-sandbox in to steamwebhelper
2022-06-01 10:19:07 +03:00
Ryan Houdek 95efd18b73 AppConfig: Inject --no-sandbox in to steamwebhelper
Steam's webhelper has started enabling its sandbox which completely
breaks under FEX since we don't support seccomp.

Curiously the Chromium code is actually supposed to support a fallback
namespace only mode, which is used in glibc 2.34 environments.

The startup script for this will try to use this namespace only mode,
but Chromium developers never tested this in an environment that doesn't
support seccomp.

Due to an early check in their sandbox code, it checks for bpf support
before checking for which sandbox mode it is entering. This returns
early with a false statement which brings the entire browser instance
down with an assert.

Inject the --no-sandbox argument so we get around this and the sandbox
is disabled.
2022-05-31 17:12:08 -07:00
Ryan Houdek c9319a768f Config: Add the ability to inject command line arguments
Simple enough since we control the full emulation
2022-05-31 17:11:42 -07:00
Mai M da48020882 Merge pull request #1730 from Sonicadvance1/support_pause
OpcodeDispatcher: Implements support for PAUSE
2022-05-26 19:20:04 -04:00
Ryan Houdek ce4380e136 unittests: Adds basic PAUSE test 2022-05-26 08:51:00 -07:00
Ryan Houdek c633661121 OpcodeDispatcher: Implements support for PAUSE
The pause instruction is architecturally defined to be the `REP NOP`
instruction.

This allows people to use this instruction as a backwards compatible
pause without checking for CPUID support. In fact there is no way to
check if the hardware implements this as a `REP NOP` or a `PAUSE`.

If you have new enough hardware then this just ends up being a PAUSE.

Pass this PAUSE over to our host to help out applications that are
writing spin loops with a PAUSE in it, which we were deleting
previously.
2022-05-26 08:47:56 -07:00
Ryan Houdek 4bfd1dde1f IR: Implements support for Yield IR op
This is an IR op that produces nor consumes any SSA values, but has side
effects.

Turns in to the pause instruction on x86 and yield instruction on
AArch64.
2022-05-26 08:46:14 -07:00
Mai M fe11bd2242 Merge pull request #1726 from Sonicadvance1/fix_pextrb
OpcodeDispatcher: Fixes pextrb with high registers
2022-05-24 23:05:51 -04:00
Ryan Houdek 814f0c3c93 unittests: Adds unit test for the high pextrb encoding
Previous unit tests didn't cover this edge case.
2022-05-24 16:31:09 -07:00
Ryan Houdek 7e904056d3 OpcodeDispatcher: Fixes pextrb with high registers
Previously if the instruction was encoded to use rsp, rbp, rsi, or rdi
then due to how these were encoded in modrm this would hit the frontend
path for writing to the high 8 bits of a 16bit register.

This is because it's instruction specific if an 8-bit modrm instruction
chooses to use the high 8-bit region or the upper 4 registers.
See the ModRM.reg section of `ModRM.reg and .r/m Field Encodings`
specifically to see what each encoding stands for. Has four different
meanings per encoding depending on instruction.

This instruction doesn't actually write to registers at 8-bit size, it
extracts an element at 8-bit size and then zero extends it to the full
GPR.

When storing to memory it always stores to memory at the size of the
element extracted.

I grepped around the instruction tables to see if there were any other
instances of this mistake. This was the only one.

Fixes #1472 and also gets Psychonauts 2 running.
2022-05-24 16:23:22 -07:00
Mai M 969d8f866c Merge pull request #1724 from Sonicadvance1/v5.18
v5.18 support
2022-05-23 14:43:11 -04:00
Ryan Houdek a2d0b7d7c4 Linux: Expose v5.18 host to guest 2022-05-23 11:00:16 -07:00
Ryan Houdek 97a8fa77bc Ioctl: Update drm msm for v5.18
Fixes #1602
2022-05-23 10:59:29 -07:00
Ryan Houdek c1296cc64d Update external drm-headers 2022-05-23 10:58:50 -07:00
Ryan Houdek 1dee54a9d8 Merge pull request #1722 from Hypnotron/main
Fix dangling curl hyphen
2022-05-21 11:24:16 -07:00
The Hypnotron 17e5d73e64 Fix dangling curl hyphen
Fixes a regression in fa87c73b9ee60a334eace2cdc3097725cbaf5b88; curl complains "curl: option -: is unknown" when trying to fetch a RootFS without this.
2022-05-21 14:12:48 -04:00
Mai M d523b7a6c7 Merge pull request #1720 from Sonicadvance1/workaround_libstdcxx_bug
FEXLogServer: Stop improper use of std::erase_if
2022-05-20 09:02:38 -04:00
Ryan Houdek 24ad208778 FEXLogServer: Stop improper use of std::erase_if
std::erase_if shouldn't allow you to modify the object passed in to the
predicate.
libstdc++ hasn't always enforced this but now it does with libstdc++12
2022-05-20 03:47:24 -07:00
Ryan Houdek fa87c73b9e Merge pull request #1719 from Sonicadvance1/fex_rootfs_fetcher_no_rety
FEXRootFSFetcher: Don't continue download
2022-05-20 03:13:38 -07:00
Ryan Houdek 0ed96544e1 Merge pull request #1721 from Sonicadvance1/fix_clone3_stack
Syscalls: Fixes clone3 stack pointer
2022-05-20 02:46:05 -07:00
Ryan Houdek d28ccc59ac Syscalls: Fixes clone3 stack pointer
clone2 stack pointer passed in points to the highest address for the
stack.

clone3 switches this around and gives us a base pointer and a size.

glibc started using clone3 for its thread cloning which finally caught
this bug. Necessary to run any application under the Ubuntu 22.04 rootfs
since that uses a new enough glibc to encounter this.
2022-05-19 22:51:17 -07:00
Ryan Houdek eaa75c1ed2 FEXRootFSFetcher: Don't continue download
While our CDN supports download continue, the backblaze storage backing
does not.
2022-05-19 22:44:22 -07:00
Ryan Houdek 4f4263263b Merge pull request #1718 from FEX-Emu/skmp/fix-x86tables-leave
X86Tables: Leave shouldn't end block
2022-05-19 00:22:27 -07:00
Stefanos Kornilios Misis Poiitidis ae00654694 X86Tables: Leave shouldn't end block 2022-05-19 08:51:02 +03:00
Stefanos Kornilios Mitsis Poiitidis a7156276e9 Merge pull request #1716 from FEX-Emu/skmp/jitsymbols-file-offsets
JitSymbols: Print file+offset if possible
2022-05-17 16:06:18 +03:00
Stefanos Kornilios Misis Poiitidis 29859d2491 JitSymbols: Print file offsets if possible 2022-05-17 15:05:47 +03:00
Stefanos Kornilios Mitsis Poiitidis 5460a24ea9 Merge pull request #1558 from FEX-Emu/skmp/smc-memtrack
SMC detection via segfaults
2022-05-16 17:37:07 +03:00
Stefanos Kornilios Misis Poiitidis a284adcd19 SMC: Add mprotect based tracking, --smc=mtrack, make default 2022-05-16 15:51:09 +03:00
Stefanos Kornilios Mitsis Poiitidis 73d43c1d55 Merge pull request #1700 from FEX-Emu/skmp/standarized-todo
Standarized TODO markers: FEX_TODO, FEX_TODO_ISSUE
2022-05-16 14:17:16 +03:00
Stefanos Kornilios Misis Poiitidis 256df76674 FEX_TODO: Convert some XXX to FEX_TODO 2022-05-16 12:22:42 +03:00
Ryan Houdek c8dc663b0b Merge pull request #1709 from Sonicadvance1/remove_debug_statement
OpcodeDispatcher: Remove debugging dump statement
2022-05-14 19:31:01 -07:00
Ryan Houdek ba78dff1f8 Merge pull request #1707 from Sonicadvance1/non_temporal
OpcodeDispatcher: Adds support for non-temporal loadstores
2022-05-14 19:30:53 -07:00
Ryan Houdek 1e597bfbed Merge pull request #1706 from Sonicadvance1/ref_count_shared_mutex
FEXCore: Adds refcount_shared_mutex class
2022-05-14 19:30:43 -07:00
Ryan Houdek b78af2fdaf Merge pull request #1684 from Sonicadvance1/testharness_named_regions
TestHarnessRunner: Use guest mapper for test harness files
2022-05-14 19:18:13 -07:00
Ryan Houdek 5379f0a9c7 Merge pull request #1677 from neobrain/refactor_scopedsignalmask
Clean up and document ScopedSignalMask
2022-05-14 19:13:37 -07:00
Ryan Houdek f1f523e525 OpcodeDispatcher: Remove debugging dump statement 2022-05-14 19:09:10 -07:00
Ryan Houdek c79d79e08b TestHarnessRunner: Use guest mapper for test harness files
This will allow it to get picked up for named region handling. Thus
ending up in the code caching for testing.
2022-05-14 19:08:32 -07:00
Ryan Houdek ee2d417d21 Merge pull request #1691 from Sonicadvance1/object_cache_named_region
Object cache named region no-op implementation
2022-05-14 18:20:05 -07:00
Ryan Houdek cebdde599a ObjectCache: Adds no-op named region object loading
This does the setup for handling the named region object loading and
closing using the async interface.

This exercises the async interface while the async thread itself only
does the minimum no-op steps required to fake loading and saving.
2022-05-14 18:09:38 -07:00
Ryan Houdek c3ac72a01e Core: Do named region async code object cache usage 2022-05-14 18:02:43 -07:00
Ryan Houdek 13f3c6e75a Merge pull request #1690 from Sonicadvance1/job_handler
Core: Adds Code Object Cache service
2022-05-14 18:00:40 -07:00
Ryan Houdek b3cd4edb3b Core: Clear relocations after the cache service had a chance to copy them
Can't clear the relocations vector until after the code object cache
service has consumed them.
2022-05-14 17:49:18 -07:00
Ryan Houdek afe10c1666 Core: Adds Code Object Cache service
The no-op interface is hooked up to the point of exercising it in the
most minimal of sense.

If the configuration is set to enable read-only or read/write object
code then it will spin up the async worker thread as well, but it
doesn't do anything yet.
2022-05-14 17:38:20 -07:00
Ryan Houdek d9d30916ba ObjectCache: Adds no-op object cache
Currently unused.

Showcases the main interface in to the service implemented as no-ops
currently.
2022-05-14 17:38:20 -07:00
Ryan Houdek 3f6c1c0e68 JITs: Return pointer to internal relocation vectors
Currently unused.

This will be used by the Code Object Serialization service soon.
2022-05-14 17:36:15 -07:00
Ryan Houdek 5b2cc77109 CPUBackend: Adds RelocateJITObjectCode virtual function
When a backend supports relocations it will override this function.
returning nullptr meaning no relocation done.

Currently unused but will be soon.
2022-05-14 17:36:14 -07:00
Ryan Houdek b824023ec6 InternalThreadState: Adds Object Cache job ref counter mutex
Currently unused but will be used soon.

Removes old `IsCompileService` bool as well.
2022-05-14 17:36:14 -07:00
Ryan Houdek 6ce1be0880 InternalThreadState: Adds Relocations pointer to DebugData
This will be used soon to pass relocation data to the JIT object cache.
2022-05-14 17:36:14 -07:00
Ryan Houdek c5dacab2ee Merge pull request #1688 from Sonicadvance1/jit_relocations
JIT relocation handling support
2022-05-14 17:25:34 -07:00
Ryan Houdek 4d24b85d57 OpcodeDispatcher: Adds support for non-temporal loadstores
x86 has eight instructions that are non-temporal.
Only one of which is a load-NT.

Adds a memory access type classification to our LoadSource/StoreResult
helpers.

This lets us explicitly choose Default, TSO, NonTSO, and Stream.

Stream currently just behaves like NonTSO so at some point in the future
we can add non-temporal loadstores to the IR.

Main thing is to move these NT accesses to non-TSO.
2022-05-14 02:53:44 -07:00
Ryan Houdek e967b447e6 FEXCore: Adds refcount_shared_mutex class
This class is similar to std::shared_mutex except it is safe for the
same thread to increment or decrement the ref counter multiple times.

This can be passed to regular std locks.

This will be required with the code object cache service soon.
2022-05-14 00:51:15 -07:00
Ryan Houdek 9bc631a427 Merge pull request #1705 from Sonicadvance1/fix_fsgsbase
32-bit FSGS instruction fixes.
2022-05-14 00:04:42 -07:00
Ryan Houdek c480ef137d unittests: Only disable fsgs tests on host
Since the x86 CI machine doesn't have a new enough kernel for this.
2022-05-13 02:43:28 -07:00
Ryan Houdek 15629e790e unittests: Adds fsgsbase 32-bit tests
Ensures that the upper 32-bits are zero'd rather than inserted.
2022-05-13 02:42:28 -07:00
Ryan Houdek 3f08d8b691 OpcodeDispatcher: Fixes 32-bit fs/gs write instructions
Documentation claims that these insert the lower 32-bits leaving the
upper bits unaffected.
Hardware testing proves that the upper 32-bits of the base registers are
zero'd.

Additional documentation also concurs that this is the case.
2022-05-13 02:42:28 -07:00
Ryan Houdek 9d9d171aad OpcodeDispatcher: Only expose fsgs instructions in 64-bit
These aren't supported in 32-bit
2022-05-13 02:42:28 -07:00
Ryan Houdek 65218c8285 unittests: Update tests to use canonical addresses 2022-05-13 02:34:02 -07:00
Ryan Houdek 27f2e0b06d Merge pull request #1704 from Sonicadvance1/fix_instruction_rerun
Arm64: Fix LDAPUR/STLUR DMB backpatch
2022-05-13 00:49:21 -07:00
Ryan Houdek ae75983b54 Arm64: Fix LDAPUR/STLUR DMB backpatch
This ended up in the wrong commit. We need to rerun the DMB that we
patched in.
2022-05-12 08:03:18 -07:00
Ryan Houdek f8ba373e18 Merge pull request #1702 from Sonicadvance1/support_rcpc2
Arm64: Adds support for RCPC2 extension
2022-05-12 07:33:02 -07:00
Ryan Houdek 2feae06209 Arm64: Adds support for RCPC2 extension
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.

This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.

Apple M1 supports this extension, didn't test with Cortex-X2/A710.
2022-05-12 07:13:47 -07:00
Stefanos Kornilios Mitsis Poiitidis 2e0534924a Merge pull request #1699 from FEX-Emu/skmp/add-fwrapv
CMake: C/C++ flags for defined singed overflow warping
2022-05-11 11:10:46 +03:00
Stefanos Kornilios Misis Poiitidis ad1fd7f54b FexHeaderUtils: Add TodoDefines 2022-05-11 11:08:28 +03:00
Tony Wasserka 9aaace51e1 ScopedSignalMask: Add usage guidelines 2022-05-10 17:18:36 +02:00
Tony Wasserka 58841142ee Merge pull request #1693 from neobrain/feature_linker
CMake: Add option to use the mold linker
2022-05-10 16:29:39 +02:00
Stefanos Kornilios Misis Poiitidis a6a816fb38 CMake: C/C++ flags for defined singed overflow warping 2022-05-10 17:14:17 +03:00
Ryan Houdek 70988ccfee Merge pull request #1694 from Sonicadvance1/fix_RCPC
Arm64: Fixes AtomicSwap
2022-05-09 23:30:34 -07:00
Ryan Houdek a8d9caf0c0 x64Jit: Adds relocation handling support
The JIT currently doesn't use this. This is just the handling code
itself.

One line disabled handling Guest RIP move relocations until the JIT
object cache is enabled.
2022-05-09 19:56:33 -07:00
Ryan Houdek bc22186093 Arm64: Adds relocation handling support
The JIT currently doesn't use this. This is just the handling code
itself.

One line disabled handling Guest RIP move relocations until the JIT
Object cache is enabled.
2022-05-09 19:56:33 -07:00
Ryan Houdek 317416b2e0 Context: Adds Cache object code config option 2022-05-09 19:56:32 -07:00
Ryan Houdek b5ae9e4c97 Merge pull request #1686 from Sonicadvance1/add_relocation_definitions
ArchHelpers: Adds relocation struct defines
2022-05-09 19:49:02 -07:00
Ryan Houdek 099737ca05 ArchHelpers: Adds relocation struct defines
Pulled from #1548 with one of the unused relocation types removed.

Unused for now.
2022-05-09 19:30:09 -07:00
Ryan Houdek e90164b519 Arm64: Fixes AtomicSwap
It wasn't using acquire semantics, only release semantics.
This was causing the swap to load data from a stale cacheline, causing
the futex system in glibc to break.

This break only occured if you tried going down the RCPC codepath
because of edge case memory ordering problems.

This then enables the RCPC code path now since it works.
2022-05-09 16:50:43 -07:00
Tony Wasserka 933c1af7e8 CMake: Add option to use the mold linker 2022-05-09 17:00:10 +02:00
Stefanos Kornilios Mitsis Poiitidis b9d878b1f4 Merge pull request #1672 from FEX-Emu/skmp/add-guest-mmap-munmap
Syscalls/Linux: Add guest[Mmap/Munmap]
2022-05-09 12:55:40 +03:00
Parallels 45a9a83c79 Loaders: Use bind_front instead of lambdas to bind GuestM(un)map 2022-05-09 12:39:48 +03:00
Parallels 560cfc757c HarnessHelper: Allocations need to be MAP_FIXED 2022-05-09 11:57:19 +03:00
Stefanos Kornilios Misis Poiitidis e92f51e415 Syscalls/Linux: Add GuestMmap & GuestMunmap, update code to use it 2022-05-09 11:57:08 +03:00
Ryan Houdek 278ca52d97 Merge pull request #1683 from Sonicadvance1/code_cache_config
Config: Adds code cache config option
2022-05-08 18:44:32 -07:00
Ryan Houdek 912dbfe5bd Merge pull request #1689 from Sonicadvance1/AArch64_MoveConstant_ADR
Arm64Emitter: Optimize constants with ADRP and ADR
2022-05-08 18:43:36 -07:00
Ryan Houdek 687f46fc71 Arm64Emitter: Optimize constants with ADRP and ADR
In a large number of cases we are moving pointers within a 4GB region
and some marginal pointers that are within 1MB.

This is only used in the case that MOVZ can't be used.

NOP padding still occurs after these instructions to ensure that if they
are being used with relocations it will still get padded to a full 4
instruction length.

Not all hardware fuses these and LLVM claims that Cortex beyond A72 even
doesn't, but it'll still be faster.
2022-05-08 18:33:43 -07:00
Ryan Houdek 6e9e5b3bd6 Config: Adds code object cache config option
This will be used soon
2022-05-06 10:26:44 -07:00
Stefanos Kornilios Mitsis Poiitidis 4fbc266b18 Merge pull request #1685 from Sonicadvance1/fix_tmp_file_flags
EmulatedFiles: Fixes temporary file flags
2022-05-06 10:38:42 +03:00
Ryan Houdek d1ac406895 EmulatedFiles: Fixes temporary file flags
mode and flags were being combined incorrectly.
2022-05-05 21:50:34 -07:00
Stefanos Kornilios Mitsis Poiitidis ce0f5db6f7 Merge pull request #1671 from FEX-Emu/skmp/refactor-guest-mman-tracking
Syscalls/Linux: Refactor guest mman tracking
2022-05-03 15:37:55 +03:00
Stefanos Kornilios Misis Poiitidis efb42c1ad1 Syscalls/Linux: Refactor guest mman tracking 2022-05-03 15:26:32 +03:00
Stefanos Kornilios Mitsis Poiitidis d8109880f4 Merge pull request #1670 from FEX-Emu/skmp/processwide-code-invalidations
Core: context-wide guest code invalidations
2022-05-03 15:23:10 +03:00
Stefanos Kornilios Misis Poiitidis 09be28a443 LookupCache: Cleanups 2022-05-02 16:59:50 +03:00
Tony Wasserka fb0bb8dd2c ScopedSignalMask: Unify implementation 2022-05-02 11:27:26 +02:00
Ryan Houdek 8e36f5331f Merge pull request #1669 from FEX-Emu/skmp/movable-lock-guards
ScopedSignalMask: Add shared mutex support, move constructors
2022-05-01 16:27:47 -07:00
Stefanos Kornilios Misis Poiitidis a365a70275 Review feedback 2022-05-02 01:51:17 +03:00
Stefanos Kornilios Misis Poiitidis b5a4e5920d Core: Rename FlushCodeRange to InvalidateGuestCodeRange 2022-05-02 01:49:54 +03:00
Stefanos Kornilios Misis Poiitidis 94d2ed85a7 Core: Add support for process-wide code invalidation, rename IR invalidate op to do thread specific invalidation 2022-05-02 01:49:49 +03:00
Ryan Houdek 90f338d7db Merge pull request #1674 from FEX-Emu/skmp/fexloader-fix-aotir-create_directories
FEXLoader: Fix create_directories check for aotir .path file writting
2022-05-01 15:36:28 -07:00
Stefanos Kornilios Misis Poiitidis df78f5d50e FEXLoader: Fix create_directories check for aotir .path file writting 2022-05-02 01:25:02 +03:00
Ryan Houdek b2b4c2bdcf Merge pull request #1673 from FEX-Emu/skmp/shmdt-fixes
Linux/MemAllocator32Bit: Add missing lock to shmdt, fix error returns
2022-05-01 15:22:08 -07:00
Stefanos Kornilios Misis Poiitidis a888da436b Linux/MemAllocator32Bit: Add missing lock to shmdt, fix error returns 2022-05-02 01:00:53 +03:00
Stefanos Kornilios Misis Poiitidis 3cb8ae9a9c ScopedSignalMask: Add shared mutex support, move constructors 2022-04-30 16:00:00 +03:00
Ryan Houdek db3854e391 Merge pull request #1664 from CallumDev/f64-fldcw-impl
F64: Implement FCW using host rounding mode
2022-04-28 15:11:43 -07:00
CallumDev 05b4b095fe F64: Set host RoundingMode for all FCW loads 2022-04-29 05:07:44 +09:30
CallumDev 93926641d9 F64: Implement FLDCW using host rounding mode 2022-04-29 04:39:57 +09:30
Ryan Houdek 89d6752d3d Merge pull request #1662 from CallumDev/f64-int-fixes
F64: Fix FILD and FIST for Size < 8
2022-04-28 08:28:11 -07:00
CallumDev e4f95fec79 F64: Fix FILD and FIST for Size < 8 2022-04-29 00:42:56 +09:30
Ryan Houdek 8a7f39559c Merge pull request #1627 from Sonicadvance1/wip_reclaimable_pool_allocator
FEXCore: Reclaimable thread pool allocator
2022-04-26 10:21:36 -07:00
Ryan Houdek 753d0ede6c FEXCore: Reclaimable thread pool allocator
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.

A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.

Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
 these stale allocations and reclaim it from the idling thread. Saving
further memory.

This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.

Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
2022-04-26 10:01:56 -07:00
Ryan Houdek da2e44d024 Merge pull request #1658 from wannacu/main
AOTIR: copy RAData and IRList, make sure data is accessible
2022-04-25 19:04:23 -07:00
wannacu 7b379fc3cf AOTIR: copy RAData and IRList, make sure data is accessible 2022-04-26 09:27:18 +08:00
Ryan Houdek ec38d58b37 Merge pull request #1659 from Sonicadvance1/fexbash_ps1
FEXBash: Set PS1 to make it more obvious when running under FEX
2022-04-25 12:30:17 -07:00
Ryan Houdek f6a74a710d FEXBash: Set PS1 to make it more obvious when running under FEX
This requires us to pass in --no-rc to bash since otherwise PS1 gets
overwritten by shell variables and nothing happens.

Which this is fine for the common use case of just wanting to run a
basic bash script under emulation.
2022-04-25 11:59:04 -07:00
Ryan Houdek 3fd136b0da Merge pull request #1657 from Sonicadvance1/fix_32bit_mmap
Linux: Fixes 32-bit mmap
2022-04-24 11:17:09 -07:00
Ryan Houdek 72e82d0304 Linux: Fixes 32-bit mmap
This went unnoticed for so long since most applications are new enough
to use mmap2 instead of mmap.
This was just completely broken.

Fixes #1630
2022-04-24 10:56:56 -07:00
Ryan Houdek 2f7dcb8d93 Merge pull request #1656 from Sonicadvance1/v5.17_support
V5.17 support
2022-04-24 10:55:39 -07:00
Ryan Houdek 128a24d699 Linux: Updates supported guest Linux version to v5.17 2022-04-23 11:59:18 -07:00
Ryan Houdek cd94a8f0ac Linux: Adds support for new v5.17 virtio IOCTL 2022-04-23 11:58:56 -07:00
Ryan Houdek df5e0e5df9 Linux: Adds support for new v5.17 syscall 2022-04-23 11:58:38 -07:00
Ryan Houdek 458bbf4ef7 Linux: Updates syscalls for v5.17 2022-04-23 11:57:37 -07:00
Ryan Houdek 82319c9deb Scripts: Updates generate syscall numbers to support renaming
Instead of manually renaming the three syscalls each time, let the
script do it automatically.
2022-04-23 11:56:33 -07:00
Ryan Houdek 3bc4df7295 Updates drm headers to v5.17 2022-04-23 11:56:06 -07:00
Ryan Houdek 42a6320935 Merge pull request #1585 from CallumDev/x87f64
Emulate reduced-precision X87 with 64-bit host FPU ops
2022-04-21 09:22:03 -07:00
CallumDev 843fe378db Document that X87ReducedPrecision reduces accuracy 2022-04-22 01:38:18 +09:30
CallumDev 3679673d5b Implement FNSAVE and FRSTOR in F64 2022-04-22 01:36:56 +09:30
CallumDev 3a269f04d2 Add remaining possible X87F64 tests. Tweak FPREM 2022-04-22 01:36:56 +09:30
CallumDev 0021723b50 Fix F64 FSCALE 2022-04-22 01:36:56 +09:30
CallumDev 1c4b0272e8 X87F64: Implement FXTRACT using bit ops, add test 2022-04-22 01:36:56 +09:30
CallumDev 9640216124 X87F64: Add working BCD test 2022-04-22 01:36:56 +09:30
CallumDev 0f8f2bf2b4 X87F64 fix integer load, add tests 2022-04-22 01:36:56 +09:30
CallumDev ef899b7b1a F64: Working FABS and FCHS 2022-04-22 01:36:56 +09:30
CallumDev b85abf725d Implement FCOM in 64-bit ops 2022-04-22 01:36:56 +09:30
CallumDev 3f89e46d66 Add X87ReducedPrecision to FEXConfig 2022-04-22 01:36:56 +09:30
CallumDev d02712ddc6 X87F64: Basic implementation of FRNDINT 2022-04-22 01:36:56 +09:30
CallumDev 91f48c63ff Implement F64 ops in Interpreter 2022-04-22 01:36:56 +09:30
CallumDev cd16769e57 F64: Fix Arm64 JIT compile error 2022-04-22 01:36:56 +09:30
CallumDev 0a778e802f Unit Tests for x87F64 2022-04-22 01:36:56 +09:30
CallumDev d4d5f4d1dd Introduce F64 codegen for reduced precision X87 2022-04-22 01:36:44 +09:30
Mai M b1033ed7c6 Merge pull request #1652 from Sonicadvance1/remove_compile_service
CompileService: Removes no longer necessary service thread
2022-04-19 22:34:15 -04:00
Ryan Houdek 253333a4cf CompileService: Removes no longer necessary service thread
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.

This makes the compile service never be invoked so we can just remove
it.

We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
2022-04-19 18:52:01 -07:00
Ryan Houdek 50595ac3a9 X86Dispatcher: Disable signals when compiling just like on AArch64 2022-04-19 18:29:21 -07:00
Ryan Houdek 37f1e55ed5 Docs: Update for release FEX-2204 2022-04-19 01:19:00 -07:00
Ryan Houdek 8ad14728f6 Merge pull request #1644 from Sonicadvance1/ldiv_minor_opt
JITArm64: Get long divide out of the hot path
2022-04-01 18:23:30 -07:00
Ryan Houdek 6b3cd3d31d Merge pull request #1645 from Sonicadvance1/update_aarch64_fit
Scripts: Updates AArch64 fit for Clang 14
2022-04-01 18:23:13 -07:00
Ryan Houdek fba698cb74 Scripts: Updates AArch64 fit for Clang 14
Clang now supports these latest ARMv9 CPUs
2022-04-01 18:08:02 -07:00
Ryan Houdek 0946b123bb JITArm64: Get long divide out of the hot path
For 128-bit divides, we can very quickly check at runtime if we can
avoid the long divide and just do a 64-bit divide.

For unsigned just check if the top bits are all zero.
For signed just check if the top bits match bit 63 of the lower bits.

Additionally, keep the long divide handlers inside of the dispatcher.
This keeps the majority of the code bloat out of the code block itself,
significantly reducing block size for something doing these divides.
Also a fairly large icache improvement from this.

Hard performance number improvements here are hard to get since it
heavily depends on the application, also only occurs on x86-64.

Seems to have helped FTL and Dead Cells performance quite a bit though.
2022-03-31 09:33:19 -07:00
Ryan Houdek b43937a7a1 Merge pull request #1643 from Sonicadvance1/fix_termux
SignalDelegator: Adds missing include
2022-03-29 21:05:28 -07:00
Ryan Houdek 4564eba20d SignalDelegator: Adds missing include
Fixes Termux building.
Fixes #1642
2022-03-29 20:47:22 -07:00
Ryan Houdek 5cc0c0a3da Merge pull request #1641 from philpax/docs-remove-stale-text
docs: Remove stale text
2022-03-29 02:56:20 -07:00
Philpax f8e7c75f86 docs: Remove stale text 2022-03-29 11:27:43 +02:00
Ryan Houdek 042cd354dc Merge pull request #1633 from Sonicadvance1/disable_instructions_on_host_missing
OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
2022-03-23 13:48:04 -07:00
Ryan Houdek 977bda97b2 Merge pull request #1635 from Sonicadvance1/4000_0001h
CPUID: Adds 4000_0001h function
2022-03-23 13:42:45 -07:00
Ryan Houdek 4cf48ca9bb CPUID: Adds 4000_0001h function
Exposes the host architecture through this CPUID function. Only exposes
the architectures we support. Not burning 16-bits on using ELF machine
definitions here.

Uses 4 bits still for future expansion.
2022-03-22 16:53:44 -07:00
Mai M a247df50ea Merge pull request #1624 from Sonicadvance1/cleanup_ir_after_use
FEXCore: Delete IR after it is used
2022-03-22 13:23:05 -04:00
Mai M 1f1c214944 Merge pull request #1634 from Sonicadvance1/cpuid_documentation
Documentation: Adds hypervisor CPUID information
2022-03-22 12:57:15 -04:00
Ryan Houdek d16db4ebde OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
If the host doesn't support the instructions required for implementing
an instruction then don't even add them to the opcodedispatcher.

This means that we will never try emitting instructions that the host
doesn't support (For these instructions anyway) and successfully passes
the guest SIGILL for these particular instructions.

Fixes #1631
2022-03-21 23:03:04 -07:00
Ryan Houdek ae1c563082 Documentation: Adds hypervisor CPUID information
Currently we only implement function 4000_0000h. This will expand in the
future but this is all we have right now.
2022-03-21 22:46:48 -07:00
Ryan Houdek ebd0edbab7 Merge pull request #1632 from FEX-Emu/skmp/flush-test-harness
TestHarnessRunner: Flush log on asserts
2022-03-21 12:50:43 -07:00
Stefanos Kornilios Misis Poiitidis e87e9d269a TestHarnessRunner: Flush log on asserts 2022-03-21 21:32:35 +02:00
Ryan Houdek 187c64182b Merge pull request #1628 from Sonicadvance1/fix_finit
OpcodeDispatcher: Fixes FNINIT
2022-03-17 20:37:57 -07:00
Ryan Houdek 60c7ea6e5f Merge pull request #1620 from Sonicadvance1/fix_1618
FEXCore: Fixes #1618
2022-03-17 20:36:06 -07:00
Ryan Houdek 6f1b4b0eee OpcodeDispatcher: Fixes FNINIT
Was incorrectly setting the FCW to 037h when it was supposed to be
037Fh.

Fixes a bug in a visual novel where its CPUID state wouldn't initialize
if this was set incorrectly.
2022-03-17 20:27:22 -07:00
Ryan Houdek fb69300397 FEXCore: Delete IR after it is used
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.

This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.

This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
2022-03-13 19:01:40 -07:00
Ryan Houdek 5677924525 Merge pull request #1621 from Sonicadvance1/fix_1584
Softfloat: Fixes FSCALE
2022-03-13 18:57:06 -07:00
Ryan Houdek 8422fc632d Merge pull request #1623 from Sonicadvance1/remove_unused_debug_data
FEXCore: Removes unused debug data
2022-03-13 18:56:50 -07:00
Ryan Houdek d33cd744fb FEXCore: Removes unused debug data
This isn't used anywhere. Just remove these.
If we get the imgui debugger running again then we can add even more
stats to sort block costs by.
2022-03-13 18:40:49 -07:00
Ryan Houdek 3b0fb27ae9 Softfloat: Fixes FSCALE
I misread the implementation details of this instruction when
implementing.

The pseudocode says `ST(0) = ST(0) ∗ 2^rndint(ST(1))` so I understood
the instruction to use the current rounding mode of the host to extract
the integer portion of `ST(1)`.

The actual implementation is in the details of the statement `the
integer portion of the floating- point value in ST(1).`

This behaves like round towards zero/truncate, additional hardware
testing and documentation reading confirms this.

Fixes #1584
2022-03-13 14:11:31 -07:00
Ryan Houdek 4603e09a04 FEXCore: Fixes #1618 2022-03-13 13:42:37 -07:00
Ryan Houdek 7b0265ffe2 Merge pull request #1617 from Sonicadvance1/gdbstub_improvements4
GDBServer improvements: Three's a crowd
2022-03-13 13:24:41 -07:00
Ryan Houdek ec54560a38 GDBServer: Fixes memory reading
memory-map is not something we want to use. Adds a comment about it and
disables it.

Also changes core events to wait for an event from GDBStub for waking up
which fixes a hang.
2022-03-13 13:05:00 -07:00
Ryan Houdek 6a5abd3672 Merge pull request #1616 from Sonicadvance1/gdbstub_improvements3
Gdbstub improvements: The sequel
2022-03-13 13:03:25 -07:00
Ryan Houdek 3e6af39c42 GDBServer: Zero initialize some variables to fix connection stability
Otherwise you always had to attempt connecting twice in a row
2022-03-13 12:50:01 -07:00
Ryan Houdek cf82ffc052 GDBServer: Let gdb know when the library map has updated
We need to fetch the full list of map files from the memory map and hand
it over to gdb.
It will then fetch all the libraries from the remote host and give us
backtraces
2022-03-13 12:50:01 -07:00
Ryan Houdek b190150281 FEXCore: Merges redundant string trimming implementations 2022-03-13 12:50:01 -07:00
Ryan Houdek 53ffe5df43 Merge pull request #1613 from Sonicadvance1/gdbstub_improvements2
GDBServer improvements
2022-03-13 12:49:17 -07:00
Ryan Houdek d39df8d3ed Merge pull request #1614 from Sonicadvance1/add_comment
JIT: Adds comment to EmitDetectionString
2022-03-10 14:53:41 -08:00
Ryan Houdek eeb2b928b9 JIT: Adds comment to EmitDetectionString 2022-03-10 14:28:45 -08:00
Ryan Houdek 6cf24a748f GdbServer: Document what PassSignals is for 2022-03-10 14:23:21 -08:00
Ryan Houdek 3b4fd180de GDBServer: Pass auxv better
Fixes 32-bit auxv as well.
2022-03-10 14:21:46 -08:00
Ryan Houdek 7300c7a853 SignalDelegator: Remove anti-pattern usage 2022-03-10 14:00:51 -08:00
Ryan Houdek 376f6db3ac GDBServer: Reformat code to two space tabs
No functional change
2022-03-10 14:00:51 -08:00
Ryan Houdek ad3a960717 GDBServer: Support sending gdb the correct signal
Instead of just sending SIGSEGV, pass the real signal
2022-03-10 13:57:17 -08:00
Ryan Houdek 0a1ef867ee SignalDelegator: Support multiple backend host handlers
This will be necessary for gdbserver to handle signals indepedentally of
the FEX handling.
2022-03-10 13:57:17 -08:00
Ryan Houdek a7fe69deea CPU: Stop trying to initialize signal handlers per thread
These are static per process and only need to be initialized once.
We are going to support multiple signal handlers from the backend after
this, so can only install once.
2022-03-10 13:57:17 -08:00
Ryan Houdek a8c3b6d46f GDBServer: Capture signal capture numbers
This will allow us to wire this to a future signal handler for gdbserver
2022-03-10 13:57:17 -08:00
Ryan Houdek 60db2655bb GDBServer: Encode the return to pread correctly
This encodes the resulting data as raw binary rather than any special
escaped encoding
2022-03-10 13:57:10 -08:00
Ryan Houdek fd717b6995 GDBServer: Expose program offsets better
We were incorrectly returning programing offsets
Get the program offset from the frontend so we can know what to give gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek ba37388fe3 GDBServer: Expose auxv values
We already expose these in the code loader, pump it through gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 68f32d85c9 CodeLoader: Expose base ELF loaded offset
Useful for gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 2aa77e85de LinuxSyscalls: Expose CodeLoader through syscall interface 2022-03-09 19:07:45 -08:00
Ryan Houdek 91665fdf0e Merge pull request #1610 from wannacu/main
FileManager: Fix realpath failed on debian buster
2022-03-09 18:58:52 -08:00
Ryan Houdek 23a1c64bf7 Merge pull request #1612 from Sonicadvance1/tag_memory_allocations
JITs: Emit identification string in the code buffers
2022-03-09 18:52:46 -08:00
Ryan Houdek c2dcf06632 JITs: Emit identification string in the code buffers
At the start of each code buffer, emit a small string for letting memory
inspection know if a code region is for the JITs.
2022-03-09 18:29:37 -08:00
Ryan Houdek fad91bb818 Merge pull request #1609 from Sonicadvance1/fix_map_32bit
Linux: Fixes MAP_32BIT supported range
2022-03-08 17:33:34 -08:00
wannacu 0e769ece26 docs: Update Readme_CN.md 2022-03-08 18:10:08 +08:00
wannacu 9a780b40a2 Docs: Add Chinese README 2022-03-08 18:04:09 +08:00
wannacu 898873e9e3 FileManager: Fix realpath failed on debian buster
This happend on debian buster when run realpath(i386) on arm64 host.
2022-03-08 16:42:05 +08:00
Ryan Houdek 52292e5f7e Linux: Fixes MAP_32BIT supported range
I accidentally committed a 32-bit range that was significantly smaller
than what it should be.
While the minimal range worked for simple cases, it didn't work for
anything complex.
Give it the full range it needs.

Fixes #1600
2022-03-06 17:46:17 -08:00
Ryan Houdek 5de6c866b7 Merge pull request #1608 from Sonicadvance1/termux_build_option
Adds a cmake option for forcing a termux build
2022-03-06 13:54:20 -08:00
Ryan Houdek ec0cd3aec4 Adds a cmake option for forcing a termux build
This is necessary when cross-compiling rather than building on-device
2022-03-06 12:56:25 -08:00
Mai M f5f9512d9a Merge pull request #1606 from Sonicadvance1/fhu_page_size
Change page define usages over to self-defined
2022-03-06 15:53:15 -05:00
Mai M fb27cb4356 Merge pull request #1607 from Sonicadvance1/disable_guis_termux
Disables GUI applications in a Termux build
2022-03-06 15:52:37 -05:00
Mai M 94664580c8 Merge pull request #1605 from Sonicadvance1/update_docs_termux
Update ReleaseProcess docs for Termux
2022-03-06 15:52:10 -05:00
Ryan Houdek 99a93fa9ea Disables GUI applications in a Termux build 2022-03-06 08:09:55 -08:00
Ryan Houdek 4cb6918506 Change page define usages over to self-defined
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.

Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
2022-03-06 07:33:10 -08:00
Ryan Houdek 9cc743bf84 Update ReleaseProcess docs for Termux
FEX hardly works on Termux as-is, but we should make sure to document
how to update the packages otherwise we will quickly become outdated on
their package management.
2022-03-06 05:59:33 -08:00
444 changed files with 20551 additions and 7507 deletions

No files matched your search

+26 -1
View File
@@ -29,6 +29,20 @@ jobs:
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
@@ -51,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -153,6 +167,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+26 -61
View File
@@ -1,16 +1,22 @@
cmake_minimum_required(VERSION 3.14)
project(FEX)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
option(ENABLE_LLD "Enable linking with LLD" FALSE)
option(ENABLE_LLD "Enable linking with lld" FALSE)
option(ENABLE_MOLD "Enable linking with mold" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
@@ -21,9 +27,9 @@ option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
set (X86_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86.cmake" CACHE FILEPATH "Toolchain file for the x86 (cross-)compiler")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
# These options are meant for package management
@@ -41,6 +47,12 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_GDB_SYMBOLS)
message(STATUS "GDBSymbols support enabled")
add_definitions(-DGDB_SYMBOLS_ENABLED=1)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
@@ -73,6 +85,7 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
set (X86_TOOLCHAIN_FILE "")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
@@ -99,9 +112,14 @@ if (ENABLE_COMPILE_TIME_TRACE)
endif()
set (PTHREAD_LIB pthread)
if (ENABLE_LLD)
if (ENABLE_LLD AND ENABLE_MOLD)
message (FATAL_ERROR "Cannot enable both lld and mold")
elseif (ENABLE_LLD)
set (LD_OVERRIDE "-fuse-ld=lld")
link_libraries(${LD_OVERRIDE})
add_link_options(${LD_OVERRIDE})
elseif (ENABLE_MOLD)
add_link_options("-fuse-ld=mold")
endif()
if (ENABLE_LIBCXX)
@@ -115,59 +133,7 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
# Check if the build target page size is 4096
include(CheckCSourceRuns)
check_c_source_runs(
"#include <unistd.h>
int main(int argc, char* argv[])
{
return getpagesize() == 4096 ? 0 : 1;
}"
PAGEFILE_RESULT
)
if (NOT ${PAGEFILE_RESULT})
message(FATAL_ERROR "Host PAGE_SIZE is not 4096. Can't build on this target")
endif()
include(CheckCXXSourceCompiles)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_SIZE;
}
"
HAS_PAGESIZE)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_SHIFT;
}
"
HAS_PAGESHIFT)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_MASK;
}
"
HAS_PAGEMASK)
if (NOT HAS_PAGESIZE)
add_definitions(-DPAGE_SIZE=4096)
endif()
if (NOT HAS_PAGESHIFT)
add_definitions(-DPAGE_SHIFT=12)
endif()
if (NOT HAS_PAGEMASK)
add_definitions("-DPAGE_MASK=(~(PAGE_SIZE-1))")
endif()
if(DEFINED ENV{TERMUX_VERSION})
if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
add_definitions(-DTERMUX_BUILD=1)
set(TERMUX_BUILD 1)
# Termux doesn't support Jemalloc due to bad interactions between emutls, jemalloc, and scudo
@@ -176,7 +142,7 @@ endif()
if (ENABLE_STATIC_PIE)
if (_M_ARM_64 AND ENABLE_LLD)
message (FATAL_ERROR "Static linking does not currently work with AArch64+LLD. Use GNU ld for now.")
message (FATAL_ERROR "Static linking does not currently work with AArch64+lld. Use GNU ld for now.")
endif()
file(WRITE ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt.c
@@ -558,8 +524,7 @@ if (BUILD_THUNKS)
BINARY_DIR "Guest"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DX86_C_COMPILER:STRING=${X86_C_COMPILER}"
"-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"AdditionalArguments": "--no-sandbox"
}
}
+60 -239
View File
@@ -6,18 +6,10 @@
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGL.so",
"/usr/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/lib/x86_64-linux-gnu/libGL.so",
"/lib/x86_64-linux-gnu/libGL.so.1",
"/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/lib/x86_64-linux-gnu/libGL.so.1.7.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.2.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.7.0"
]
},
"GLESv2": {
@@ -26,322 +18,151 @@
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/lib/x86_64-linux-gnu/libGLESv2.so",
"/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libX11.so",
"/usr/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libX11.so",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/lib/x86_64-linux-gnu/libX11.so",
"/lib/x86_64-linux-gnu/libX11.so.6",
"/lib/x86_64-linux-gnu/libX11.so.6.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6.4.0"
]
},
"Vulkan-radeon": {
"Library": "libvulkan_radeon-guest.so",
"Vulkan": {
"Library": "libvulkan-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/lib/x86_64-linux-gnu/libvulkan_radeon.so"
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so.1",
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
],
"Comment": [
"Vulkan library relies on xcb, otherwise it crashes with jemalloc"
]
},
"Vulkan-lavapipe": {
"Library": "libvulkan_lvp-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/lib/x86_64-linux-gnu/libvulkan_lvp.so"
]
},
"Vulkan-freedreno": {
"Library": "libvulkan_freedreno-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/lib/x86_64-linux-gnu/libvulkan_freedreno.so"
]
},
"Vulkan-intel": {
"Library": "libvulkan_intel-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/lib/x86_64-linux-gnu/libvulkan_intel.so"
]
},
"Vulkan-panfrost": {
"Library": "libvulkan_panfrost-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/lib/x86_64-linux-gnu/libvulkan_panfrost.so"
]
},
"Vulkan-nvidia": {
"Library": "libvulkan_nvidia-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/usr/local/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/lib/x86_64-linux-gnu/libGLX_nvidia.so.0"
],
"Comment": [
"Not currently wired up"
]
},
"Vulkan-virtio": {
"Library": "libvulkan_virtio-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/lib/x86_64-linux-gnu/libvulkan_virtio.so"
]
},
"xcb": {
"Library": "libxcb-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb.so",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/lib/x86_64-linux-gnu/libxcb.so",
"/lib/x86_64-linux-gnu/libxcb.so.1",
"/lib/x86_64-linux-gnu/libxcb.so.1.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb_dri2-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb_dri3-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb_xfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb_shm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb_sync-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/lib/x86_64-linux-gnu/libxcb-sync.so",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb_randr-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb_present-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-present.so",
"/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb_glx-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libshmfence-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/lib/x86_64-linux-gnu/libxshmfence.so",
"/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libdrm.so",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/lib/x86_64-linux-gnu/libdrm.so",
"/lib/x86_64-linux-gnu/libdrm.so.2",
"/lib/x86_64-linux-gnu/libdrm.so.2.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2.4.0"
]
},
"asound": {
"Library": "libasound-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libasound.so",
"/usr/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libasound.so",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/lib/x86_64-linux-gnu/libasound.so",
"/lib/x86_64-linux-gnu/libasound.so.2",
"/lib/x86_64-linux-gnu/libasound.so.2.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2.0.0"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXrender.so",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/lib/x86_64-linux-gnu/libXrender.so",
"/lib/x86_64-linux-gnu/libXrender.so.1",
"/lib/x86_64-linux-gnu/libXrender.so.1.3.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXext.so",
"/usr/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libXext.so",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/lib/x86_64-linux-gnu/libXext.so",
"/lib/x86_64-linux-gnu/libXext.so.6",
"/lib/x86_64-linux-gnu/libXext.so.6.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/lib/x86_64-linux-gnu/libXfixes.so",
"/lib/x86_64-linux-gnu/libXfixes.so.3",
"/lib/x86_64-linux-gnu/libXfixes.so.3.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3.1.0"
]
},
"":{}
+2 -12
View File
@@ -18,14 +18,8 @@ This project aims to provide a fast and functional x86-64 emulation library that
* Portable library implementation in order to support easy integration in to applications
### Target Host Architecture
The target host architecture for this library is AArch64. Specifically the ARMv8.1 version or newer.
The CPU IR is designed with AArch64 in mind but there is a desire to run the recompiled code on other architectures as well.
Multiple architecture support is desired for easier bringup and debugging, performance isn't as much of a priority there (ex. x86-64(guest) translated to x86-64(host))
### Not currently goals but will be in the future
* 32bit x86 support
* This will be a desire in the future, but to lower the amount of work required, decided to push this off for now.
* Integration in to WINE
* Later generation of x86-64 instruction sets
* Including AVX, F16C, XOP, FMA, AVX2, etc
The CPU IR is designed with AArch64 in mind but should allow for other architectures as well.
x86-64 host support is available for ease of development, but is not a priority.
### Not desired
* Kernel space emulation
* CPL0-2 emulation
@@ -33,7 +27,3 @@ Multiple architecture support is desired for easier bringup and debugging, perfo
* IRQs
* SVM
* "Cycle Accurate" emulation
### Dependencies
* clang-tidy if you want to ensure the code stays tidy
* cmake
* A C++17 compliant compiler (There are assumptions made about using Clang and LTO)
+15 -10
View File
@@ -80,16 +80,20 @@ set (SRCS
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
Interface/Core/CompileService.cpp
Interface/Core/Core.cpp
Interface/Core/CPUBackend.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/GdbServer.cpp
Interface/Core/HostFeatures.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
Interface/Core/OpcodeDispatcher/Vector.cpp
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/SignalDelegator.cpp
Interface/Core/X86Tables.cpp
@@ -114,6 +118,7 @@ set (SRCS
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/HLE/Thunks/Thunks.cpp
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
@@ -184,7 +189,9 @@ if (ENABLE_JIT_X86_64)
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
Interface/Core/JIT/x86_64/VectorOps.cpp
Interface/Core/JIT/x86_64/x64Relocations.cpp
)
list(APPEND DEFINES -DJIT_X86_64)
endif()
@@ -201,7 +208,9 @@ if (ENABLE_JIT_ARM64)
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
)
endif()
set (LIBS vixl dl xxhash tiny-json)
@@ -219,13 +228,11 @@ set(OUTPUT_IR_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/IR")
set(OUTPUT_NAME "${OUTPUT_IR_FOLDER}/IRDefines.inc")
set(INPUT_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/IR/IR.json")
add_custom_target(CREATE_IR_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_IR_FOLDER}")
file(MAKE_DIRECTORY "${OUTPUT_IR_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_NAME}"
DEPENDS "${INPUT_NAME}"
DEPENDS CREATE_IR_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py" "${INPUT_NAME}" "${OUTPUT_NAME}"
)
@@ -239,7 +246,6 @@ set(OUTPUT_IR_DOC "${CMAKE_BINARY_DIR}/IR.md")
add_custom_command(
OUTPUT "${OUTPUT_IR_DOC}"
DEPENDS "${INPUT_NAME}"
DEPENDS CREATE_IR_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py" "${INPUT_NAME}" "${OUTPUT_IR_DOC}"
)
@@ -260,15 +266,13 @@ set(INPUT_CONFIG_NAME "${CMAKE_BINARY_DIR}/generated/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
set(OUTPUT_MAN_NAME_COMPRESS "${CMAKE_BINARY_DIR}/generated/FEX.1.gz")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
file(MAKE_DIRECTORY "${OUTPUT_CONFIG_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_CONFIG_NAME}"
OUTPUT "${OUTPUT_CONFIG_OPTION_NAME}"
OUTPUT "${OUTPUT_MAN_NAME}"
DEPENDS "${INPUT_CONFIG_NAME}"
DEPENDS CREATE_CONFIG_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py" "${INPUT_CONFIG_NAME}" "${OUTPUT_CONFIG_NAME}" "${OUTPUT_MAN_NAME}"
"${OUTPUT_CONFIG_OPTION_NAME}"
@@ -329,6 +333,7 @@ function(AddDefaultOptionsToTarget Name)
-Wno-trigraphs
-ffunction-sections
-fwrapv
)
if (GCC_COLOR)
+8
View File
@@ -34,6 +34,14 @@ namespace FEXCore {
fmt::print(fp.get(), "{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
+1
View File
@@ -13,6 +13,7 @@ public:
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
+5 -1
View File
@@ -188,6 +188,10 @@ struct X80SoftFloat {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -257,7 +261,7 @@ struct X80SoftFloat {
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs);
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
+29
View File
@@ -0,0 +1,29 @@
#pragma once
#include <string>
namespace FEXCore::StringUtils {
// Trim the left side of the string of whitespace and new lines
[[maybe_unused]] static std::string LeftTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
// Trim the right side of the string of whitespace and new lines
[[maybe_unused]] static std::string RightTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
// Trim both the left and right of the string of whitespace and new lines
[[maybe_unused]] static std::string Trim(std::string String, std::string TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(String, TrimTokens), TrimTokens);
}
}
+25 -21
View File
@@ -1,4 +1,5 @@
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
@@ -370,29 +371,22 @@ namespace JSON {
return {};
}
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
std::string FindContainer() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
}
}
return String;
return {};
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
@@ -401,7 +395,7 @@ namespace JSON {
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
@@ -445,6 +439,16 @@ namespace JSON {
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
if (CacheObjectCodeCompilation() && Core() == FEXCore::Config::CONFIG_INTERPRETER) {
// If running the interpreter then disable cache code compilation
FEXCore::Config::Erase(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION);
}
}
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
+50 -15
View File
@@ -38,6 +38,17 @@
"Number of physical hardware threads to tell the process we have.",
"0 will auto detect."
]
},
"CacheObjectCodeCompilation": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE",
"TextDefault": "none",
"Choices": [ "none", "read", "readwrite" ],
"ArgumentHandler": "CacheObjectCodeHandler",
"Desc": [
"Cache JIT object code to drive.",
"Allows JIT code to be shared between applications"
]
}
},
"Emulation": {
@@ -104,6 +115,13 @@
"This can be useful for setting environment variables that thunks can pick up.",
"Typically isn't necessary since the guest libc isn't thunked. But is possible."
]
},
"AdditionalArguments": {
"Type": "strarray",
"Default": "",
"Desc": [
"Allows the user to pass additional arguments to the application"
]
}
},
"Debug": {
@@ -190,6 +208,16 @@
"Useful for determining hot blocks of code",
"Has some file writing overhead per JIT block"
]
},
"GDBSymbols": {
"Type": "bool",
"Default": "false",
"Desc": [
"Integrates with GDB using the JIT interface.",
"Needs the fex jit loader in GDB, which can be loaded via `jit-reader-load libFEXGDBReader.so.`",
"Also needs x86_64-linux-gnu-objdump in PATH.",
"Can be very slow."
]
}
},
"Logging": {
@@ -201,36 +229,28 @@
"Disables logging"
]
},
"OutputSocket": {
"Type": "str",
"Default": "",
"Desc": [
"Socket to connect to",
"eg: localhost:8087",
"If set will override the OutputLog location"
]
},
"OutputLog": {
"Type": "str",
"Default": "stderr",
"Default": "server",
"ShortArg": "o",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, <Filename>]"
"[stdout, stderr, server, <Filename>]"
]
}
},
"Hacks": {
"SMCChecks": {
"Type": "uint8",
"Default": "FEXCore::Config::CONFIG_SMC_MMAN",
"TextDefault": "mman",
"Default": "FEXCore::Config::CONFIG_SMC_MTRACK",
"TextDefault": "mtrack",
"ArgumentHandler": "SMCCheckHandler",
"Desc": [
"Checks code for modification before execution.",
"\tnone: No checks",
"\tmman: Invalidate on mmap, mprotect, munmap",
"\tfull: Validate code before every run (slow)"
"\tmtrack: Page tracking based invalidation",
"\tfull: Validate code before every run (slow)",
"\tmman: Invalidate on mmap, mprotect, munmap (deprecated, use mtrack)"
]
},
"TSOEnabled": {
@@ -241,6 +261,21 @@
"Highly likely to break any multithreaded application if disabled."
]
},
"TSOAutoMigration": {
"Type": "bool",
"Default": "true",
"Desc": [
"Automatically enables TSO when shared memory is used.",
"Should work without issues in most cases."
]
},
"X87ReducedPrecision": {
"Type": "bool",
"Default": "false",
"Desc": [
"Emulates X87 floating point using 64-bit precision. This reduces emulation accuracy and may result in rendering bugs."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
+12 -7
View File
@@ -43,8 +43,8 @@ namespace FEXCore::Context {
delete CTX;
}
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader) {
return CTX->InitCore(Loader);
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, uint64_t InitialRIP, uint64_t StackPointer) {
return CTX->InitCore(InitialRIP, StackPointer);
}
void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler) {
@@ -149,13 +149,14 @@ namespace FEXCore::Context {
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
CTX->SignalDelegation = SignalDelegation;
}
void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler) {
CTX->SyscallHandler = Handler;
CTX->SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
@@ -186,11 +187,15 @@ namespace FEXCore::Context {
CTX->WriteFilesWithCode(Writer);
}
void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name) {
return CTX->AddNamedRegion(Base, Length, Offset, Name);
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(FEXCore::Context::Context *CTX, const std::string &Name) {
return CTX->LoadAOTIRCacheEntry(Name);
}
void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length) {
return CTX->RemoveNamedRegion(Base, Length);
void UnloadAOTIRCacheEntry(FEXCore::Context::Context *CTX, IR::AOTIRCacheEntry *Entry) {
return CTX->UnloadAOTIRCacheEntry(Entry);
}
CustomIRResult AddCustomIREntrypoint(FEXCore::Context::Context *CTX, uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
return CTX->AddCustomIREntrypoint(Entrypoint, Handler, Creator, Data);
}
namespace Debug {
+62 -33
View File
@@ -4,6 +4,8 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -34,13 +36,21 @@ class CodeLoader;
class ThunkHandler;
class GdbServer;
namespace CodeSerialize {
class CodeObjectSerializeService;
}
namespace CPU {
class Arm64JITCore;
class X86JITCore;
class InterpreterCore;
class Dispatcher;
}
namespace HLE {
struct SyscallArguments;
class SyscallHandler;
class SourcecodeResolver;
struct SourcecodeMap;
}
}
@@ -67,6 +77,7 @@ namespace FEXCore::Context {
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::CPU::InterpreterCore;
friend class FEXCore::IR::Validation::IRValidation;
struct {
@@ -81,6 +92,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(TSOAutoMigration, TSOAUTOMIGRATION);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(ABINoPF, ABINOPF);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
@@ -97,16 +109,15 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
} Config;
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
IntCallbackReturn InterpreterCallbackReturn;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
uint64_t ThreadID{};
FEXCore::Core::InternalThreadState* ParentThread;
std::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
@@ -118,9 +129,13 @@ namespace FEXCore::Context {
Event PauseWait;
bool Running{};
std::shared_mutex CodeInvalidationMutex;
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
std::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
@@ -135,7 +150,7 @@ namespace FEXCore::Context {
Context();
~Context();
FEXCore::Core::InternalThreadState* InitCore(FEXCore::CodeLoader *Loader);
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
@@ -155,13 +170,20 @@ namespace FEXCore::Context {
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
static void RemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Must be called from owning thread
static void RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState
static void RemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveCodeEntry(Frame->Thread, GuestRIP);
// Must be called from owning thread
static void RemoveThreadCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveThreadCodeEntry(Frame->Thread, GuestRIP);
}
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
@@ -171,21 +193,19 @@ namespace FEXCore::Context {
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
// User's responsibility to deallocate this.
FEXCore::IR::RegisterAllocationData* RAData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo);
struct CompileCodeResult {
void* CompiledCode;
FEXCore::IR::IRListView* IRData;
FEXCore::Core::DebugData* DebugData;
// User's responsibility to deallocate this.
FEXCore::IR::RegisterAllocationData* RAData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
bool GeneratedIR;
uint64_t StartAddr;
uint64_t Length;
@@ -196,21 +216,9 @@ namespace FEXCore::Context {
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
* @param CompileThread Is this for the compile service or not?
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
* This is exposed because the CompileService needs to initialize compilers while copying data from
* the paired InternalThreadState that it is compiling code for
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread);
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
@@ -233,7 +241,7 @@ namespace FEXCore::Context {
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
/**
* @brief Initializes the TLS data for a thread
* @brief Initializes TID, PID and TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
@@ -269,8 +277,8 @@ namespace FEXCore::Context {
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry);
FEXCore::JITSymbols Symbols;
@@ -297,8 +305,15 @@ namespace FEXCore::Context {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
void MarkMemoryShared();
bool IsTSOEnabled() { return (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled; }
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread);
private:
/**
@@ -310,22 +325,36 @@ namespace FEXCore::Context {
*/
void InitializeThreadData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void WaitForIdleWithTimeout();
void NotifyPause();
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length);
FEXCore::CodeLoader *LocalLoader{};
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
// Entry Cache
uint64_t StartingRIP;
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
IR::AOTIRCaptureCache IRCaptureCache;
std::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::unordered_map<uint64_t, std::tuple<std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)>, void *, void *>> CustomIRHandlers;
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
FEXCore::CPU::DispatcherConfig DispatcherConfig;
};
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args);
+159 -112
View File
@@ -1,14 +1,14 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <aarch64/cpu-aarch64.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <atomic>
#include <stdint.h>
#include <signal.h>
#include "aarch64/cpu-aarch64.h"
#include <csignal>
#include <cstdint>
namespace FEXCore::ArchHelpers::Arm64 {
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock, TYPE_HAS_SPLIT_LOCKS);
@@ -513,7 +513,8 @@ uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
//Only 32-bit pairs
for(int i = 1; i < 10; i++) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST) {
ExpectedReg1 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::CCMP_MASK) == FEXCore::ArchHelpers::Arm64::CCMP_INST) {
ExpectedReg2 = GetRmReg(NextInstr);
@@ -1579,7 +1580,7 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1592,7 +1593,7 @@ bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
uint32_t ResultReg = Instr & 0b11111;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = mcontext->regs[AddressReg];
uint64_t Addr = mcontext->regs[AddressReg] + Offset;
if (Size == 2) {
auto Res = DoLoad16(Addr);
@@ -1622,7 +1623,7 @@ bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr) {
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1635,7 +1636,7 @@ bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr) {
uint32_t DataReg = Instr & 0x1F;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = mcontext->regs[AddressReg];
uint64_t Addr = mcontext->regs[AddressReg] + Offset;
constexpr bool DoRetry = false;
if (Size == 2) {
@@ -1745,7 +1746,8 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
#endif
DesiredReg = GetRdReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST) {
ExpectedReg = GetRmReg(NextInstr);
}
}
@@ -1828,11 +1830,13 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
// Scan forward at most five instructions to find our instructions
for (size_t i = 1; i < 6; ++i) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_INST) {
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_SHIFT_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_ADD;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_INST) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_SHIFT_INST) {
uint32_t RnReg = GetRnReg(NextInstr);
if (RnReg == REGISTER_MASK) {
// Zero reg means neg
@@ -1843,21 +1847,34 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
return HandleCAS_NoAtomics(_ucontext, _info); //ARMv8.0 CAS
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST ) {
return HandleCAS_NoAtomics(_ucontext, _info); //ARMv8.0 CAS
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::AND_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_AND;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::BIC_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_BIC;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::OR_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_OR;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ORN_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_ORN;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::EOR_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_EOR;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::EON_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_EON;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
@@ -1888,40 +1905,53 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
uint32_t Size = 1 << (Instr >> 30);
constexpr bool DoRetry = true;
auto NOPExpected = []<typename AtomicType>(AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto BICDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & ~Desired;
};
auto ORDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto ORNDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | ~Desired;
};
auto EORDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto EONDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ ~Desired;
};
auto NEGDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
if (Size == 2) {
using AtomicType = uint16_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -1937,12 +1967,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -1966,38 +2005,6 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
else if (Size == 4) {
using AtomicType = uint32_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -2013,12 +2020,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -2042,38 +2058,6 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
else if (Size == 8) {
using AtomicType = uint64_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -2089,12 +2073,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -2140,7 +2133,7 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
if ((Instr & 0x3F'FF'FC'00) == 0x08'DF'FC'00 || // LDAR*
(Instr & 0x3F'FF'FC'00) == 0x38'BF'C0'00) { // LDAPR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr)) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr, 0)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
@@ -2164,7 +2157,7 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
}
else if ( (Instr & 0x3F'FF'FC'00) == 0x08'9F'FC'00) { // STLR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr)) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr, 0)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
@@ -2186,6 +2179,60 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & RCPC2_MASK) == LDAPUR_INST) { // LDAPUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr, Offset)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAPUR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t LDUR = 0b0011'1000'0100'0000'0000'0000'0000'0000;
LDUR |= Size << 30;
LDUR |= AddrReg << 5;
LDUR |= DataReg;
LDUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB;
PC[0] = LDUR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & RCPC2_MASK) == STLUR_INST) { // STLUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr, Offset)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDLUR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t STUR = 0b0011'1000'0000'0000'0000'0000'0000'0000;
STUR |= Size << 30;
STUR |= AddrReg << 5;
STUR |= DataReg;
STUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB;
PC[0] = STUR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXP_MASK) == FEXCore::ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
//Should be compare and swap pair only. LDAXP not used elsewhere
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleCASPAL_ARMv8(ucontext, info, Instr);
+22 -9
View File
@@ -12,6 +12,10 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t ATOMIC_MEM_MASK = 0x3B200C00;
constexpr uint32_t ATOMIC_MEM_INST = 0x38200000;
constexpr uint32_t RCPC2_MASK = 0x3F'E0'0C'00;
constexpr uint32_t LDAPUR_INST = 0x19'40'00'00;
constexpr uint32_t STLUR_INST = 0x19'00'00'00;
constexpr uint32_t LDAXP_MASK = 0xBF'FF'80'00;
constexpr uint32_t LDAXP_INST = 0x88'7F'80'00;
@@ -27,13 +31,19 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t CBNZ_MASK = 0x7F'00'00'00;
constexpr uint32_t CBNZ_INST = 0x35'00'00'00;
constexpr uint32_t ALU_OP_MASK = 0x7F'00'00'00;
constexpr uint32_t ADD_INST = 0x0B'00'00'00;
constexpr uint32_t SUB_INST = 0x4B'00'00'00;
constexpr uint32_t CMP_INST = 0x6B'00'00'00;
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t ALU_OP_MASK = 0x7F'20'00'00;
constexpr uint32_t ADD_INST = 0x0B'00'00'00;
constexpr uint32_t SUB_INST = 0x4B'00'00'00;
constexpr uint32_t ADD_SHIFT_INST = 0x0B'20'00'00;
constexpr uint32_t SUB_SHIFT_INST = 0x4B'20'00'00;
constexpr uint32_t CMP_INST = 0x6B'00'00'00;
constexpr uint32_t CMP_SHIFT_INST = 0x6B'20'00'00;
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t BIC_INST = 0x0A'20'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t ORN_INST = 0x2A'20'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t EON_INST = 0x4A'20'00'00;
constexpr uint32_t CCMP_MASK = 0x7F'E0'0C'10;
constexpr uint32_t CCMP_INST = 0x7A'40'00'00;
@@ -46,8 +56,11 @@ namespace FEXCore::ArchHelpers::Arm64 {
TYPE_ADD,
TYPE_SUB,
TYPE_AND,
TYPE_BIC,
TYPE_OR,
TYPE_ORN,
TYPE_EOR,
TYPE_EON,
TYPE_NEG, // This is just a sub with zero. Need to know the differences
};
@@ -83,8 +96,8 @@ namespace FEXCore::ArchHelpers::Arm64 {
return (Instr >> RM_OFFSET) & REGISTER_MASK;
}
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset);
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset);
bool HandleAtomicLoad128(void *_ucontext, void *_info, uint32_t Instr);
uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
@@ -1,22 +1,28 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include "aarch64/cpu-aarch64.h"
#include "cpu-features.h"
#include "aarch64/instructions-aarch64.h"
#include "utils-vixl.h"
#include <aarch64/cpu-aarch64.h>
#include <aarch64/instructions-aarch64.h>
#include <cpu-features.h>
#include <utils-vixl.h>
#include <array>
#include <tuple>
#include <utility>
namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size)
: vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode)
, EmitterCTX {ctx} {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
@@ -42,12 +48,57 @@ void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant,
}
int NumMoves = 1;
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
int RequiredMoveSegments{};
// Count the number of move segments
// We only want to use ADRP+ADD if we have more than 1 segment
for (size_t i = 0; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
++NumMoves;
if (Part != 0) {
++RequiredMoveSegments;
}
}
// ADRP+ADD is specifically optimized in hardware
// Check if we can use this
auto PC = GetCursorAddress<uint64_t>();
// PC aligned to page
uint64_t AlignedPC = PC & ~0xFFFULL;
// Offset from aligned PC
int64_t AlignedOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(AlignedPC);
// If the aligned offset is within the 4GB window then we can use ADRP+ADD
// and the number of move segments more than 1
if (RequiredMoveSegments > 1 && vixl::IsInt32(AlignedOffset)) {
// If this is 4k page aligned then we only need ADRP
if ((AlignedOffset & 0xFFF) == 0) {
adrp(Reg, AlignedOffset >> 12);
}
else {
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
if (vixl::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
}
else {
// Need to use ADRP + ADD
adrp(Reg, AlignedOffset >> 12);
add(Reg, Reg, Constant & 0xFFF);
NumMoves = 2;
}
}
}
else {
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
++NumMoves;
}
}
}
@@ -260,23 +311,10 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
add(sp, sp, SPOffset);
}
void Arm64Emitter::ResetStack() {
if (SpillSlots == 0)
return;
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
nop();
nop();
}
}
@@ -1,15 +1,19 @@
#pragma once
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "platform-vixl.h"
#include "FEXCore/Config/Config.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/constants-aarch64.h>
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <platform-vixl.h>
#include <FEXCore/Config/Config.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <cstddef>
#include <cstdint>
#include <utility>
namespace FEXCore::CPU {
@@ -60,6 +64,7 @@ class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
FEXCore::Context::Context *EmitterCTX;
vixl::aarch64::CPU CPU;
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
@@ -77,11 +82,8 @@ protected:
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void ResetStack();
void Align16B();
uint32_t SpillSlots{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
+87
View File
@@ -0,0 +1,87 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
namespace FEXCore {
namespace CPU {
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
}
return CurrentCodeBuffer;
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::mmap(nullptr, Buffer.Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (ThreadState->CTX->Config.GlobalJITNaming()) {
ThreadState->CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto &Buffer: CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
}
}
return false;
}
}
}
+24 -4
View File
@@ -416,7 +416,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
Res.ecx =
(1 << 0) | // SSE3
(0 << 1) | // PCLMULQDQ
(1 << 1) | // PCLMULQDQ
(1 << 2) | // DS area supports 64bit layout
(1 << 3) | // MWait
(0 << 4) | // DS-CPL
@@ -446,7 +446,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(SUPPORTS_AVX << 28) | // AVX
(0 << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(0 << 31); // Hypervisor always returns zero
(1 << 31); // Hypervisor always returns one
Res.edx =
(1 << 0) | // FPU
@@ -658,7 +658,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 28) | // Reserved
(0 << 29) | // SHA instructions
(1 << 29) | // SHA instructions
(0 << 30) | // Reserved
(0 << 31); // Reserved
@@ -826,7 +826,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
// CPUID documentation information:
// 4000_0000h - 4FFF_FFFFh - No existing or future CPU will return information in this range
// Reserved entirely for VMs to do whatever they want.
Res.eax = 0x40000000;
Res.eax = 0x40000001;
// EBX, EDX, ECX become the hypervisor ID signature
constexpr static char HypervisorID[12] = "FEXIFEXIEMU";
@@ -834,6 +834,25 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
#ifdef _M_X86_64
// EAX[3:0] = 1 = x86_64 host architecture
Res.eax |= 0b0001;
#elif defined(_M_ARM_64)
// EAX[3:0] = 2 = AArch64 host architecture
Res.eax |= 0b0010;
#else
// EAX[3:0] = 0 = Unknown architecture
#endif
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
@@ -1228,6 +1247,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#endif
// Hypervisor CPUID information leaf
RegisterFunction(0x4000'0000, &CPUIDEmu::Function_4000_0000h);
RegisterFunction(0x4000'0001, &CPUIDEmu::Function_4000_0001h);
// Largest extended function number
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
+1
View File
@@ -80,6 +80,7 @@ private:
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf);
@@ -1,165 +0,0 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include "FEXCore/HLE/Linux/ThreadManagement.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <memory>
#include <pthread.h>
#include <stdio.h>
namespace FEXCore {
static void* ThreadHandler(void *Arg) {
FEXCore::CompileService *This = reinterpret_cast<FEXCore::CompileService*>(Arg);
This->ExecutionThread();
return nullptr;
}
CompileService::CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ParentThread {Thread} {
CompileThreadData = std::make_unique<FEXCore::Core::InternalThreadState>();
CompileThreadData->IsCompileService = true;
// We need a compiler for this work thread
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CompileService::Initialize() {
// Share CompileService which = this
CompileThreadData->CompileService = ParentThread->CompileService;
}
void CompileService::Shutdown() {
ShuttingDown = true;
// Kick the working thread
StartWork.NotifyAll();
WorkerThread->join(nullptr);
}
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread) {
// On cache clear we need to spin down the execution thread to ensure it isn't trying to give us more work items
if (CompileMutex.try_lock()) {
// We can only clear these things if we pulled the compile mutex
// Grab the work queue and clear it
// We don't need to grab the queue mutex since this thread will no longer receive any work events
// Threads are bounded 1:1
while (!WorkQueue.empty()) {
WorkQueue.pop();
}
// Go through the garbage collection array and clear it
// It's safe to clear things that aren't marked safe since we are clearing cache
GCArray.clear();
LOGMAN_THROW_A_FMT(CompileThreadData->LocalIRCache.empty(), "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
// Clear the inverse cache of what is calling us from the Context ClearCache routine
auto SelectedThread = Thread->IsCompileService ? ParentThread : Thread;
SelectedThread->LookupCache->ClearCache();
SelectedThread->CPUBackend->ClearCache();
}
CompileService::WorkItem *CompileService::CompileCode(uint64_t RIP) {
WorkItem* ResultItem = nullptr;
{
// Tell the worker thread to compile code for us
auto Item = std::make_unique<WorkItem>();
Item->RIP = RIP;
// Fill the threads work queue
std::scoped_lock lk(QueueMutex);
ResultItem = WorkQueue.emplace(std::move(Item)).get();
}
// Notify the thread that it has more work
StartWork.NotifyAll();
return ResultItem;
}
void CompileService::ExecutionThread() {
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
pthread_setname_np(pthread_self(), ThreadName);
while (true) {
// Wait for work
StartWork.Wait();
if (ShuttingDown.load()) {
break;
}
std::scoped_lock lk(CompileMutex);
size_t WorkItems{};
do {
// Grab a work item
std::unique_ptr<WorkItem> Item{};
{
std::scoped_lock lk(QueueMutex);
WorkItems = WorkQueue.size();
if (WorkItems != 0) {
Item = std::move(WorkQueue.front());
WorkQueue.pop();
}
}
// If we had a work item then work on it
if (Item) {
// Make sure it's not in lookup cache by accident
LOGMAN_THROW_A_FMT(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->CurrentFrame->State.rip = Item->RIP;
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
LOGMAN_THROW_A_FMT(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
ERROR_AND_DIE_FMT("Couldn't compile code for thread at RIP: 0x{:x}", Item->RIP);
}
Item->CodePtr = CodePtr;
Item->IRList = IRList;
Item->DebugData = DebugData;
Item->RAData = RAData;
Item->StartAddr = StartAddr;
Item->Length = Length;
auto& GCItem = GCArray.emplace_back(std::move(Item));
GCItem->ServiceWorkDone.NotifyAll();
}
} while (WorkItems != 0);
// Clean up any safe entries in our GC array if we have any.
std::erase_if(GCArray, [](const auto& Entry) {
return Entry->SafeToClear.load(std::memory_order_relaxed);
});
}
}
}
-68
View File
@@ -1,68 +0,0 @@
#pragma once
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <memory>
#include <mutex>
#include <queue>
#include <stdint.h>
#include <vector>
namespace FEXCore {
namespace Context {
struct Context;
}
namespace IR {
class IRListView;
class RegisterAllocationData;
};
class CompileService final {
public:
CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
void Initialize();
void Shutdown();
struct WorkItem {
// Incoming
uint64_t RIP{};
// Outgoing
void *CodePtr{};
FEXCore::IR::IRListView *IRList{};
FEXCore::IR::RegisterAllocationData *RAData{};
FEXCore::Core::DebugData *DebugData{};
uint64_t StartAddr;
uint64_t Length;
// Communication
Event ServiceWorkDone{};
std::atomic_bool SafeToClear{};
};
WorkItem *CompileCode(uint64_t RIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread);
// Public for threading
void ExecutionThread();
bool IsAddressInJITCode(uint64_t Address) const {
return CompileThreadData->CPUBackend->IsAddressInJITCode(Address, false, false);
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::unique_ptr<FEXCore::Core::InternalThreadState> CompileThreadData;
std::mutex QueueMutex{};
std::mutex CompileMutex{};
std::queue<std::unique_ptr<WorkItem>> WorkQueue{};
std::vector<std::unique_ptr<WorkItem>> GCArray{};
Event StartWork{};
std::atomic_bool ShuttingDown{false};
};
}
File diff suppressed because it is too large. Load diff
@@ -16,20 +16,23 @@
#include <array>
#include <bit>
#include <cmath>
#include <cstddef>
#include <cstdint>
#include <memory>
#include <stddef.h>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/constants-aarch64.h>
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <code-buffer-vixl.h>
#include <platform-vixl.h>
#include <sys/syscall.h>
#include <unistd.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
MemOperand(STATE, offsetof(FEXCore::Core::STATE_TYPE, FIELD))
namespace FEXCore::CPU {
using namespace vixl;
@@ -38,12 +41,12 @@ using namespace vixl::aarch64;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
, config(config) {
SetAllowAssembler(true);
DispatchPtr = GetCursorAddress<CPUBackend::AsmDispatch>();
DispatchPtr = GetCursorAddress<AsmDispatch>();
// while (true) {
// Ptr = FindBlock(RIP)
@@ -53,12 +56,9 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Ptr();
// }
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
Literal l_ExitFunctionLink {config.ExitFunctionLink};
Literal l_ExitFunctionLinkThis {config.ExitFunctionLinkThis};
// Push all the register we need to save
PushCalleeSavedRegisters();
@@ -71,11 +71,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
str(x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled) {
if (config.StaticRegisterAllocation) {
FillStaticRegs();
}
@@ -92,11 +92,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
ldr(x2, STATE_PTR(CpuStateFrame, State.rip));
auto RipReg = x2;
// L1 Cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -104,21 +104,17 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
b(&CallBlock);
}
br(x3);
// L1C check failed, do a full lookup
bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, &l_PagePtr);
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, VirtualMemorySize - 1);
}
@@ -157,44 +153,21 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
// Jump to the block
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
bind(&CallBlock);
mov(x0, STATE);
blr(x3);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, &l_CTX);
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
} else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
br(x3);
}
}
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
@@ -209,7 +182,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
constexpr bool SignalSafeCompile = true;
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
if (SignalSafeCompile) {
@@ -231,11 +204,10 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
svc(0);
}
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
mov(x0, STATE);
mov(x1, lr);
ldr(x3, &l_ExitFunctionLink);
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
blr(x3);
if (SignalSafeCompile) {
@@ -256,7 +228,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
mov(x0, x4);
}
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
br(x0);
}
@@ -265,7 +237,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
bind(&NoBlock);
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
if (SignalSafeCompile) {
@@ -312,7 +284,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
add(sp, sp, 16);
}
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
b(&LoopTop);
@@ -331,7 +303,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
hlt(0);
@@ -342,18 +314,17 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
LoadConstant(x0, reinterpret_cast<uint64_t>(&SynchronousFaultData));
LoadConstant(w1, 1);
strb(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)));
strb(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException));
LoadConstant(w1, X86State::X86_TRAPNO_OF);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo));
LoadConstant(w1, 0x80);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code));
LoadConstant(x1, 0);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code));
// hlt/udf = SIGILL
// brk = SIGTRAP
@@ -364,7 +335,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
bind(&ThreadPauseHandler);
@@ -399,7 +370,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// On return to the thunk, the thunk can get whatever its return value is from the thread context depending on ABI handling on its end
// When the thunk itself returns, it'll do its regular return logic there
// void ReentrantCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
CallbackPtr = GetCursorAddress<CPUBackend::JITCallback>();
CallbackPtr = GetCursorAddress<JITCallback>();
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
@@ -408,40 +379,112 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
mov(STATE, x0);
// Make sure to adjust the refcounter so we don't clear the cache now
LoadConstant(x0, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
ldr(w2, MemOperand(x0));
ldr(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
add(w2, w2, 1);
str(w2, MemOperand(x0));
str(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
ldr(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
str(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
str(x1, STATE_PTR(CpuStateFrame, State.rip));
// load static regs
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
place(&l_PagePtr);
{
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
place(&l_ExitFunctionLink);
place(&l_ExitFunctionLinkThis);
FinalizeCode();
@@ -457,51 +500,114 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Pointers.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext, uint32_t IgnoreMask) {
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline, destination buffer is set before use
static thread_local vixl::aarch64::Assembler emit((uint8_t*)&emit, 1);
size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxGDBPauseCheckSize);
vixl::CodeBufferCheckScope scope(&emit, MaxGDBPauseCheckSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(FEXCore::Context::Context::Config.RunningMode) == 4, "This is expected to be size of 4");
emit.ldr(x0, STATE_PTR(CpuStateFrame, Thread)); // Get thread
emit.ldr(x0, MemOperand(x0, offsetof(FEXCore::Core::InternalThreadState, CTX))); // Get Context
emit.ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then we don't need to stop
emit.cbz(w0, &RunBlock);
{
Literal l_GuestRIP {GuestRIP};
// Make sure RIP is syncronized to the context
emit.ldr(x0, &l_GuestRIP);
emit.str(x0, STATE_PTR(CpuStateFrame, State.rip));
// Stop the thread
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.br(x0);
emit.place(&l_GuestRIP);
}
emit.bind(&RunBlock);
emit.FinalizeCode();
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
return UsedBytes;
}
size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "GenerateInterpreterTrampoline dispatcher does not support SRA");
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
vixl::CodeBufferCheckScope scope(&emit, MaxInterpreterTrampolineSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label InlineIRData;
emit.mov(x0, STATE);
emit.adr(x1, &InlineIRData);
emit.ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.blr(x3);
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.br(x0);
emit.bind(&InlineIRData);
emit.FinalizeCode();
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
return UsedBytes;
}
void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {
for(int i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
// Skip this one, it's already spilled
continue;
}
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
for(int i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&ThreadState->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
memcpy(&Thread->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
}
}
#ifdef _M_ARM_64
void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Dispatcher = std::make_unique<Arm64Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
auto &AArch64 = Thread->CurrentFrame->Pointers.AArch64;
AArch64.LUDIVHandler = LUDIVHandlerAddress;
AArch64.LDIVHandler = LDIVHandlerAddress;
AArch64.LUREMHandler = LUREMHandlerAddress;
AArch64.LREMHandler = LREMHandlerAddress;
}
}
#endif
std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<Arm64Dispatcher>(CTX, Config);
}
}
@@ -15,10 +15,21 @@ namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
protected:
void SpillSRA(void *ucontext, uint32_t IgnoreMask) override;
void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) override;
private:
// Long division helpers
uint64_t LUDIVHandlerAddress{};
uint64_t LDIVHandlerAddress{};
uint64_t LUREMHandlerAddress{};
uint64_t LREMHandlerAddress{};
DispatcherConfig config;
};
}
@@ -1,8 +1,8 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/CompileService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
@@ -40,7 +40,7 @@ void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuS
ctx->IdleWaitCV.notify_all();
}
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, void *ucontext) {
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -65,7 +65,7 @@ ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, vo
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, ThreadState->CurrentFrame, sizeof(FEXCore::Core::CPUState));
memcpy(&Context->GuestState, Thread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
@@ -82,13 +82,13 @@ ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, vo
Context->SigInfoLocation = 0;
// Store fault to top status and then reset it
Context->FaultToTopAndGeneratedException = SynchronousFaultData.FaultToTopAndGeneratedException;
SynchronousFaultData.FaultToTopAndGeneratedException = false;
Context->FaultToTopAndGeneratedException = Thread->CurrentFrame->SynchronousFaultData.FaultToTopAndGeneratedException;
Thread->CurrentFrame->SynchronousFaultData.FaultToTopAndGeneratedException = false;
return Context;
}
void Dispatcher::RestoreThreadState(void *ucontext) {
void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext) {
uint64_t OldSP{};
if (CTX->Config.Core() == FEXCore::Config::CONFIG_IRJIT) {
OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -103,13 +103,13 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(ThreadState->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
memcpy(Thread->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
if (Context->UContextLocation) {
auto Frame = ThreadState->CurrentFrame;
auto Frame = Thread->CurrentFrame;
if (Context->Flags &ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT) {
// XXX: Unsupported since it needs state reconstruction
@@ -133,7 +133,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP];
// XXX: Full context setting
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
@@ -191,7 +191,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
// XXX: Full context setting
// First 32-bytes of flags is EFLAGS broken out
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
@@ -219,7 +219,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&Frame->State.mm[i], &fpstate->_st[i], 10);
}
@@ -274,14 +274,14 @@ static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
return 0;
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
auto ContextBackup = StoreThreadState(Signal, ucontext);
bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
auto ContextBackup = StoreThreadState(Thread, Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
auto Frame = Thread->CurrentFrame;
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
++Thread->CurrentFrame->SignalHandlerRefCounter;
uint64_t OldPC = ArchHelpers::Context::GetPc(ucontext);
// Set the new PC
@@ -300,7 +300,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// We are going to be returning to the top of the dispatcher which will fill again
// Otherwise we might load garbage
if (SRAEnabled) {
if (IsAddressInJITCode(OldPC, false)) {
if (Thread->CPUBackend->IsAddressInCodeBuffer(OldPC)) {
uint32_t IgnoreMask{};
#ifdef _M_ARM_64
if (Frame->InSyscallInfo != 0) {
@@ -324,11 +324,11 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#endif
// We are in jit, SRA must be spilled
SpillSRA(ucontext, IgnoreMask);
SpillSRA(Thread, ucontext, IgnoreMask);
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT;
} else {
if (!IsAddressInJITCode(OldPC, true)) {
if (!IsAddressInDispatcher(OldPC)) {
// This is likely to cause issues but in some cases it isn't fatal
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
@@ -410,11 +410,11 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
*guest_siginfo = *HostSigInfo;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = SynchronousFaultData.err_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = Frame->SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = Frame->SynchronousFaultData.err_code;
// Overwrite si_code
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_siginfo->si_code = Thread->CurrentFrame->SynchronousFaultData.si_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
@@ -501,9 +501,9 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = SynchronousFaultData.err_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = Frame->SynchronousFaultData.TrapNo;
guest_siginfo->si_code = Frame->SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = Frame->SynchronousFaultData.err_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
@@ -529,7 +529,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#undef COPY_REG
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&fpstate->_st[i], &Frame->State.mm[i], 10);
}
@@ -634,44 +634,44 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
return true;
}
bool Dispatcher::HandleSIGILL(int Signal, void *info, void *ucontext) {
bool Dispatcher::HandleSIGILL(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) {
if (ArchHelpers::Context::GetPc(ucontext) == SignalHandlerReturnAddress) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
return true;
}
if (ArchHelpers::Context::GetPc(ucontext) == PauseReturnInstruction) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
return true;
}
return false;
}
bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = ThreadState->SignalReason.load();
auto Frame = ThreadState->CurrentFrame;
bool Dispatcher::HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = Thread->SignalReason.load();
auto Frame = Thread->CurrentFrame;
if (SignalReason == FEXCore::Core::SignalEvent::Pause) {
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
StoreThreadState(Thread, Signal, ucontext);
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
@@ -682,9 +682,9 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
++Thread->CurrentFrame->SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -695,16 +695,16 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
ArchHelpers::Context::SetSp(ucontext, Frame->ReturningStackLocation);
// Our ref counting doesn't matter anymore
SignalHandlerRefCounter = 0;
Thread->CurrentFrame->SignalHandlerRefCounter = 0;
// Set the new PC
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
@@ -713,24 +713,24 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
// We need to be a little bit careful here
// If we were already paused (due to GDB) and we are immediately stopping (due to gdb kill)
// Then we need to ensure we don't double decrement our idle thread counter
if (ThreadState->RunningEvents.ThreadSleeping) {
if (Thread->RunningEvents.ThreadSleeping) {
// If the thread was sleeping then its idle counter was decremented
// Reincrement it here to not break logic
++ThreadState->CTX->IdleWaitRefCount;
++Thread->CTX->IdleWaitRefCount;
}
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::Return) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -749,31 +749,4 @@ uint64_t Dispatcher::GetCompileBlockPtr() {
return CompileBlockPtr.Data;
}
void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
for (auto iter = CodeBuffers.begin(); iter != CodeBuffers.end(); ++iter) {
auto [start, end] = *iter;
if (start == reinterpret_cast<uint64_t>(start_to_remove)) {
CodeBuffers.erase(iter);
return;
}
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher, bool IncludeCompileService) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher && IsAddressInDispatcher(Address)) {
return true;
}
if (IncludeCompileService && ThreadState->CompileService && ThreadState->CompileService->IsAddressInJITCode(Address)) {
return true;
}
return false;
}
}
@@ -1,8 +1,6 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <cstdint>
@@ -21,22 +19,20 @@ struct CpuStateFrame;
struct InternalThreadState;
}
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::CPU {
struct DispatcherConfig {
bool ExecuteBlocksWithCall = false;
uintptr_t ExitFunctionLink = 0;
uintptr_t ExitFunctionLinkThis = 0;
bool StaticRegisterAssignment = false;
bool StaticRegisterAllocation = false;
};
class Dispatcher {
public:
virtual ~Dispatcher() = default;
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
/**
* @name Dispatch Helper functions
* @{ */
@@ -50,59 +46,66 @@ public:
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint64_t IntCallbackReturnAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
} SynchronousFaultData;
uint64_t Start{};
uint64_t End{};
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
bool HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
void RegisterCodeBuffer(uint8_t* start, size_t size) {
CodeBuffers.emplace_back(reinterpret_cast<uint64_t>(start),
reinterpret_cast<uint64_t>(start + size));
}
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
protected:
Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ThreadState {Thread} {}
virtual void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) = 0;
ArchHelpers::Context::ContextBackup* StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
// These are across all arches for now
static constexpr size_t MaxGDBPauseCheckSize = 128;
static constexpr size_t MaxInterpreterTrampolineSize = 128;
virtual size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) = 0;
virtual size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) = 0;
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
protected:
Dispatcher(FEXCore::Context::Context *ctx)
: CTX {ctx}
{}
ArchHelpers::Context::ContextBackup* StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext);
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext);
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext, uint32_t IgnoreMask) {}
virtual void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
private:
std::vector<std::tuple<uint64_t, uint64_t>> CodeBuffers; // Start, End
using AsmDispatch = void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
};
}
@@ -18,21 +18,26 @@
#include <stddef.h>
#include <stdint.h>
#include <sys/mman.h>
#include "xbyak/xbyak.h"
#include <xbyak/xbyak.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
[STATE + offsetof(FEXCore::Core::STATE_TYPE, FIELD)]
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread)
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: Dispatcher(ctx)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE,
FEXCore::Allocator::mmap(nullptr, MAX_DISPATCHER_CODE_SIZE, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0),
nullptr) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "X86 dispatcher does not support SRA");
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
DispatchPtr = getCurr<AsmDispatch>();
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
@@ -78,11 +83,10 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)], rsp);
mov(qword STATE_PTR(CpuStateFrame, ReturningStackLocation), rsp);
Label LoopTop;
Label FullLookup;
Label CallBlock;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
@@ -92,29 +96,24 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
{
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
mov(rdx, qword STATE_PTR(CPUState, rip));
// L1 Cache
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
cmp(qword[r13 + rax + offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)], rdx);
jne(FullLookup);
if (!config.ExecuteBlocksWithCall) {
jmp(qword[r13 + rax + 0]);
} else {
mov(rax, qword[r13 + rax + 0]);
jmp(CallBlock);
}
jmp(qword[r13 + rax + offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)]);
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Full lookup
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
mov(rax, rdx);
mov(rbx, VirtualMemorySize - 1);
and_(rax, rbx);
@@ -143,7 +142,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
je(NoBlock);
// Update L1
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
@@ -151,30 +150,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(qword[r13 + rcx*8 + 0], rax);
// Real block if we made it here
if (!config.ExecuteBlocksWithCall) {
jmp(rax);
} else {
L(CallBlock);
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
jmp(rax);
}
{
@@ -193,10 +169,37 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
ret();
}
constexpr bool SignalSafeCompile = true;
// Block creation
{
L(NoBlock);
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rdx
mov(r9, rdx);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rdx, r9);
}
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
@@ -204,20 +207,84 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
call(rax);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rdx
mov(r9, rdx);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
// Bring stack back
add(rsp, 16);
mov(rdx, r9);
}
// rdx already contains RIP here
jmp(LoopTop);
}
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
// {rdi, rsi, rdx}
mov(rdi, config.ExitFunctionLinkThis);
mov(rsi, STATE);
mov(rdx, rax); // rax is set at the block end
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
mov(rax, config.ExitFunctionLink);
call(rax);
jmp(rax);
// Backup rax
mov(r9, rax);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rax, r9);
}
// {rdi, rsi}
mov(rdi, STATE);
mov(rsi, rax); // rax is set at the block end
call(qword STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rax
mov(r9, rax);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
// Bring stack back
add(rsp, 16);
jmp(r9);
}
else {
jmp(rax);
}
}
{
@@ -237,7 +304,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
CallbackPtr = getCurr<JITCallback>();
push(rbx);
push(rbp);
@@ -252,7 +319,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
add(qword STATE_PTR(CpuStateFrame, SignalHandlerRefCounter), 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
@@ -260,12 +327,12 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])]);
sub(qword STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]), 16);
mov(rbx, qword STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], rsi);
mov(qword STATE_PTR(CpuStateFrame, State.rip), rsi);
// Back to the loop top now
jmp(LoopTop);
@@ -292,18 +359,17 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
mov(rax, reinterpret_cast<uint64_t>(&SynchronousFaultData));
add(byte [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)], 1);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)], X86State::X86_TRAPNO_OF);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)], 0);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)], 0x80);
add(byte STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException), 1);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo), X86State::X86_TRAPNO_OF);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code), 0);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code), 0x80);
hlt();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
IntCallbackReturnAddress = getCurr<uint64_t>();
// using CallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
@@ -337,39 +403,91 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(Start), End-Start);
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
}
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandler = ThreadStopHandlerAddress;
Pointers.ThreadPauseHandler = ThreadPauseHandlerAddress;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline
static thread_local Xbyak::CodeGenerator emit(1, &emit); // actual emit target set with setNewBuffer
size_t X86Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
using namespace Xbyak;
using namespace Xbyak::util;
emit.setNewBuffer(CodeBuffer, MaxGDBPauseCheckSize);
Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
emit.mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then we don't need to stop
emit.cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
emit.je(RunBlock);
{
// Make sure RIP is syncronized to the context
emit.mov(rax, GuestRIP);
emit.mov(qword STATE_PTR(CpuStateFrame, State.rip), rax);
// Stop the thread
emit.mov(rax, qword STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.jmp(rax);
}
emit.L(RunBlock);
emit.ready();
return emit.getSize();
}
size_t X86Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
using namespace Xbyak;
using namespace Xbyak::util;
emit.setNewBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
Label InlineIRData;
emit.mov(rdi, STATE);
emit.lea(rsi, ptr[rip + InlineIRData]);
emit.call(qword STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.jmp(qword STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.L(InlineIRData);
emit.ready();
return emit.getSize();
}
X86Dispatcher::~X86Dispatcher() {
FEXCore::Allocator::munmap(top_, MAX_DISPATCHER_CODE_SIZE);
}
#ifdef _M_X86_64
void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddress;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddress;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Dispatcher = std::make_unique<X86Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
(uintptr_t&)Interpreter.CallbackReturn = IntCallbackReturnAddress;
}
}
#endif
std::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<X86Dispatcher>(CTX, Config);
}
}
@@ -17,7 +17,10 @@ namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
virtual ~X86Dispatcher() override;
};
+37 -10
View File
@@ -19,6 +19,7 @@ $end_info$
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <set>
#include <sys/mman.h>
@@ -177,17 +178,12 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
Decoder::Decoder(FEXCore::Context::Context *ctx)
: CTX {ctx}
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN } {
// Using mmap is a start-up time optimization
// Take advantage of page faulting to reduce startup time for minimal runtime cost
DecodedBuffer =
reinterpret_cast<FEXCore::X86Tables::DecodedInst *>(
FEXCore::Allocator::mmap(0, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize,
PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN }
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
}
Decoder::~Decoder() {
FEXCore::Allocator::munmap(DecodedBuffer, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize);
PoolObject.UnclaimBuffer();
}
uint8_t Decoder::ReadByte() {
@@ -1140,7 +1136,7 @@ const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage) {
Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
@@ -1148,6 +1144,7 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
DecodedSize = 0;
MaxCondBranchForward = 0;
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// XXX: Load symbol data
SymbolAvailable = false;
@@ -1169,6 +1166,13 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
// Entry is a jump target
BlocksToDecode.emplace(PC);
uint64_t CurrentCodePage = PC & FHU::FEX_PAGE_MASK;
std::set<uint64_t> CodePages = { CurrentCodePage };
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
while (!BlocksToDecode.empty()) {
auto BlockDecodeIt = BlocksToDecode.begin();
uint64_t RIPToDecode = *BlockDecodeIt;
@@ -1185,10 +1189,33 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
while (1) {
// MAX_INST_SIZE assumes worst case
auto OpMinAddress = RIPToDecode + PCOffset;
auto OpMaxAddress = OpMinAddress + MAX_INST_SIZE;
auto OpMinPage = OpMinAddress & FHU::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FHU::FEX_PAGE_MASK;
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
}
if (OpMaxPage != CurrentCodePage) {
CurrentCodePage = OpMaxPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
}
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", PC + PCOffset, PC);
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", RIPToDecode + PCOffset, PC);
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
// Error while decoding instruction. We don't know the table or instruction size
+7 -1
View File
@@ -27,7 +27,7 @@ public:
Decoder(FEXCore::Context::Context *ctx);
~Decoder();
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
@@ -38,6 +38,11 @@ public:
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
void DelayedDisownBuffer() {
PoolObject.DelayedDisownBuffer();
}
private:
// To pass any information from instruction prefixes
// down into the actual instruction handling machinery.
@@ -63,6 +68,7 @@ private:
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
FEXCore::X86Tables::DecodedInst *DecodedBuffer{};
Utils::FixedSizePooledAllocation<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint8_t const *InstStream;
File diff suppressed because it is too large. Load diff
+14 -1
View File
@@ -6,8 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <istream>
#include <memory>
#include <mutex>
@@ -27,6 +29,10 @@ public:
// Public for threading
void GdbServerLoop();
void AlertLibrariesChanged() {
LibraryMapChanged = true;
}
private:
void Break(int signal);
@@ -38,6 +44,9 @@ private:
void SendACK(std::ostream &stream, bool NACK);
Event ThreadBreakEvent{};
void WaitForThreadWakeup();
struct HandledPacketType {
std::string Response{};
enum ResponseType {
@@ -74,9 +83,13 @@ private:
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string MemoryMapString{};
std::string OSDataString{};
void buildLibraryMap();
std::atomic<bool> LibraryMapChanged = true;
std::string LibraryMapString{};
// Used to keep track of which signals to pass to the guest
std::array<bool, SignalDelegator::MAX_SIGNALS + 1> PassSignals{};
uint32_t CurrentDebuggingThread{};
int ListenSocket{};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
+14 -5
View File
@@ -59,11 +59,8 @@ HostFeatures::HostFeatures() {
// Only supported when FEAT_AFP is supported
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsRCPC = Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsTSOImm9 = Features.Has(vixl::CPUFeatures::Feature::kRCpcImm);
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
@@ -83,6 +80,18 @@ HostFeatures::HostFeatures() {
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
SupportsRCPC = true;
SupportsTSOImm9 = true;
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
+1
View File
@@ -19,6 +19,7 @@ class HostFeatures final {
bool SupportsCLZERO{};
bool SupportsAtomics{};
bool SupportsRCPC{};
bool SupportsTSOImm9{};
bool SupportsRAND{};
// Float exception behaviour
@@ -4,6 +4,7 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
@@ -25,20 +26,13 @@ static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
DEF_OP(CallbackReturn) {
Data->State->CTX->InterpreterCallbackReturn(Data->State, Data->StackEntry);
Data->State->CurrentFrame->Pointers.Interpreter.CallbackReturn(Data->State, Data->StackEntry);
}
DEF_OP(ExitFunction) {
@@ -147,8 +141,8 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
Data->State->CTX->RemoveCodeEntry(Data->State, Data->CurrentEntry);
DEF_OP(RemoveThreadCodeEntry) {
Data->State->CTX->RemoveThreadCodeEntry(Data->State, Data->CurrentEntry);
}
DEF_OP(CPUID) {
@@ -513,6 +513,44 @@ DEF_OP(CRC32) {
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Selector = Op->Selector;
auto* Dst = GetDest<uint64_t*>(Data->SSAData, Node);
auto* Src1 = GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
auto* Src2 = GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const uint64_t TMP1 = (Selector & 0x01) == 0 ? Src1[0] : Src1[1];
const uint64_t TMP2 = (Selector & 0x10) == 0 ? Src2[0] : Src2[1];
const auto make_lo = [](uint64_t lhs, uint64_t rhs) {
uint64_t result = 0;
for (size_t i = 0; i < 64; i++) {
if ((lhs & (1ULL << i)) != 0) {
result ^= rhs << i;
}
}
return result;
};
const auto make_hi = [](uint64_t lhs, uint64_t rhs) {
uint64_t result = 0;
for (size_t i = 1; i < 64; i++) {
if ((lhs & (1ULL << i)) != 0) {
result ^= rhs >> (64 - i);
}
}
return result;
};
Dst[0] = make_lo(TMP1, TMP2);
Dst[1] = make_hi(TMP1, TMP2);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -356,6 +356,81 @@ DEF_OP(F80BCDSTORE) {
memcpy(GDP, BCD, 10);
}
DEF_OP(F64SIN) {
auto Op = IROp->C<IR::IROp_F64SIN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = sin(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64COS) {
auto Op = IROp->C<IR::IROp_F64COS>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = cos(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64TAN) {
auto Op = IROp->C<IR::IROp_F64TAN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = tan(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64F2XM1) {
auto Op = IROp->C<IR::IROp_F64F2XM1>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = exp2(Src) - 1.0;
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64ATAN) {
auto Op = IROp->C<IR::IROp_F64ATAN>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = atan2(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM) {
auto Op = IROp->C<IR::IROp_F64FPREM>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = fmod(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM1) {
auto Op = IROp->C<IR::IROp_F64FPREM1>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = remainder(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FYL2X) {
auto Op = IROp->C<IR::IROp_F64FYL2X>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = Src2 * log2(Src1);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64SCALE) {
auto Op = IROp->C<IR::IROp_F64SCALE>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double trunc = (double)(int64_t)(Src2); //truncate
double Tmp = Src1 * exp2(trunc);
memcpy(GDP, &Tmp, sizeof(double));
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -222,6 +222,73 @@ struct OpHandlers<IR::OP_F80SCALE> {
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
static double handle(double src) {
return sin(src);
}
};
template<>
struct OpHandlers<IR::OP_F64COS> {
static double handle(double src) {
return cos(src);
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
static double handle(double src) {
return tan(src);
}
};
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
static double handle(double src) {
return exp2(src) - 1.0;
}
};
template<>
struct OpHandlers<IR::OP_F64ATAN> {
static double handle(double src1, double src2) {
return atan2(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM> {
static double handle(double src1, double src2) {
return fmod(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
static double handle(double src1, double src2) {
return remainder(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
static double handle(double src1, double src2) {
return src2 * log2(src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(double src1, double src2) {
double trunc = (double)(int64_t)(src2); //truncate
return src1 * exp2(trunc);
}
};
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
static X80SoftFloat handle(X80SoftFloat Src1) {
@@ -8,6 +8,7 @@
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::CPU {
class Dispatcher;
class X86DispatchGenerator;
class Arm64DispatchGenerator;
@@ -20,28 +21,27 @@ using DestMapType = std::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
explicit InterpreterCore(Dispatcher *Dispatch,
FEXCore::Core::InternalThreadState *Thread);
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearCache() override;
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
std::unique_ptr<Dispatcher> Dispatcher{};
size_t BufferUsed;
Dispatcher *Dispatch;
};
template<typename T>
@@ -9,66 +9,101 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <memory>
#include <signal.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include "InterpreterOps.h"
#if defined(_M_X86_64)
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#elif defined(_M_ARM_64)
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#else
#error missing arch
#endif
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
namespace FEXCore::IR {
class IRListView;
class RegisterAllocationData;
}
namespace FEXCore::CPU {
class CPUBackend;
static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
InterpreterCore::InterpreterCore(Dispatcher *Dispatcher, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Dispatch(Dispatcher)
{
auto LocalEntry = Thread->LocalIRCache.find(Thread->CurrentFrame->State.rip);
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
Interpreter.FragmentExecuter = reinterpret_cast<uint64_t>(&InterpreterOps::InterpretIR);
ClearCache();
}
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: CTX {ctx}
, State {Thread} {
if (!CompileThread &&
CTX->Config.Core == FEXCore::Config::CONFIG_INTERPRETER) {
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
}
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
const auto IRSize = AlignUp(IR->GetInlineSize(), 16);
const auto MaxSize = IRSize + Dispatcher::MaxInterpreterTrampolineSize + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((BufferUsed + MaxSize) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState);
}
const auto BufferStart = CurrentCodeBuffer->Ptr + BufferUsed;
auto DestBuffer = BufferStart;
if (GDBEnabled) {
const auto GDBSize = Dispatch->GenerateGDBPauseCheck(DestBuffer, Entry);
DestBuffer += GDBSize;
BufferUsed += GDBSize;
}
const auto TrampolineSize = Dispatch->GenerateInterpreterTrampoline(DestBuffer);
DestBuffer += TrampolineSize;
BufferUsed += TrampolineSize;
IR->Serialize(DestBuffer);
DestBuffer += IRSize;
BufferUsed += IRSize;
return BufferStart;
}
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
return reinterpret_cast<void*>(InterpreterExecution);
void InterpreterCore::ClearCache() {
// Calling this one is needed to setup the initial CurrentCodeBuffer
[[maybe_unused]] auto CodeBuffer = GetEmptyCodeBuffer();
BufferUsed = 0;
}
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<InterpreterCore>(ctx, Thread, CompileThread);
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<InterpreterCore>(ctx->Dispatcher.get(), Thread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
CPUBackendFeatures GetInterpreterBackendFeatures() {
return CPUBackendFeatures { };
}
}
@@ -12,9 +12,11 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
struct DispatcherConfig;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetInterpreterBackendFeatures();
} // namespace FEXCore::CPU
@@ -47,6 +47,16 @@ FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat), FEXCore::Core::FallbackH
return {FABI_F64_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(double,double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F64_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I16_F80, (void*)fn, HandlerIndex};
@@ -122,6 +132,18 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM1>::handle, Core::OPINDEX_F80FPREM1).fn);
Info[Core::OPINDEX_F80FPREM] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM>::handle, Core::OPINDEX_F80FPREM).fn);
Info[Core::OPINDEX_F80SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SCALE>::handle, Core::OPINDEX_F80SCALE).fn);
// Double Precision
Info[Core::OPINDEX_F64SIN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle, Core::OPINDEX_F64SIN).fn);
Info[Core::OPINDEX_F64COS] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle, Core::OPINDEX_F64COS).fn);
Info[Core::OPINDEX_F64TAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle, Core::OPINDEX_F64TAN).fn);
Info[Core::OPINDEX_F64ATAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64ATAN>::handle, Core::OPINDEX_F64ATAN).fn);
Info[Core::OPINDEX_F64F2XM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle, Core::OPINDEX_F64F2XM1).fn);
Info[Core::OPINDEX_F64FYL2X] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle, Core::OPINDEX_F64FYL2X).fn);
Info[Core::OPINDEX_F64FPREM] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM>::handle, Core::OPINDEX_F64FPREM).fn);
Info[Core::OPINDEX_F64FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle, Core::OPINDEX_F64FPREM1).fn);
Info[Core::OPINDEX_F64SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle, Core::OPINDEX_F64SCALE).fn);
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
@@ -238,6 +260,12 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
return true; \
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
@@ -261,6 +289,19 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
// Double Precision Unary
COMMON_F64_OP(F2XM1)
COMMON_F64_OP(TAN)
COMMON_F64_OP(SIN)
COMMON_F64_OP(COS)
// Double Precision Binary
COMMON_F64_OP(FYL2X)
COMMON_F64_OP(ATAN)
COMMON_F64_OP(FPREM1)
COMMON_F64_OP(FPREM)
COMMON_F64_OP(SCALE)
default:
break;
}
@@ -112,8 +112,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -123,7 +121,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
// Conversion ops
@@ -166,6 +164,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -176,6 +175,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
@@ -283,6 +283,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
@@ -311,6 +312,17 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
// F64 ops
REGISTER_OP(F64SIN, F64SIN);
REGISTER_OP(F64COS, F64COS);
REGISTER_OP(F64TAN, F64TAN);
REGISTER_OP(F64F2XM1, F64F2XM1);
REGISTER_OP(F64ATAN, F64ATAN);
REGISTER_OP(F64FPREM, F64FPREM);
REGISTER_OP(F64FPREM1, F64FPREM1);
REGISTER_OP(F64FYL2X, F64FYL2X);
REGISTER_OP(F64SCALE, F64SCALE);
return Handlers;
}();
@@ -321,15 +333,9 @@ void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
void InterpreterOps::InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *CurrentIR) {
volatile void *StackEntry = alloca(0);
// Debug data is only passed in debug builds
#ifndef NDEBUG
// TODO: should be moved to an IR Op
Thread->Stats.InstructionsExecuted.fetch_add(DebugData->GuestInstructionCount);
#endif
uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
@@ -338,9 +344,9 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
auto BlockEnd = CurrentIR->GetBlocks().end();
InterpreterOps::IROpData OpData{};
OpData.State = Thread;
OpData.State = Frame->Thread;
OpData.SSAData = alloca(ListSize * 16);
OpData.CurrentEntry = Entry;
OpData.CurrentEntry = Frame->State.rip;
OpData.CurrentIR = CurrentIR;
OpData.StackEntry = StackEntry;
OpData.BlockIterator = CurrentIR->GetBlocks().begin();
@@ -28,6 +28,8 @@ namespace FEXCore::CPU {
FABI_F80_I32,
FABI_F32_F80,
FABI_F64_F80,
FABI_F64_F64,
FABI_F64_F64_F64,
FABI_I16_F80,
FABI_I32_F80,
FABI_I64_F80,
@@ -45,14 +47,14 @@ namespace FEXCore::CPU {
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static void InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *IR);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
uint64_t CurrentEntry{};
FEXCore::IR::IRListView *CurrentIR{};
FEXCore::IR::IRListView const *CurrentIR{};
volatile void *StackEntry{};
void *SSAData{};
struct {
@@ -140,8 +142,6 @@ namespace FEXCore::CPU {
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -151,7 +151,7 @@ namespace FEXCore::CPU {
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -197,6 +197,7 @@ namespace FEXCore::CPU {
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -301,6 +302,7 @@ namespace FEXCore::CPU {
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
///< F80 ops
DEF_OP(F80LOADFCW);
@@ -328,6 +330,17 @@ namespace FEXCore::CPU {
DEF_OP(F80CMP);
DEF_OP(F80BCDLOAD);
DEF_OP(F80BCDSTORE);
//< F64 ops
DEF_OP(F64SIN);
DEF_OP(F64COS);
DEF_OP(F64TAN);
DEF_OP(F64F2XM1);
DEF_OP(F64ATAN);
DEF_OP(F64FPREM);
DEF_OP(F64FPREM1);
DEF_OP(F64FYL2X);
DEF_OP(F64SCALE);
#undef DEF_OP
template<typename unsigned_type, typename signed_type, typename float_type>
[[nodiscard]] static bool IsConditionTrue(uint8_t Cond, uint64_t Src1, uint64_t Src2) {
@@ -394,7 +407,7 @@ namespace FEXCore::CPU {
return CompResult;
}
static uint8_t GetOpSize(FEXCore::IR::IRListView *CurrentIR, IR::OrderedNodeWrapper Node) {
static uint8_t GetOpSize(FEXCore::IR::IRListView const *CurrentIR, IR::OrderedNodeWrapper Node) {
auto IROp = CurrentIR->GetOp<FEXCore::IR::IROp_Header>(Node);
return IROp->Size;
}
@@ -4,6 +4,7 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
@@ -4,6 +4,8 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
@@ -157,6 +159,11 @@ DEF_OP(RDRAND) {
// Second result is if we managed to read a valid random number or not
DstPtr[1] = Result == 8 ? 1 : 0;
}
DEF_OP(Yield) {
// Nop implementation
}
#undef DEF_OP
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,131 @@
/*
$info$
tags: backend|arm64
desc: relocation logic of the arm64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.GetCode();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
.Lit = Literal(Pointer),
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
place(&Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
void Arm64JITCore::InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.GetCode();
LoadConstant(Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
Literal<uint64_t> Lit(Pointer);
place(&Lit);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
}
return true;
}
}
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
@@ -509,10 +510,10 @@ DEF_OP(AtomicSwap) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: swplb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swplh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpl(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpl(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: swpalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swpalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpal(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpal(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Core/LookupCache.h"
@@ -19,13 +20,6 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// First we must reset the stack
@@ -33,7 +27,7 @@ DEF_OP(SignalReturn) {
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalReturnHandler)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)));
br(x0);
}
@@ -46,10 +40,10 @@ DEF_OP(CallbackReturn) {
ResetStack();
// We can now lower the ref counter again
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalHandlerRefCountPointer)));
ldr(w2, MemOperand(x0));
ldr(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
sub(w2, w2, 1);
str(w2, MemOperand(x0));
str(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
@@ -73,7 +67,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchHost{ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker};
Literal l_BranchGuest{NewRIP};
ldr(x0, &l_BranchHost);
@@ -82,10 +76,10 @@ DEF_OP(ExitFunction) {
place(&l_BranchHost);
place(&l_BranchGuest);
} else {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
RipReg = GetReg<RA_64>(Op->NewRIP.ID());
// L1 Cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -96,7 +90,7 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.DispatcherLoopTop)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop)));
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
@@ -104,9 +98,9 @@ DEF_OP(ExitFunction) {
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
const auto Target = Op->TargetBlock.ID();
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
@@ -199,8 +193,8 @@ DEF_OP(Syscall) {
str(GetReg<RA_64>(Op->Header.Args[i].ID()), MemOperand(sp, i * 8));
}
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerFunc)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
blr(x3);
@@ -383,7 +377,7 @@ DEF_OP(Thunk) {
PushDynamicRegsAndLR();
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x0, GetReg<RA_64>(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
@@ -442,7 +436,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
DEF_OP(RemoveThreadCodeEntry) {
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
@@ -452,7 +446,7 @@ DEF_OP(RemoveCodeEntry) {
mov(x0, STATE);
LoadConstant(x1, Entry);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.RemoveCodeEntryFromJIT)));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -469,10 +463,10 @@ DEF_OP(CPUID) {
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Function.ID()));
mov(x2, GetReg<RA_64>(Op->Leaf.ID()));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
@@ -489,8 +483,6 @@ DEF_OP(CPUID) {
#undef DEF_OP
void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -500,7 +492,7 @@ void Arm64JITCore::RegisterBranchHandlers() {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -13,22 +13,22 @@ using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
mov(GetDst(Node), GetSrc(Op->DestVector.ID()));
switch (Op->Header.ElementSize) {
case 1: {
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 2: {
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 4: {
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 8: {
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Src.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -39,18 +39,18 @@ DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
switch (Op->Header.ElementSize) {
case 1:
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
uxtb(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 2:
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
uxth(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 4:
fmov(GetDst(Node).S(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
fmov(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()).W());
break;
case 8:
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Header.Args[0].ID()).X());
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()).X());
break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
}
@@ -58,22 +58,22 @@ DEF_OP(VCastFromGPR) {
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()));
break;
}
case 0x0408: { // Float <- int64_t
scvtf(GetDst(Node).S(), GetReg<RA_64>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).S(), GetReg<RA_64>(Op->Src.ID()));
break;
}
case 0x0804: { // Double <- int32_t
scvtf(GetDst(Node).D(), GetReg<RA_32>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).D(), GetReg<RA_32>(Op->Src.ID()));
break;
}
case 0x0808: { // Double <- int64_t
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()));
break;
}
}
@@ -81,14 +81,14 @@ DEF_OP(Float_FromGPR_S) {
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // Double <- Float
fcvt(GetDst(Node).D(), GetSrc(Op->Header.Args[0].ID()).S());
fcvt(GetDst(Node).D(), GetSrc(Op->Scalar.ID()).S());
break;
}
case 0x0408: { // Float <- Double
fcvt(GetDst(Node).S(), GetSrc(Op->Header.Args[0].ID()).D());
fcvt(GetDst(Node).S(), GetSrc(Op->Scalar.ID()).D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
@@ -99,10 +99,10 @@ DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
switch (Op->Header.ElementSize) {
case 4:
scvtf(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
scvtf(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
scvtf(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
scvtf(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", Op->Header.ElementSize);
}
@@ -112,10 +112,10 @@ DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
switch (Op->Header.ElementSize) {
case 4:
fcvtzs(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
fcvtzs(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", Op->Header.ElementSize);
}
@@ -125,11 +125,11 @@ DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
fcvtzs(GetDst(Node).V4S(), GetDst(Node).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetDst(Node).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", Op->Header.ElementSize);
@@ -142,11 +142,11 @@ DEF_OP(Vector_FToF) {
switch (Conv) {
case 0x0804: { // Double <- Float
fcvtl(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2S());
fcvtl(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2S());
break;
}
case 0x0408: { // Float <- Double
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Vector.ID()).V2D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
@@ -159,50 +159,50 @@ DEF_OP(Vector_FToI) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
case 4:
frintn(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintn(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintn(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintn(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintm(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintm(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintm(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintm(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintp(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintp(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintp(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintp(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
case 4:
frintz(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintz(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintz(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintz(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
@@ -14,41 +14,41 @@ using namespace vixl::aarch64;
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
aesimc(GetDst(Node).V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
aesimc(GetDst(Node).V16B(), GetSrc(Op->Vector.ID()).V16B());
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
aesmc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
aesimc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESKeyGenAssist) {
@@ -59,7 +59,7 @@ DEF_OP(AESKeyGenAssist) {
// Do a "regular" AESE step
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->Src.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
// Do a table shuffle to undo ShiftRows
@@ -102,16 +102,45 @@ DEF_OP(CRC32) {
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
auto Dst = GetDst(Node).Q();
auto Src1 = GetSrc(Op->Src1.ID()).V2D();
auto Src2 = GetSrc(Op->Src2.ID()).V2D();
switch (Op->Selector) {
case 0b00000000:
pmull(Dst, Src1, Src2);
break;
case 0b00000001:
mov(VTMP1.V1D(), Src1, 1);
pmull(Dst, VTMP1.V2D(), Src2);
break;
case 0b00010000:
mov(VTMP1.V1D(), Src2, 1);
pmull(Dst, VTMP1.V2D(), Src1);
break;
case 0b00010001:
pmull2(Dst, Src1, Src2);
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
#undef REGISTER_OP
}
}
@@ -13,7 +13,7 @@ using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
+204 -248
View File
@@ -35,6 +35,10 @@ $end_info$
#include <unistd.h>
#include <string.h>
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
namespace {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
@@ -71,11 +75,6 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
namespace FEXCore::CPU {
void Arm64JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
}
using namespace vixl;
using namespace vixl::aarch64;
@@ -93,7 +92,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -108,7 +107,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -127,7 +126,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -152,7 +151,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
else {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -173,7 +172,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -192,7 +191,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -203,6 +202,43 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
break;
case FABI_F64_F64: {
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetDst(Node).D(), v0.D());
}
break;
case FABI_F64_F64_F64: {
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
mov(v1.D(), GetSrc(IROp->Args[1].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetDst(Node).D(), v0.D());
}
break;
case FABI_I16_F80:{
SpillStaticRegs();
@@ -211,7 +247,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -229,7 +265,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -247,7 +283,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -268,7 +304,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -286,7 +322,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -309,7 +345,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -332,81 +368,63 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
//fmt::print("ExitFunctionLink: Aborting, {:X} not in cache\n", GuestRip);
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
auto offset = HostCode/4 - branch/4;
if (IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.b(offset);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
emit.ldr(x0, &l_BranchHost);
emit.blr(x0);
emit.place(&l_BranchHost);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
});
} else {
// fallback case - do a soft-er link by patching the pointer
record[0] = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
}
return HostCode;
}
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
Arm64JITCore::CodeBuffer Arm64JITCore::AllocateNewCodeBuffer(size_t Size) {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(
FEXCore::Allocator::mmap(nullptr,
Buffer.Size,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS,
-1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
Dispatcher->RegisterCodeBuffer(Buffer.Ptr, Buffer.Size);
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
Dispatcher->RemoveCodeBuffer(Buffer.Ptr);
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
CurrentCodeBuffer = &InitialCodeBuffer;
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx, 0)
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -447,95 +465,76 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RegisterVectorHandlers();
RegisterEncryptionHandlers();
if (!CompileThread) {
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalReturnInstruction = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
{
// Set up pointers that the JIT needs to load
ThreadSharedData.Dispatcher = Dispatcher.get();
// Common
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore_ExitFunctionLink);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
// Platform Specific
auto &AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
AArch64.LUREM = reinterpret_cast<uint64_t>(LUREM);
AArch64.LREM = reinterpret_cast<uint64_t>(LREM);
}
// Must be done after Dispatcher init
SetAllowAssembler(true);
ClearCache();
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
if (!Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Thread->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
auto Buffer = GetBuffer();
Buffer->EmitString(JITString);
Buffer->Align();
}
void Arm64JITCore::ClearCache() {
// Get the backing code buffer
auto Buffer = GetBuffer();
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
// Set the current code buffer to the initial
*Buffer = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
CurrentCodeBuffer = &InitialCodeBuffer;
}
if (CurrentCodeBuffer->Size == MAX_CODE_SIZE) {
// Rewind to the start of the code cache start
Buffer->Reset();
}
else {
FreeCodeBuffer(InitialCodeBuffer);
// Resize the code buffer and reallocate our code size
InitialCodeBuffer.Size *= 1.5;
InitialCodeBuffer.Size = std::min(InitialCodeBuffer.Size, MAX_CODE_SIZE);
InitialCodeBuffer = AllocateNewCodeBuffer(InitialCodeBuffer.Size);
*Buffer = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
}
else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = Arm64JITCore::AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
*Buffer = vixl::CodeBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
auto CodeBuffer = GetEmptyCodeBuffer();
*GetBuffer() = vixl::CodeBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
FreeCodeBuffer(InitialCodeBuffer);
}
IR::PhysicalRegister Arm64JITCore::GetPhys(IR::NodeID Node) const {
@@ -666,13 +665,14 @@ bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
using namespace aarch64;
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
#ifndef NDEBUG
LoadConstant(x0, Entry);
@@ -681,9 +681,9 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
this->IR = IR;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((GetCursorOffset() + BufferRange) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState, false);
CTX->ClearCodeCache(ThreadState);
}
// AAPCS64
@@ -706,31 +706,11 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
// X1-X3 = Temp
// X4-r18 = RA
auto GuestEntry = GetCursorAddress<uint64_t>();
GuestEntry = GetCursorAddress<uint8_t *>();
if (CTX->GetGdbServerStatus()) {
aarch64::Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Thread))); // Get thread
ldr(x0, MemOperand(x0, offsetof(FEXCore::Core::InternalThreadState, CTX))); // Get Context
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then we don't need to stop
cbz(w0, &RunBlock);
{
// Make sure RIP is syncronized to the context
LoadConstant(x0, Entry);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// Stop the thread
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(x0);
}
bind(&RunBlock);
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
GetBuffer()->CursorForward(GDBSize);
}
//LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
@@ -755,6 +735,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = GetCursorAddress<uint8_t *>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -769,10 +750,6 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({GetCursorAddress<uintptr_t>(), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
@@ -782,7 +759,10 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = GetCursorAddress<uintptr_t>() - DebugData->Subblocks.back().HostCodeStart;
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t *>() - BlockStartHostCode)
});
}
}
@@ -795,68 +775,44 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
FinalizeCode();
auto CodeEnd = GetCursorAddress<uint64_t>();
CPU.EnsureIAndDCacheCoherency(reinterpret_cast<void*>(GuestEntry), CodeEnd - reinterpret_cast<uint64_t>(GuestEntry));
auto CodeEnd = GetCursorAddress<uint8_t *>();
CPU.EnsureIAndDCacheCoherency(GuestEntry, CodeEnd - GuestEntry);
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(CodeEnd) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->HostCodeSize = CodeEnd - GuestEntry;
DebugData->Relocations = &Relocations;
}
this->IR = nullptr;
return reinterpret_cast<void*>(GuestEntry);
return GuestEntry;
}
uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
void Arm64JITCore::ResetStack() {
if (SpillSlots == 0)
return;
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
//fmt::print("ExitFunctionLink: Aborting, {:X} not in cache\n", GuestRip);
Frame->State.rip = GuestRip;
return core->ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = core->ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress;
auto offset = HostCode/4 - branch/4;
if (IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.b(offset);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
emit.ldr(x0, &l_BranchHost);
emit.blr(x0);
emit.place(&l_BranchHost);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
});
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// fallback case - do a soft-er link by patching the pointer
record[0] = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
return HostCode;
}
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<Arm64JITCore>(ctx, Thread, CompileThread);
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<Arm64JITCore>(ctx, Thread);
}
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
}
CPUBackendFeatures GetArm64JITBackendFeatures() {
return CPUBackendFeatures {
.SupportsStaticRegisterAllocation = true
};
}
}
+83 -60
View File
@@ -6,17 +6,23 @@ $end_info$
#pragma once
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <array>
#include <cstdint>
#include <map>
#include <utility>
#include <vector>
#define STATE x28
#define TMP1 x0
#define TMP2 x1
@@ -37,14 +43,8 @@ using namespace vixl::aarch64;
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
~Arm64JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
@@ -52,7 +52,7 @@ public:
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -60,21 +60,15 @@ public:
void ClearCache() override;
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
void ClearRelocations() override { Relocations.clear(); }
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
@@ -145,46 +139,77 @@ private:
vixl::aarch64::Decoder Decoder;
#endif
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
void FreeCodeBuffer(CodeBuffer Buffer);
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
CodeBuffer InitialCodeBuffer{};
// This is the array of /additional/ code buffers that we may need to allocate
// Allocation only occurs when we've hit signals and need to clear code cache
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
// We don't want to mvoe above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
#if DEBUG
vixl::aarch64::Disassembler Disasm;
#endif
static uint64_t ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
FEXCore::Core::DebugData *DebugData;
void ResetStack();
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Literal<uint64_t> Lit;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
@@ -275,8 +300,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -286,7 +309,7 @@ private:
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -326,7 +349,7 @@ private:
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
DEF_OP(GuestOpcode);
DEF_OP(Fence);
DEF_OP(Break);
DEF_OP(Phi);
@@ -336,6 +359,7 @@ private:
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -442,9 +466,8 @@ private:
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
#undef DEF_OP
};
}
} // namespace FEXCore::CPU
@@ -4,6 +4,8 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include <FEXCore/Utils/CompilerDefs.h>
@@ -58,26 +60,26 @@ DEF_OP(LoadContext) {
DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1:
strb(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
strb(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 2:
strh(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
strh(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 4:
str(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
str(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 8:
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
str(GetReg<RA_64>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
}
}
else {
auto Src = GetSrc(Op->Header.Args[0].ID());
auto Src = GetSrc(Op->Value.ID());
switch (OpSize) {
case 1:
str(Src.B(), MemOperand(STATE, Op->Offset));
@@ -104,7 +106,7 @@ DEF_OP(LoadRegister) {
auto Op = IROp->C<IR::IROp_LoadRegister>();
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0])) / 8;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
@@ -135,7 +137,7 @@ DEF_OP(LoadRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "out of range regId");
@@ -189,7 +191,7 @@ DEF_OP(StoreRegister) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
if (Op->Class == IR::GPRClass) {
auto regId = Op->Offset / 8 - 1;
auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
@@ -219,7 +221,7 @@ DEF_OP(StoreRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "regId out of range");
@@ -261,8 +263,8 @@ DEF_OP(StoreRegister) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[0].ID());
const size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Index.ID());
if (Op->Class == FEXCore::IR::GPRClass) {
switch (Op->Stride) {
@@ -349,11 +351,11 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[1].ID());
const size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Index.ID());
if (Op->Class == FEXCore::IR::GPRClass) {
auto value = GetReg<RA_64>(Op->Header.Args[0].ID());
auto value = GetReg<RA_64>(Op->Value.ID());
switch (Op->Stride) {
case 1:
@@ -392,7 +394,7 @@ DEF_OP(StoreContextIndexed) {
}
}
else {
auto value = GetSrc(Op->Header.Args[0].ID());
auto value = GetSrc(Op->Value.ID());
switch (Op->Stride) {
case 1:
@@ -441,25 +443,25 @@ DEF_OP(StoreContextIndexed) {
DEF_OP(SpillRegister) {
auto Op = IROp->C<IR::IROp_SpillRegister>();
uint8_t OpSize = IROp->Size;
uint32_t SlotOffset = Op->Slot * 16;
const uint8_t OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * 16;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
strb(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 2: {
strh(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
strh(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 4: {
str(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetReg<RA_32>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 8: {
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -467,15 +469,15 @@ DEF_OP(SpillRegister) {
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
case 4: {
str(GetSrc(Op->Header.Args[0].ID()).S(), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()).S(), MemOperand(sp, SlotOffset));
break;
}
case 8: {
str(GetSrc(Op->Header.Args[0].ID()).D(), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()).D(), MemOperand(sp, SlotOffset));
break;
}
case 16: {
str(GetSrc(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -539,7 +541,7 @@ DEF_OP(LoadFlag) {
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
strb(GetReg<RA_64>(Op->Value.ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
}
MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
@@ -570,7 +572,7 @@ MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Registe
DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -617,13 +619,43 @@ DEF_OP(LoadMem) {
DEF_OP(LoadMemTSO) {
auto Op = IROp->C<IR::IROp_LoadMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("LoadMemTSO: No offset allowed");
if (CTX->HostFeatures.SupportsTSOImm9) {
// RCPC2 means that the offset must be an inline constant
LOGMAN_THROW_A_FMT(MemSrc.IsRegisterOffset() == false, "RCPC2 doesn't support register offset. Only Immediate offset");
}
else {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid(), "LoadMemTSO: No offset allowed");
}
if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldapurb(Dst, MemSrc);
}
else {
// Aligned
nop();
auto Dst = GetReg<RA_64>(Node);
switch (IROp->Size) {
case 2:
ldapurh(Dst, MemSrc);
break;
case 4:
ldapur(Dst.W(), MemSrc);
break;
case 8:
ldapur(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", IROp->Size);
}
nop();
}
}
else if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
@@ -744,13 +776,41 @@ DEF_OP(StoreMem) {
DEF_OP(StoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("StoreMemTSO: No offset allowed");
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (CTX->HostFeatures.SupportsTSOImm9) {
// RCPC2 means that the offset must be an inline constant
LOGMAN_THROW_A_FMT(MemSrc.IsRegisterOffset() == false, "RCPC2 doesn't support register offset. Only Immediate offset");
}
else {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid(), "StoreMemTSO: No offset allowed");
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlurb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
nop();
switch (IROp->Size) {
case 2:
stlurh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlur(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlur(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
nop();
}
}
else if (Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
@@ -934,7 +994,7 @@ DEF_OP(VStoreMemElement) {
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
@@ -949,7 +1009,7 @@ DEF_OP(CacheLineClear) {
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use this instruction directly
+25 -11
View File
@@ -4,13 +4,21 @@ tags: backend|arm64
$end_info$
*/
#include <syscall.h>
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - GuestEntry});
}
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -36,7 +44,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.OverflowExceptionHandler)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)));
br(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
@@ -46,13 +54,13 @@ DEF_OP(Break) {
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadStopHandlerSpillSRA)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)));
br(TMP1);
break;
}
@@ -60,7 +68,7 @@ DEF_OP(Break) {
{
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.UnimplementedInstructionHandler)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)));
br(TMP1);
break;
@@ -97,7 +105,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Src = GetReg<RA_64>(Op->RoundMode.ID());
// Setup the rounding flags correctly
and_(TMP1, Src, 0b11);
@@ -132,15 +140,15 @@ DEF_OP(Print) {
PushDynamicRegsAndLR();
if (IsGPR(Op->Header.Args[0].ID())) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintValue)));
if (IsGPR(Op->Value.ID())) {
mov(x0, GetReg<RA_64>(Op->Value.ID()));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)));
}
else {
fmov(x0, GetSrc(Op->Header.Args[0].ID()).V1D());
fmov(x0, GetSrc(Op->Value.ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Header.Args[0].ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintVectorValue)));
fmov(x1, GetSrc(Op->Value.ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)));
}
SpillStaticRegs();
blr(x3);
@@ -219,6 +227,10 @@ DEF_OP(RDRAND) {
cset(Dst.second, Condition::ne);
}
DEF_OP(Yield) {
hint(SystemHint::YIELD);
}
#undef DEF_OP
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -227,6 +239,7 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, GuestOpcode);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -237,6 +250,7 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
#undef REGISTER_OP
}
@@ -15,13 +15,13 @@ DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
case 4: {
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_32>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_32>(Node), Regs[Op->Element]);
break;
}
case 8: {
auto Src = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_64>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_64>(Node), Regs[Op->Element]);
break;
@@ -40,15 +40,15 @@ DEF_OP(CreateElementPair) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_32>(Op->Header.Args[1].ID());
RegFirst = GetReg<RA_32>(Op->Lower.ID());
RegSecond = GetReg<RA_32>(Op->Upper.ID());
RegTmp = w0;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetReg<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_64>(Op->Header.Args[1].ID());
RegFirst = GetReg<RA_64>(Op->Lower.ID());
RegSecond = GetReg<RA_64>(Op->Upper.ID());
RegTmp = x0;
break;
}
@@ -70,7 +70,7 @@ DEF_OP(CreateElementPair) {
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
mov(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()));
}
#undef DEF_OP
File diff suppressed because it is too large. Load diff
+7 -4
View File
@@ -14,10 +14,13 @@ namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetX86JITBackendFeatures();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -29,13 +29,6 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
@@ -43,7 +36,7 @@ DEF_OP(SignalReturn) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
}
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalReturnHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)]);
}
DEF_OP(CallbackReturn) {
@@ -53,7 +46,7 @@ DEF_OP(CallbackReturn) {
}
// Make sure to adjust the refcounter so we don't clear the cache now
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
@@ -91,14 +84,15 @@ DEF_OP(ExitFunction) {
jmp(qword[rax]);
L(l_BranchHost);
dq(ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress);
//FEX_TODO(this is not per thread)
dq(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
L(l_BranchGuest);
dq(NewRIP);
} else {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer)]);
mov(rax, RipReg);
@@ -113,7 +107,7 @@ DEF_OP(ExitFunction) {
L(FullLookup);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.DispatcherLoopTop)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop)]);
}
#ifdef BLOCKSTATS
@@ -181,13 +175,13 @@ DEF_OP(Syscall) {
}
mov(rsi, STATE); // Move thread in to rsi
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerObj)]);
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj)]);
mov(rdx, rsp);
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerFunc)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -259,7 +253,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
DEF_OP(RemoveThreadCodeEntry) {
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -272,7 +266,7 @@ DEF_OP(RemoveCodeEntry) {
mov(rax, Entry); // imm64 move
mov(rsi, rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.RemoveCodeEntryFromJIT)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -296,7 +290,7 @@ DEF_OP(CPUID) {
// rsi can be in the source registers, so copy argument to edx first
mov (edx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov (esi, GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDObj)]);
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)]);
auto NumPush = RA64.size();
@@ -304,7 +298,7 @@ DEF_OP(CPUID) {
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDFunction)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -320,8 +314,6 @@ DEF_OP(CPUID) {
#undef DEF_OP
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -330,7 +322,7 @@ void X86JITCore::RegisterBranchHandlers() {
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -75,16 +75,37 @@ DEF_OP(CRC32) {
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
auto Dst = GetDst(Node);
auto Src1 = GetSrc(Op->Src1.ID());
auto Src2 = GetSrc(Op->Src2.ID());
switch (Op->Selector) {
case 0b00000000:
case 0b00000001:
case 0b00010000:
case 0b00010001:
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
#undef DEF_OP
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
#undef REGISTER_OP
}
}
+134 -190
View File
@@ -43,6 +43,9 @@ $end_info$
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
namespace {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
@@ -55,31 +58,6 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
namespace FEXCore::CPU {
CodeBuffer AllocateNewCodeBuffer(FEXCore::Context::Context *CTX, size_t Size) {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(
FEXCore::Allocator::mmap(nullptr,
Buffer.Size,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS,
-1, 0));
LOGMAN_THROW_A_FMT(Buffer.Ptr != reinterpret_cast<uint8_t*>(~0ULL), "Couldn't allocate code buffer");
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
void X86JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
}
void X86JITCore::PushRegs() {
sub(rsp, 16 * RAXMM_x.size());
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
@@ -120,7 +98,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_VOID_U16: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
break;
@@ -129,7 +107,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movss(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -143,7 +121,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -158,7 +136,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -174,7 +152,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -188,7 +166,34 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
movsd(GetDst(Node), xmm0);
}
break;
case FABI_F64_F64: {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
movsd(GetDst(Node), xmm0);
}
break;
case FABI_F64_F64_F64: {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
movsd(xmm1, GetSrc(IROp->Args[1].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -202,7 +207,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -215,7 +220,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -228,7 +233,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -244,7 +249,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -257,7 +262,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -275,7 +280,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -295,41 +300,34 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
}
static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->CurrentFrame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
}
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
record[0] = HostCode;
return HostCode;
}
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
: CodeGenerator(Buffer.Size, Buffer.Ptr, nullptr)
, CTX {ctx}
, ThreadState {Thread}
, InitialCodeBuffer {Buffer}
{
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
CurrentCodeBuffer = &InitialCodeBuffer;
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, CodeGenerator(0, this, nullptr) // this is not used here
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -358,92 +356,52 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
RegisterVectorHandlers();
RegisterEncryptionHandlers();
if (!CompileThread) {
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
{
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&X86JITCore_ExitFunctionLink);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
}
// Must be done after Dispatcher init
ClearCache();
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
}
X86JITCore::~X86JITCore() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
void X86JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::X86JITCore::";
for (char c : JITString) {
db(c);
}
CodeBuffers.clear();
FreeCodeBuffer(InitialCodeBuffer);
}
void X86JITCore::ClearCache() {
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
// Set the current code buffer to the initial
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
CurrentCodeBuffer = &InitialCodeBuffer;
}
if (CurrentCodeBuffer->Size == MAX_CODE_SIZE) {
// Rewind to the start of the code cache start
reset();
}
else {
FreeCodeBuffer(InitialCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MAX_CODE_SIZE);
InitialCodeBuffer = AllocateNewCodeBuffer(CTX, CurrentCodeBuffer->Size);
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
}
else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(CTX, X86JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
auto CodeBuffer = GetEmptyCodeBuffer();
setNewBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
}
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
@@ -611,39 +569,27 @@ std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::G
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((getSize() + BufferRange) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState, false);
CTX->ClearCodeCache(ThreadState);
}
void *GuestEntry = getCurr<void*>();
GuestEntry = getCurr<uint8_t*>();
CursorEntry = getSize();
this->IR = IR;
if (CTX->GetGdbServerStatus()) {
Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(RunBlock);
// Else we need to pause now
mov(rax, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddress);
jmp(rax);
ud2();
L(RunBlock);
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
setSize(getSize() + GDBSize);
}
LOGMAN_THROW_A_FMT(RAData != nullptr, "Needs RA");
@@ -701,12 +647,13 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
{
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = getCurr<uint8_t *>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -761,6 +708,13 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(getCurr<uint8_t *>() - BlockStartHostCode)
});
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
@@ -777,32 +731,22 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(GuestExit) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->Relocations = &Relocations;
}
return GuestEntry;
}
uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->CurrentFrame->State.rip = GuestRip;
return core->ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress;
}
auto LinkerAddress = core->ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
record[0] = HostCode;
return HostCode;
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread);
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
CPUBackendFeatures GetX86JITBackendFeatures() {
return CPUBackendFeatures { };
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
}
+80 -50
View File
@@ -6,8 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#define XBYAK64
#include <xbyak/xbyak.h>
@@ -24,13 +26,6 @@ using namespace Xbyak;
#include <tuple>
namespace FEXCore::CPU {
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void FreeCodeBuffer(CodeBuffer Buffer);
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
@@ -57,9 +52,7 @@ const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
CodeBuffer Buffer,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
~X86JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
@@ -67,7 +60,7 @@ public:
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -75,20 +68,75 @@ public:
void ClearCache() override;
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
void ClearRelocations() override { Relocations.clear(); }
private:
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
void LoadConstantWithPadding(Xbyak::Reg Reg, uint64_t Constant);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Label Offset;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(Xbyak::Reg Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(Xbyak::Reg Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/**
* @brief Current guest RIP entrypoint
*/
uint64_t CursorEntry{};
/** @} */
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
uint64_t Entry;
std::unordered_map<IR::NodeID, Label> JumpTargets;
@@ -141,41 +189,23 @@ private:
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
FEXCore::Core::DebugData *DebugData;
#ifdef BLOCKSTATS
bool GetSamplingData {true};
#endif
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
CodeBuffer InitialCodeBuffer{};
// This is the array of /additional/ code buffers that we may need to allocate
// Allocation only occurs when we've hit signals and need to clear code cache
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (X86JITCore::*)(const Label& label, LabelType type);
@@ -274,8 +304,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -284,7 +312,7 @@ private:
DEF_OP(Syscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -319,7 +347,7 @@ private:
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
DEF_OP(GuestOpcode);
DEF_OP(Fence);
DEF_OP(Break);
DEF_OP(Phi);
@@ -329,6 +357,7 @@ private:
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -434,6 +463,7 @@ private:
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
#undef DEF_OP
};
@@ -4,6 +4,7 @@ tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include <FEXCore/Core/CoreState.h>
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
@@ -20,6 +21,12 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - GuestEntry});
}
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -45,7 +52,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.OverflowExceptionHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)]);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
@@ -53,7 +60,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
break;
}
case FEXCore::IR::Break_Interrupt3: // INT3
@@ -65,7 +72,7 @@ DEF_OP(Break) {
}
// This jump target needs to be a constant offset here
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadPauseHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)]);
}
else {
// If we don't have a gdb server attached then....crash?
@@ -73,7 +80,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
}
break;
}
@@ -84,7 +91,7 @@ DEF_OP(Break) {
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.UnimplementedInstructionHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
@@ -130,13 +137,13 @@ DEF_OP(Print) {
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintValue)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)]);
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintVectorValue)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)]);
}
PopRegs();
@@ -166,6 +173,10 @@ DEF_OP(RDRAND) {
setc(Dst.second.cvt8());
}
DEF_OP(Yield) {
pause();
}
#undef DEF_OP
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -174,6 +185,7 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, GuestOpcode);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -184,6 +196,7 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
#undef REGISTER_OP
}
}
@@ -0,0 +1,139 @@
/*
$info$
tags: backend|x86-64
desc: relocation logic of the x86-64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
namespace FEXCore::CPU {
uint64_t X86JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void X86JITCore::LoadConstantWithPadding(Xbyak::Reg Reg, uint64_t Constant) {
// The maximum size a move constant can be in bytes
// Need to NOP pad to this size to ensure backpatching is always the same size
// Calculated as:
// [Rex]
// [Mov op]
// [8 byte constant]
//
// All other move types are smaller than this. xbyak will use a NOP slide which is quite quick
constexpr static size_t MAX_MOVE_SIZE = 10;
auto StartingOffset = getSize();
mov(Reg, Constant);
auto MoveSize = getSize() - StartingOffset;
auto NOPPadSize = MAX_MOVE_SIZE - MoveSize;
nop(NOPPadSize);
}
X86JITCore::NamedSymbolLiteralPair X86JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
NamedSymbolLiteralPair Lit {
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void X86JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = getSize();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CursorEntry;
uint64_t Pointer = GetNamedSymbolLiteral(Lit.MoveABI.NamedSymbolLiteral.Symbol);
L(Lit.Offset);
dq(Pointer);
Relocations.emplace_back(Lit.MoveABI);
}
void X86JITCore::InsertGuestRIPMove(Xbyak::Reg Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = getSize();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CursorEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.getIdx();
if (CTX->Config.CacheObjectCodeCompilation()) {
LoadConstantWithPadding(Reg, Constant);
}
else {
mov(Reg, Constant);
}
Relocations.emplace_back(MoveABI);
}
bool X86JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Place the pointer
dq(Pointer);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(CTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstantWithPadding(Xbyak::Reg64(Reloc->NamedThunkMove.RegisterIndex), Pointer);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstantWithPadding(Xbyak::Reg64(Reloc->GuestRIPMove.RegisterIndex), Pointer);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
return true;
}
}
+3 -8
View File
@@ -52,15 +52,8 @@ LookupCache::~LookupCache() {
FEXCore::Allocator::munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
}
void LookupCache::HintUsedRange(uint64_t Address, uint64_t Size) {
// Tell the kernel we will definitely need [Address, Address+Size) mapped for the page pointer
// Page Pointer is allocated per page, so shift by page size
Address >>= 12;
Size >>= 12;
madvise(reinterpret_cast<void*>(PagePointer + Address), Size, MADV_WILLNEED);
}
void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear out the page memory
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8, MADV_DONTNEED);
madvise(reinterpret_cast<void*>(PageMemory), CODE_SIZE, MADV_DONTNEED);
@@ -68,6 +61,8 @@ void LookupCache::ClearL2Cache() {
}
void LookupCache::ClearCache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear L1
madvise(reinterpret_cast<void*>(L1Pointer), L1_SIZE, MADV_DONTNEED);
// Clear L2
+76 -56
View File
@@ -7,6 +7,7 @@
#include <stddef.h>
#include <utility>
#include <vector>
#include <mutex>
namespace FEXCore {
namespace Context {
@@ -24,38 +25,73 @@ public:
LookupCache(FEXCore::Context::Context *CTX);
~LookupCache();
using LookupCacheIter = uintptr_t;
uintptr_t End() { return 0; }
uintptr_t FindBlock(uint64_t Address) {
auto HostCode = FindCodePointerForAddress(Address);
if (HostCode) {
return HostCode;
} else {
auto HostCode = BlockList.find(Address);
// Try L1, no lock needed
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
} else {
return 0;
// L2 and L3 need to be locked
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Try L2
const auto PageIndex = (Address & (VirtualMemSize -1)) >> 12;
const auto PageOffset = Address & (0x0FFF);
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
auto LocalPagePointer = Pointers[PageIndex];
// Do we a page pointer for this address?
if (LocalPagePointer) {
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == Address)
{
L1Entry.GuestCode = Address;
L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
return L1Entry.HostCode;
}
}
// Try L3
auto HostCode = BlockList.find(Address);
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
}
// Failed to find
return 0;
}
std::map<uint64_t, std::vector<uint64_t>> CodePages;
void AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto InsertPoint =
#endif
BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A_FMT(InsertPoint.second == true, "Dupplicate block mapping added");
// Appends Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
CodePages[CurrentPage].push_back(Address);
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length -1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto &CodePage = CodePages[CurrentPage];
rv |= CodePage.size() == 0;
CodePage.push_back(Address);
}
return rv;
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void *HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_A_FMT(Inserted, "Duplicate block mapping added");
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
@@ -65,6 +101,8 @@ public:
void Erase(uint64_t Address) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks.lower_bound({Address, 0});
auto upper = BlockLinks.upper_bound({Address, UINTPTR_MAX});
@@ -78,7 +116,10 @@ public:
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = L1Entry.HostCode = 0;
L1Entry.GuestCode = 0;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
}
// Do full map
@@ -101,14 +142,14 @@ public:
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks.insert({{GuestDestination, HostLink}, delinker});
}
void ClearCache();
void ClearL2Cache();
void HintUsedRange(uint64_t Address, uint64_t Size);
uintptr_t GetL1Pointer() const { return L1Pointer; }
uintptr_t GetPagePointer() const { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
@@ -116,8 +157,19 @@ public:
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::DebugStore,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
// This approach has not been fully vetted yet.
// Also note that L1 lookups might be inlined in the JIT Dispatcher and/or block ends.
std::recursive_mutex WriteLock;
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
@@ -167,38 +219,6 @@ private:
return PageMemory + NewBase;
}
uintptr_t FindCodePointerForAddress(uint64_t Address) {
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
auto FullAddress = Address;
Address = Address & (VirtualMemSize -1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t *Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// We don't have a page pointer for this address
return 0;
}
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == FullAddress)
{
L1Entry.GuestCode = FullAddress;
return L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
}
else
return 0;
}
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
@@ -0,0 +1,84 @@
#pragma once
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie{};
// Instructions per block configuration
int32_t MaxInstPerBlock{};
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Multiblock enabled
bool MultiBlock : 1;
// TSO enabled
bool TSOEnabled : 1;
// ABI local flag unsafe optimization
bool ABILocalFlags : 1;
// ABI no PF unsafe optimization
bool ABINoPF : 1;
// Static register allocation enabled
bool SRA : 1;
// Paranoid TSO mode enabled
bool ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
bool Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
bool x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 18;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
MaxInstPerBlock == other.MaxInstPerBlock &&
Arch == other.Arch &&
MultiBlock == other.MultiBlock &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
ABINoPF == other.ABINoPF &&
SRA == other.SRA &&
ParanoidTSO == other.ParanoidTSO &&
Is64BitMode == other.Is64BitMode &&
SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(CodeObjectSerializationConfig const &other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash{};
Hash <<= 32; Hash |= other.MaxInstPerBlock;
Hash <<= 1; Hash |= other.Arch;
Hash <<= 1; Hash |= other.MultiBlock;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.ABINoPF;
Hash <<= 1; Hash |= other.SRA;
Hash <<= 1; Hash |= other.ParanoidTSO;
Hash <<= 1; Hash |= other.Is64BitMode;
Hash <<= 2; Hash |= other.SMCChecks;
Hash <<= 1; Hash |= other.x87ReducedPrecision;
return Hash;
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is generated!");
}
@@ -0,0 +1,127 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <fcntl.h>
#include <filesystem>
#include <memory>
#include <string>
#include <sys/uio.h>
#include <sys/mman.h>
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
auto BaseFilename = std::filesystem::path(filename).filename().string();
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = std::make_unique<CodeRegionEntry>(
Base,
Size,
Offset,
filename,
NamedRegionHandler->DefaultCodeHeader(Base, Offset)
);
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
}
else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
// Removing a named region through the job system
// We need to find the entry that we are deleting first
std::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
// Take the pointer from the map
EntryPointer = std::move(it->second);
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Remove this from the entry map
EntryMap.erase(it);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
}
}
else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
}
void AsyncJobHandler::AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
}
@@ -0,0 +1,71 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::Context *ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ABINoPF = ctx->Config.ABINoPF;
DefaultSerializationConfig.SRA = ctx->Config.StaticRegisterAllocation;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion *>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion *>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
}
}
}
}
@@ -0,0 +1,85 @@
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <memory>
namespace {
static void* ThreadHandler(void *Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService *This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
}
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::Context *ctx)
: CTX {ctx}
, AsyncHandler { &NamedRegionHandler , this }
, NamedRegionHandler { ctx } {
Initialize();
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
WorkerThreadShuttingDown = true;
// Kick the working thread
WorkAvailable.NotifyAll();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, std::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
CodeObjectFileSection const *CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
char ThreadName[16] = "ObjectCodeSeri\0";
pthread_setname_np(pthread_self(), ThreadName);
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto &it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
}
@@ -0,0 +1,459 @@
#pragma once
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "Interface/Core/ObjectCache/CodeObjectSerializationConfig.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <map>
#include <memory>
#include <shared_mutex>
#include <string>
#include <vector>
#include <tsl/robin_map.h>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {
};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData *Data;
const char *HostCode;
uint64_t NumRelocations;
const char *Relocations;
};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase{};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset{};
// Total amount of code that should be in this file
uint64_t TotalCodeSize{};
// Used to reserve the TSL map
uint64_t NumCodeEntries{};
// The number of relocations that point to this section
uint64_t NumRelocationsTo{};
// Total relocations in this file
uint64_t TotalRelocationsCount{};
};
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base{};
// Size of this code entry
uint64_t Size{};
// The offset inside the file that is mapped to Base
uint64_t Offset{};
// Filename of the object
std::string Filename{};
CodeObjectSerializationHeader EntryHeader{};
/** @} */
// The filename of the object cache for this entry
std::string ObjectEntrySourceFilename{};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char *CodeData{};
size_t FileSize{};
std::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
tsl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap{};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base,
uint64_t Size,
uint64_t Offset,
std::string const &Filename,
CodeObjectSerializationHeader const &DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {
}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = std::map<uint64_t, std::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = std::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
void *HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex *ThreadJobRefCount;
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
std::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex *ObjectJobRefCountMutexPtr;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
AsyncJobHandler(NamedRegionObjectHandler *NamedRegionHandler, CodeObjectSerializeService *CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const { return Type; }
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry}
{}
const std::string BaseFilename;
const std::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, std::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
std::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler *NamedRegionHandler;
CodeObjectSerializeService *CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::Context *ctx);
void HandleNamedRegionObjectJobs();
CodeObjectSerializationConfig const &GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemAddNamedRegion> (
base,
filename,
executable,
entry
));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion> (
Base,
Size,
std::move(Entry)
));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs{};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex{};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
std::queue<std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue{};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::Context *ctx);
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(std::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex *ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
CodeObjectFileSection const *FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it);
CodeSerializationMutex &GetEntryMapMutex() { return EntryMapMutex; }
CodeSerializationMutex &GetUnrelocatedEntryMapMutex() { return EntryMapMutex; }
CodeRegionMapType &GetEntryMap() { return AddressToEntryMap; }
CodeRegionPtrMapType &GetUnrelocatedEntryMap() { return UnrelocatedAddressToEntryMap; }
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() { WorkAvailable.NotifyOne(); }
private:
FEXCore::Context::Context *CTX;
Event WorkAvailable{};
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
};
}
@@ -0,0 +1,78 @@
#pragma once
#include <FEXCore/IR/IR.h>
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header{};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header{};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header{};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
// The unrelocated RIP that is being moved
uint64_t GuestRIP;
};
union Relocation {
RelocationTypeHeader Header{};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
};
}
+405 -53
View File
@@ -1443,6 +1443,11 @@ void OpDispatchBuilder::XCHGOp(OpcodeArgs) {
// But this would result in a zext on 64bit, which would ruin the no-op nature of the instruction
// So x86-64 spec mandates this special case that even though it is a 32bit instruction and
// is supposed to zext the result, it is a true no-op
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX) {
// If this instruction has a REP prefix then this is architectually defined to be a `PAUSE` instruction.
// On older processors this ends up being a true `REP NOP` which is why they stuck this here.
_Yield();
}
return;
}
@@ -3353,7 +3358,9 @@ void OpDispatchBuilder::ReadSegmentReg(OpcodeArgs) {
template<OpDispatchBuilder::Segment Seg>
void OpDispatchBuilder::WriteSegmentReg(OpcodeArgs) {
auto Size = GetSrcSize(Op);
// Documentation claims that the 32-bit version of this instruction inserts in to the lower 32-bits of the segment
// This is incorrect and it instead zero extends the 32-bit value to 64-bit
auto Size = GetDstSize(Op);
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
if constexpr (Seg == Segment::FS) {
_StoreContext(Size, GPRClass, Src, offsetof(FEXCore::Core::CPUState, fs));
@@ -4604,7 +4611,7 @@ OrderedNode *OpDispatchBuilder::AppendSegmentOffset(OrderedNode *Value, uint32_t
return Value;
}
OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData, bool ForceLoad) {
OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData, bool ForceLoad, MemoryAccessType AccessType) {
LOGMAN_THROW_A_FMT(Operand.IsGPR() ||
Operand.IsLiteral() ||
Operand.IsGPRDirect() ||
@@ -4615,7 +4622,6 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
OrderedNode *Src {nullptr};
bool LoadableType = false;
bool StackAccess = false;
const uint8_t GPRSize = CTX->GetGPRSize();
const uint32_t AddrSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) != 0 ? (GPRSize >> 1) : GPRSize;
@@ -4644,7 +4650,9 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
else if (Operand.IsGPRDirect()) {
Src = _LoadContext(AddrSize, GPRClass, offsetof(FEXCore::Core::CPUState, gregs[Operand.Data.GPR.GPR]));
LoadableType = true;
StackAccess = Operand.Data.GPR.GPR == FEXCore::X86State::REG_RSP;
if (Operand.Data.GPR.GPR == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
else if (Operand.IsGPRIndirect()) {
auto GPR = _LoadContext(AddrSize, GPRClass, offsetof(FEXCore::Core::CPUState, gregs[Operand.Data.GPRIndirect.GPR]));
@@ -4653,7 +4661,9 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
Src = _Add(GPR, Constant);
LoadableType = true;
StackAccess = Operand.Data.GPRIndirect.GPR == FEXCore::X86State::REG_RSP;
if (Operand.Data.GPRIndirect.GPR == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
else if (Operand.IsRIPRelative()) {
if (CTX->Config.Is64BitMode) {
@@ -4675,7 +4685,9 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
auto Constant = _Constant(GPRSize * 8, Operand.Data.SIB.Scale);
Tmp = _Mul(Tmp, Constant);
}
StackAccess |= Operand.Data.SIB.Index == FEXCore::X86State::REG_RSP;
if (Operand.Data.SIB.Index == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
if (Operand.Data.SIB.Base != FEXCore::X86State::REG_INVALID) {
@@ -4687,7 +4699,10 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
else {
Tmp = GPR;
}
StackAccess |= Operand.Data.SIB.Base == FEXCore::X86State::REG_RSP;
if (Operand.Data.SIB.Base == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
if (Operand.Data.SIB.Offset) {
@@ -4722,7 +4737,7 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
if ((LoadableType && LoadData) || ForceLoad) {
Src = AppendSegmentOffset(Src, Flags);
if (StackAccess) {
if (AccessType == MemoryAccessType::ACCESS_NONTSO || AccessType == MemoryAccessType::ACCESS_STREAM) {
Src = _LoadMem(Class, OpSize, Src, Align == -1 ? OpSize : Align);
}
else {
@@ -4737,12 +4752,12 @@ OrderedNode *OpDispatchBuilder::GetRelocatedPC(FEXCore::X86Tables::DecodedOp con
return _EntrypointOffset(Op->PC + Op->InstSize + Offset - Entry, GPRSize);
}
OrderedNode *OpDispatchBuilder::LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData, bool ForceLoad) {
OrderedNode *OpDispatchBuilder::LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData, bool ForceLoad, MemoryAccessType AccessType) {
const uint8_t OpSize = GetSrcSize(Op);
return LoadSource_WithOpSize(Class, Op, Operand, OpSize, Flags, Align, LoadData, ForceLoad);
return LoadSource_WithOpSize(Class, Op, Operand, OpSize, Flags, Align, LoadData, ForceLoad, AccessType);
}
void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align) {
void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align, MemoryAccessType AccessType) {
LOGMAN_THROW_A_FMT(Operand.IsGPR() ||
Operand.IsLiteral() ||
Operand.IsGPRDirect() ||
@@ -4755,7 +4770,6 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
// 32bit ops ZEXT the result to 64bit
OrderedNode *MemStoreDst {nullptr};
bool MemStore = false;
bool StackAccess = false;
const uint8_t GPRSize = CTX->GetGPRSize();
const uint32_t AddrSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) != 0 ? (GPRSize >> 1) : GPRSize;
@@ -4789,7 +4803,9 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
else if (Operand.IsGPRDirect()) {
MemStoreDst = _LoadContext(AddrSize, GPRClass, offsetof(FEXCore::Core::CPUState, gregs[Operand.Data.GPR.GPR]));
MemStore = true;
StackAccess = Operand.Data.GPR.GPR == FEXCore::X86State::REG_RSP;
if (Operand.Data.GPR.GPR == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
else if (Operand.IsGPRIndirect()) {
auto GPR = _LoadContext(AddrSize, GPRClass, offsetof(FEXCore::Core::CPUState, gregs[Operand.Data.GPRIndirect.GPR]));
@@ -4797,7 +4813,9 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
MemStoreDst = _Add(GPR, Constant);
MemStore = true;
StackAccess = Operand.Data.GPRIndirect.GPR == FEXCore::X86State::REG_RSP;
if (Operand.Data.GPRIndirect.GPR == FEXCore::X86State::REG_RSP && AccessType == MemoryAccessType::ACCESS_DEFAULT) {
AccessType = MemoryAccessType::ACCESS_NONTSO;
}
}
else if (Operand.IsRIPRelative()) {
if (CTX->Config.Is64BitMode) {
@@ -4867,7 +4885,7 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
auto DestAddr = _Add(MemStoreDst, _Constant(8));
_StoreMem(GPRClass, 2, DestAddr, Upper, std::min<uint8_t>(Align, 8));
} else {
if (StackAccess) {
if (AccessType == MemoryAccessType::ACCESS_NONTSO || AccessType == MemoryAccessType::ACCESS_STREAM) {
_StoreMem(Class, OpSize, MemStoreDst, Src, Align == -1 ? OpSize : Align);
}
else {
@@ -4877,17 +4895,23 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
}
}
void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align) {
StoreResult_WithOpSize(Class, Op, Operand, Src, GetDstSize(Op), Align);
void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType) {
StoreResult_WithOpSize(Class, Op, Operand, Src, GetDstSize(Op), Align, AccessType);
}
void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align) {
StoreResult(Class, Op, Op->Dest, Src, Align);
void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType) {
StoreResult(Class, Op, Op->Dest, Src, Align, AccessType);
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::Context *ctx)
: CTX {ctx} {
: IREmitter {ctx->OpDispatcherAllocator}
, CTX {ctx} {
ResetWorkingList();
InstallHostSpecificOpcodeHandlers();
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator &Allocator)
: IREmitter {Allocator}
, CTX {nullptr} {
}
void OpDispatchBuilder::ResetWorkingList() {
@@ -4909,6 +4933,11 @@ void OpDispatchBuilder::MOVGPROp(OpcodeArgs) {
StoreResult(GPRClass, Op, Src, 1);
}
void OpDispatchBuilder::MOVGPRNTOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, 1);
StoreResult(GPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::ALUOp(OpcodeArgs) {
bool RequiresMask = false;
FEXCore::IR::IROps IROp;
@@ -5233,6 +5262,103 @@ void OpDispatchBuilder::InvalidOp(OpcodeArgs) {
#undef OpcodeArgs
void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
static bool Initialized = false;
if (!CTX || Initialized) {
// IRCompaction doesn't set a CTX and doesn't need this anyway
return;
}
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38_SHA[] = {
{OPD(PF_38_NONE, 0xC8), 1, &OpDispatchBuilder::SHA1NEXTEOp},
{OPD(PF_38_NONE, 0xC9), 1, &OpDispatchBuilder::SHA1MSG1Op},
{OPD(PF_38_NONE, 0xCA), 1, &OpDispatchBuilder::SHA1MSG2Op},
{OPD(PF_38_NONE, 0xCB), 1, &OpDispatchBuilder::SHA256RNDS2Op},
{OPD(PF_38_NONE, 0xCC), 1, &OpDispatchBuilder::SHA256MSG1Op},
{OPD(PF_38_NONE, 0xCD), 1, &OpDispatchBuilder::SHA256MSG2Op},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38_AES[] = {
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38_CRC[] = {
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
};
#undef OPD
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
#define PF_3A_NONE 0
#define PF_3A_66 1
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F3A_AES[] = {
{OPD(0, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
#undef PF_3A_NONE
#undef PF_3A_66
#undef OPD
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_6) << 5) | (prefix) << 3 | (Reg))
constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_66 = 2;
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryExtensionOp_RDRAND[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
};
#undef OPD
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOp_CLZero[] = {
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
auto InstallToTable = [](auto& FinalTable, auto& LocalTable) {
for (auto Op : LocalTable) {
auto OpNum = std::get<0>(Op);
auto Dispatcher = std::get<2>(Op);
for (uint8_t i = 0; i < std::get<1>(Op); ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].OpcodeDispatcher == nullptr, "Duplicate Entry");
FinalTable[OpNum + i].OpcodeDispatcher = Dispatcher;
}
}
};
if (CTX->HostFeatures.SupportsCRC) {
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38_CRC);
}
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38_SHA);
if (CTX->HostFeatures.SupportsAES) {
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38_AES);
InstallToTable(FEXCore::X86Tables::H0F3ATableOps, H0F3A_AES);
}
if (CTX->HostFeatures.SupportsCLZERO) {
InstallToTable(FEXCore::X86Tables::SecondModRMTableOps, SecondaryModRMExtensionOp_CLZero);
}
if (CTX->HostFeatures.SupportsRAND) {
InstallToTable(FEXCore::X86Tables::SecondInstGroupOps, SecondaryExtensionOp_RDRAND);
}
Initialized = true;
}
void InstallOpcodeHandlers(Context::OperatingMode Mode) {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable[] = {
// Instructions
@@ -5361,7 +5487,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xBD, 1, &OpDispatchBuilder::BSROp}, // BSF
{0xBE, 2, &OpDispatchBuilder::MOVSXOp},
{0xC0, 2, &OpDispatchBuilder::XADDOp},
{0xC3, 1, &OpDispatchBuilder::MOVGPROp<0>},
{0xC3, 1, &OpDispatchBuilder::MOVGPRNTOp},
{0xC4, 1, &OpDispatchBuilder::PINSROp<2>},
{0xC5, 1, &OpDispatchBuilder::PExtrOp<2>},
{0xC8, 8, &OpDispatchBuilder::BSWAPOp},
@@ -5375,7 +5501,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x17, 1, &OpDispatchBuilder::MOVUPSOp},
{0x28, 2, &OpDispatchBuilder::MOVUPSOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float<4, false>},
{0x2B, 1, &OpDispatchBuilder::MOVAPSOp},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>},
{0x2D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<4>},
@@ -5437,7 +5563,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xE3, 1, &OpDispatchBuilder::PAVGOp<2>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE7, 1, &OpDispatchBuilder::MOVUPSOp},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE8, 1, &OpDispatchBuilder::PSUBSOp<1, true>},
{0xE9, 1, &OpDispatchBuilder::PSUBSOp<2, true>},
{0xEA, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMIN, 2>},
@@ -5665,7 +5791,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x19, 7, &OpDispatchBuilder::NOPOp},
{0x28, 2, &OpDispatchBuilder::MOVAPSOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float<4, true>},
{0x2B, 1, &OpDispatchBuilder::MOVAPSOp},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<8>},
@@ -5739,7 +5865,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, false>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorOp},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE8, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSQSUB, 1>},
{0xE9, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSQSUB, 2>},
{0xEA, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMIN, 2>},
@@ -5794,15 +5920,8 @@ constexpr uint16_t PF_F2 = 3;
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
// GROUP 12
@@ -5841,10 +5960,6 @@ constexpr uint16_t PF_F2 = 3;
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_NONE, 6), 1, &OpDispatchBuilder::FenceOp<FEXCore::IR::Fence_LoadStore.Val>}, //MFENCE
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_NONE, 7), 1, &OpDispatchBuilder::StoreFenceOrCLFlush}, //SFENCE
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 0), 1, &OpDispatchBuilder::ReadSegmentReg<OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 1), 1, &OpDispatchBuilder::ReadSegmentReg<OpDispatchBuilder::Segment::GS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 2), 1, &OpDispatchBuilder::WriteSegmentReg<OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 3), 1, &OpDispatchBuilder::WriteSegmentReg<OpDispatchBuilder::Segment::GS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 5), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 6), 1, &OpDispatchBuilder::UnimplementedOp},
@@ -5860,6 +5975,15 @@ constexpr uint16_t PF_F2 = 3;
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_66, 0), 8, &OpDispatchBuilder::NOPOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_F2, 0), 8, &OpDispatchBuilder::NOPOp},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryExtensionOpTable_64[] = {
// GROUP 15
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 0), 1, &OpDispatchBuilder::ReadSegmentReg<OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 1), 1, &OpDispatchBuilder::ReadSegmentReg<OpDispatchBuilder::Segment::GS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 2), 1, &OpDispatchBuilder::WriteSegmentReg<OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 3), 1, &OpDispatchBuilder::WriteSegmentReg<OpDispatchBuilder::Segment::GS>},
};
#undef OPD
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOpTable[] = {
@@ -5868,13 +5992,242 @@ constexpr uint16_t PF_F2 = 3;
// REG /7
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
// Top bit indicating if it needs to be repeated with {0x40, 0x80} or'd in
// All OPDReg versions need it
#define OPDReg(op, reg) ((1 << 15) | ((op - 0xD8) << 8) | (reg << 3))
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> X87F64OpTable[] = {
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::FADDF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::FMULF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 2) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<32, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 3) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<32, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 4) | 0x00, 8, &OpDispatchBuilder::FSUBF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 5) | 0x00, 8, &OpDispatchBuilder::FSUBF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 6) | 0x00, 8, &OpDispatchBuilder::FDIVF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 7) | 0x00, 8, &OpDispatchBuilder::FDIVF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC0), 8, &OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC8), 8, &OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xD0), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xD8), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xE0), 8, &OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xE8), 8, &OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF0), 8, &OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF8), 8, &OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD9, 0) | 0x00, 8, &OpDispatchBuilder::FLDF64<32>},
// 1 = Invalid
{OPDReg(0xD9, 2) | 0x00, 8, &OpDispatchBuilder::FSTF64<32>},
{OPDReg(0xD9, 3) | 0x00, 8, &OpDispatchBuilder::FSTF64<32>},
{OPDReg(0xD9, 4) | 0x00, 8, &OpDispatchBuilder::X87LDENVF64},
{OPDReg(0xD9, 5) | 0x00, 8, &OpDispatchBuilder::X87FLDCWF64},
{OPDReg(0xD9, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSTENV},
{OPDReg(0xD9, 7) | 0x00, 8, &OpDispatchBuilder::X87FSTCW},
{OPD(0xD9, 0xC0), 8, &OpDispatchBuilder::FLDF64<80>},
{OPD(0xD9, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xD9, 0xD0), 1, &OpDispatchBuilder::NOPOp}, // FNOP
// D1 = Invalid
// D8 = Invalid
{OPD(0xD9, 0xE0), 1, &OpDispatchBuilder::FCHSF64},
{OPD(0xD9, 0xE1), 1, &OpDispatchBuilder::FABSF64},
// E2 = Invalid
{OPD(0xD9, 0xE4), 1, &OpDispatchBuilder::FTSTF64},
{OPD(0xD9, 0xE5), 1, &OpDispatchBuilder::X87FXAMF64},
// E6 = Invalid
{OPD(0xD9, 0xE8), 1, &OpDispatchBuilder::FLDF64_Const<0x3FF0000000000000>}, // 1.0
{OPD(0xD9, 0xE9), 1, &OpDispatchBuilder::FLDF64_Const<0x400A934F0979A372>}, // log2l(10)
{OPD(0xD9, 0xEA), 1, &OpDispatchBuilder::FLDF64_Const<0x3FF71547652B82FE>}, // log2l(e)
{OPD(0xD9, 0xEB), 1, &OpDispatchBuilder::FLDF64_Const<0x400921FB54442D18>}, // pi
{OPD(0xD9, 0xEC), 1, &OpDispatchBuilder::FLDF64_Const<0x3FD34413509F79FF>}, // log10l(2)
{OPD(0xD9, 0xED), 1, &OpDispatchBuilder::FLDF64_Const<0x3FE62E42FEFA39EF>}, // log(2)
{OPD(0xD9, 0xEE), 1, &OpDispatchBuilder::FLDF64_Const<0>}, // 0.0
// EF = Invalid
{OPD(0xD9, 0xF0), 1, &OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64F2XM1>},
{OPD(0xD9, 0xF1), 1, &OpDispatchBuilder::X87FYL2XF64},
{OPD(0xD9, 0xF2), 1, &OpDispatchBuilder::X87TANF64},
{OPD(0xD9, 0xF3), 1, &OpDispatchBuilder::X87ATANF64},
{OPD(0xD9, 0xF4), 1, &OpDispatchBuilder::FXTRACTF64},
{OPD(0xD9, 0xF5), 1, &OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM1>},
{OPD(0xD9, 0xF6), 1, &OpDispatchBuilder::X87ModifySTP<false>},
{OPD(0xD9, 0xF7), 1, &OpDispatchBuilder::X87ModifySTP<true>},
{OPD(0xD9, 0xF8), 1, &OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM>},
{OPD(0xD9, 0xF9), 1, &OpDispatchBuilder::X87FYL2XF64},
{OPD(0xD9, 0xFA), 1, &OpDispatchBuilder::FSQRTF64},
{OPD(0xD9, 0xFB), 1, &OpDispatchBuilder::X87SinCosF64},
{OPD(0xD9, 0xFC), 1, &OpDispatchBuilder::FRNDINTF64},
{OPD(0xD9, 0xFD), 1, &OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64SCALE>},
{OPD(0xD9, 0xFE), 1, &OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64SIN>},
{OPD(0xD9, 0xFF), 1, &OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64COS>},
{OPDReg(0xDA, 0) | 0x00, 8, &OpDispatchBuilder::FADDF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 1) | 0x00, 8, &OpDispatchBuilder::FMULF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 2) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<32, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 3) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<32, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 4) | 0x00, 8, &OpDispatchBuilder::FSUBF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 5) | 0x00, 8, &OpDispatchBuilder::FSUBF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 6) | 0x00, 8, &OpDispatchBuilder::FDIVF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 7) | 0x00, 8, &OpDispatchBuilder::FDIVF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDA, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
// E8 = Invalid
{OPD(0xDA, 0xE9), 1, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
// EA = Invalid
// F0 = Invalid
// F8 = Invalid
{OPDReg(0xDB, 0) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDB, 1) | 0x00, 8, &OpDispatchBuilder::FISTF64<true>},
{OPDReg(0xDB, 2) | 0x00, 8, &OpDispatchBuilder::FISTF64<false>},
{OPDReg(0xDB, 3) | 0x00, 8, &OpDispatchBuilder::FISTF64<false>},
// 4 = Invalid
{OPDReg(0xDB, 5) | 0x00, 8, &OpDispatchBuilder::FLDF64<80>},
// 6 = Invalid
{OPDReg(0xDB, 7) | 0x00, 8, &OpDispatchBuilder::FSTF64<80>},
{OPD(0xDB, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
{OPD(0xDB, 0xE2), 1, &OpDispatchBuilder::NOPOp}, // FNCLEX
{OPD(0xDB, 0xE3), 1, &OpDispatchBuilder::FNINITF64},
// E4 = Invalid
{OPD(0xDB, 0xE8), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDB, 0xF0), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
// F8 = Invalid
{OPDReg(0xDC, 0) | 0x00, 8, &OpDispatchBuilder::FADDF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 1) | 0x00, 8, &OpDispatchBuilder::FMULF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 2) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<64, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 3) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<64, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 4) | 0x00, 8, &OpDispatchBuilder::FSUBF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 5) | 0x00, 8, &OpDispatchBuilder::FSUBF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 6) | 0x00, 8, &OpDispatchBuilder::FDIVF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 7) | 0x00, 8, &OpDispatchBuilder::FDIVF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDC, 0xC0), 8, &OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xC8), 8, &OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE0), 8, &OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE8), 8, &OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF0), 8, &OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF8), 8, &OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDD, 0) | 0x00, 8, &OpDispatchBuilder::FLDF64<64>},
{OPDReg(0xDD, 1) | 0x00, 8, &OpDispatchBuilder::FISTF64<true>},
{OPDReg(0xDD, 2) | 0x00, 8, &OpDispatchBuilder::FSTF64<64>},
{OPDReg(0xDD, 3) | 0x00, 8, &OpDispatchBuilder::FSTF64<64>},
{OPDReg(0xDD, 4) | 0x00, 8, &OpDispatchBuilder::X87FRSTORF64},
// 5 = Invalid
{OPDReg(0xDD, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSAVEF64},
{OPDReg(0xDD, 7) | 0x00, 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDD, 0xC0), 8, &OpDispatchBuilder::X87FFREE},
{OPD(0xDD, 0xD0), 8, &OpDispatchBuilder::FST}, //register-register from regular X87
{OPD(0xDD, 0xD8), 8, &OpDispatchBuilder::FST}, //^
{OPD(0xDD, 0xE0), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDD, 0xE8), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 0) | 0x00, 8, &OpDispatchBuilder::FADDF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 1) | 0x00, 8, &OpDispatchBuilder::FMULF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 2) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<16, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 3) | 0x00, 8, &OpDispatchBuilder::FCOMIF64<16, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 4) | 0x00, 8, &OpDispatchBuilder::FSUBF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 5) | 0x00, 8, &OpDispatchBuilder::FSUBF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 6) | 0x00, 8, &OpDispatchBuilder::FDIVF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 7) | 0x00, 8, &OpDispatchBuilder::FDIVF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDE, 0xC0), 8, &OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xC8), 8, &OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xD9), 1, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
{OPD(0xDE, 0xE0), 8, &OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xE8), 8, &OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF0), 8, &OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF8), 8, &OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDF, 0) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDF, 1) | 0x00, 8, &OpDispatchBuilder::FISTF64<true>},
{OPDReg(0xDF, 2) | 0x00, 8, &OpDispatchBuilder::FISTF64<false>},
{OPDReg(0xDF, 3) | 0x00, 8, &OpDispatchBuilder::FISTF64<false>},
{OPDReg(0xDF, 4) | 0x00, 8, &OpDispatchBuilder::FBLDF64},
{OPDReg(0xDF, 5) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDF, 6) | 0x00, 8, &OpDispatchBuilder::FBSTPF64},
{OPDReg(0xDF, 7) | 0x00, 8, &OpDispatchBuilder::FISTF64<false>},
// XXX: This should also set the x87 tag bits to empty
// We don't support this currently, so just pop the stack
{OPD(0xDF, 0xC0), 8, &OpDispatchBuilder::X87ModifySTP<true>},
{OPD(0xDF, 0xE0), 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDF, 0xE8), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDF, 0xF0), 8, &OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> X87OpTable[] = {
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::FADD<32, false, OpDispatchBuilder::OpResult::RES_ST0>},
@@ -6110,7 +6463,6 @@ constexpr uint16_t PF_F2 = 3;
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38Table[] = {
@@ -6156,7 +6508,7 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<4, 8, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<4, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 8>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVAPSOp},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<4>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<1, 2, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<1, 4, false>},
@@ -6176,24 +6528,13 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMUL, 4>},
{OPD(PF_38_66, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_66, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66, 0xF6), 1, &OpDispatchBuilder::ADXOp},
{OPD(PF_38_F3, 0xF6), 1, &OpDispatchBuilder::ADXOp},
};
#undef OPD
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
@@ -6225,8 +6566,9 @@ constexpr uint16_t PF_F2 = 3;
{OPD(0, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<4>},
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(0, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
{OPD(0, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
};
#undef PF_3A_NONE
#undef PF_3A_66
@@ -6308,6 +6650,8 @@ constexpr uint16_t PF_F2 = 3;
{OPD(2, 0b10, 0xF7), 1, &OpDispatchBuilder::BMI2Shift},
{OPD(2, 0b11, 0xF7), 1, &OpDispatchBuilder::BMI2Shift},
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::VPCLMULQDQOp},
{OPD(3, 0b11, 0xF0), 1, &OpDispatchBuilder::RORX},
};
#undef OPD
@@ -6372,10 +6716,18 @@ constexpr uint16_t PF_F2 = 3;
InstallToTable(FEXCore::X86Tables::RepNEModOps, RepNEModOpTable);
InstallToTable(FEXCore::X86Tables::OpSizeModOps, OpSizeModOpTable);
InstallToTable(FEXCore::X86Tables::SecondInstGroupOps, SecondaryExtensionOpTable);
if (Mode == Context::MODE_64BIT) {
InstallToTable(FEXCore::X86Tables::SecondInstGroupOps, SecondaryExtensionOpTable_64);
}
InstallToTable(FEXCore::X86Tables::SecondModRMTableOps, SecondaryModRMExtensionOpTable);
InstallToX87Table(FEXCore::X86Tables::X87Ops, X87OpTable);
FEX_CONFIG_OPT(ReducedPrecision, X87REDUCEDPRECISION);
if(ReducedPrecision) {
InstallToX87Table(FEXCore::X86Tables::X87Ops, X87F64OpTable);
} else {
InstallToX87Table(FEXCore::X86Tables::X87Ops, X87OpTable);
}
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38Table);
InstallToTable(FEXCore::X86Tables::H0F3ATableOps, H0F3ATable);
+87 -8
View File
@@ -76,8 +76,10 @@ public:
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
// Used during new op bringup
bool ShouldDump {false};
struct JumpTargetInfo {
OrderedNode* BlockEntry;
bool HaveEmitted;
@@ -148,6 +150,7 @@ public:
}
OpDispatchBuilder(FEXCore::Context::Context *ctx);
OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
void ResetWorkingList();
void ResetDecodeFailure() { DecodeFailure = false; }
@@ -161,7 +164,9 @@ public:
void UnhandledOp(OpcodeArgs);
template<uint32_t SrcIndex>
void MOVGPROp(OpcodeArgs);
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
void ALUOp(OpcodeArgs);
void INTOp(OpcodeArgs);
void SyscallOp(OpcodeArgs);
@@ -463,6 +468,57 @@ public:
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMI(OpcodeArgs);
// F64 X87 Ops
template<size_t width>
void FLDF64(OpcodeArgs);
template<uint64_t num>
void FLDF64_Const(OpcodeArgs);
void FBLDF64(OpcodeArgs);
void FBSTPF64(OpcodeArgs);
void FILDF64(OpcodeArgs);
template<size_t width>
void FSTF64(OpcodeArgs);
void FSTF64(OpcodeArgs);
template<bool Truncate>
void FISTF64(OpcodeArgs);
template<size_t width, bool Integer, OpResult ResInST0>
void FADDF64(OpcodeArgs);
template<size_t width, bool Integer, OpResult ResInST0>
void FMULF64(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpResult ResInST0>
void FDIVF64(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpResult ResInST0>
void FSUBF64(OpcodeArgs);
void FCHSF64(OpcodeArgs);
void FABSF64(OpcodeArgs);
void FTSTF64(OpcodeArgs);
void FRNDINTF64(OpcodeArgs);
void FXTRACTF64(OpcodeArgs);
void FNINITF64(OpcodeArgs);
void FSQRTF64(OpcodeArgs);
template<FEXCore::IR::IROps IROp>
void X87UnaryOpF64(OpcodeArgs);
template<FEXCore::IR::IROps IROp>
void X87BinaryOpF64(OpcodeArgs);
void X87SinCosF64(OpcodeArgs);
void X87FLDCWF64(OpcodeArgs);
void X87FYL2XF64(OpcodeArgs);
void X87TANF64(OpcodeArgs);
void X87ATANF64(OpcodeArgs);
void X87FNSAVEF64(OpcodeArgs);
void X87FRSTORF64(OpcodeArgs);
void X87FXAMF64(OpcodeArgs);
void X87LDENVF64(OpcodeArgs);
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMIF64(OpcodeArgs);
void FXSaveOp(OpcodeArgs);
void FXRStoreOp(OpcodeArgs);
@@ -535,6 +591,15 @@ public:
void PSADBW(OpcodeArgs);
void SHA1NEXTEOp(OpcodeArgs);
void SHA1MSG1Op(OpcodeArgs);
void SHA1MSG2Op(OpcodeArgs);
void SHA1RNDS4Op(OpcodeArgs);
void SHA256MSG1Op(OpcodeArgs);
void SHA256MSG2Op(OpcodeArgs);
void SHA256RNDS2Op(OpcodeArgs);
void AESImcOp(OpcodeArgs);
void AESEncOp(OpcodeArgs);
void AESEncLastOp(OpcodeArgs);
@@ -558,6 +623,8 @@ public:
void DPPOp(OpcodeArgs);
void MPSADBWOp(OpcodeArgs);
void PCLMULQDQOp(OpcodeArgs);
void VPCLMULQDQOp(OpcodeArgs);
void CRC32(OpcodeArgs);
@@ -580,12 +647,22 @@ private:
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
enum class MemoryAccessType {
// Choose TSO or Non-TSO depending on access type
ACCESS_DEFAULT,
// TSO access behaviour
ACCESS_TSO,
// Non-TSO access behaviour
ACCESS_NONTSO,
// Non-temporal streaming
ACCESS_STREAM,
};
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
[[nodiscard]] static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
@@ -1163,18 +1240,20 @@ private:
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
if (CTX->IsTSOEnabled())
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
if (CTX->IsTSOEnabled())
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
void InstallHostSpecificOpcodeHandlers();
};
void InstallOpcodeHandlers(Context::OperatingMode Mode);
@@ -10,13 +10,259 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
#include <stdint.h>
#include <array>
#include <cstdint>
#include <tuple>
#include <utility>
namespace FEXCore::IR {
class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Tmp = _Ror(_VExtractToGPR(16, 4, Dest, 3), _Constant(32, 2));
auto Top = _Add(_VExtractToGPR(16, 4, Src, 3), Tmp);
auto Result = _VInsGPR(16, 4, 3, Src, Top);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W0 = _VExtractToGPR(16, 4, Dest, 3);
auto W1 = _VExtractToGPR(16, 4, Dest, 2);
auto W2 = _VExtractToGPR(16, 4, Dest, 1);
auto W3 = _VExtractToGPR(16, 4, Dest, 0);
auto W4 = _VExtractToGPR(16, 4, Src, 3);
auto W5 = _VExtractToGPR(16, 4, Src, 2);
auto D3 = _VInsGPR(16, 4, 3, Dest, _Xor(W2, W0));
auto D2 = _VInsGPR(16, 4, 2, D3, _Xor(W3, W1));
auto D1 = _VInsGPR(16, 4, 1, D2, _Xor(W4, W2));
auto D0 = _VInsGPR(16, 4, 0, D1, _Xor(W5, W3));
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// ROR by 31 is equivalent to a ROL by 1
auto ThirtyOne = _Constant(32, 31);
auto W13 = _VExtractToGPR(16, 4, Src, 2);
auto W14 = _VExtractToGPR(16, 4, Src, 1);
auto W15 = _VExtractToGPR(16, 4, Src, 0);
auto W16 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 3), W13), ThirtyOne);
auto W17 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 2), W14), ThirtyOne);
auto W18 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 1), W15), ThirtyOne);
auto W19 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 0), W16), ThirtyOne);
auto D3 = _VInsGPR(16, 4, 3, Dest, W16);
auto D2 = _VInsGPR(16, 4, 2, D3, W17);
auto D1 = _VInsGPR(16, 4, 1, D2, W18);
auto D0 = _VInsGPR(16, 4, 0, D1, W19);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(),
"Src1 needs to be literal here to indicate function and constants");
using FnType = OrderedNode* (*)(OpDispatchBuilder&, OrderedNode*, OrderedNode*, OrderedNode*);
const auto f0 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._And(B, C), Self._And(Self._Not(B), D));
};
const auto f1 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(B, C), D);
};
const auto f2 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(Self._And(B, C), Self._And(B, D)), Self._And(C, D));
};
const auto f3 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(B, C), D);
};
constexpr std::array<uint32_t, 4> k_array{
0x5A827999U,
0x6ED9EBA1U,
0x8F1BBCDCU,
0xCA62C1D6U,
};
constexpr std::array<FnType, 4> fn_array{
f0, f1, f2, f3,
};
const uint64_t Imm8 = Op->Src[1].Data.Literal.Value & 0b11;
const FnType Fn = fn_array[Imm8];
auto K = _Constant(32, k_array[Imm8]);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W0E = _VExtractToGPR(16, 4, Src, 3);
auto W1 = _VExtractToGPR(16, 4, Src, 2);
auto W2 = _VExtractToGPR(16, 4, Src, 1);
auto W3 = _VExtractToGPR(16, 4, Src, 0);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(16, 4, Dest, 3);
auto B = _VExtractToGPR(16, 4, Dest, 2);
auto C = _VExtractToGPR(16, 4, Dest, 1);
auto D = _VExtractToGPR(16, 4, Dest, 0);
auto A1 = _Add(_Add(_Add(Fn(*this, B, C, D), _Ror(A, _Constant(32, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(B, _Constant(32, 2));
auto D1 = C;
auto E1 = D;
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C,
OrderedNode *D, OrderedNode *E, OrderedNode *W) -> RoundResult {
auto ANext = _Add(_Add(_Add(_Add(Fn(*this, B, C, D), _Ror(A, _Constant(32, 27))), W), E), K);
auto BNext = A;
auto CNext = _Ror(B, _Constant(32, 2));
auto DNext = C;
auto ENext = D;
return {ANext, BNext, CNext, DNext, ENext};
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, W1);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, W2);
auto Final = Round1To3(A3, B3, C3, D3, E3, W3);
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(16, 4, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(16, 4, 1, Dest2, std::get<2>(Final));
auto Dest0 = _VInsGPR(16, 4, 0, Dest1, std::get<3>(Final));
StoreResult(FPRClass, Op, Dest0, -1);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
const auto Sigma0 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(_Xor(_Ror(W, _Constant(32, 7)), _Ror(W, _Constant(32, 18))), _Lshr(W, _Constant(32, 3)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W4 = _VExtractToGPR(16, 4, Src, 0);
auto W3 = _VExtractToGPR(16, 4, Dest, 3);
auto W2 = _VExtractToGPR(16, 4, Dest, 2);
auto W1 = _VExtractToGPR(16, 4, Dest, 1);
auto W0 = _VExtractToGPR(16, 4, Dest, 0);
auto Sig3 = _Add(W3, Sigma0(W4));
auto Sig2 = _Add(W2, Sigma0(W3));
auto Sig1 = _Add(W1, Sigma0(W2));
auto Sig0 = _Add(W0, Sigma0(W1));
auto D3 = _VInsGPR(16, 4, 3, Dest, Sig3);
auto D2 = _VInsGPR(16, 4, 2, D3, Sig2);
auto D1 = _VInsGPR(16, 4, 1, D2, Sig1);
auto D0 = _VInsGPR(16, 4, 0, D1, Sig0);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(_Xor(_Ror(W, _Constant(32, 17)), _Ror(W, _Constant(32, 19))), _Lshr(W, _Constant(32, 10)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W14 = _VExtractToGPR(16, 4, Src, 2);
auto W15 = _VExtractToGPR(16, 4, Src, 3);
auto W16 = _Add(_VExtractToGPR(16, 4, Dest, 0), Sigma1(W14));
auto W17 = _Add(_VExtractToGPR(16, 4, Dest, 1), Sigma1(W15));
auto W18 = _Add(_VExtractToGPR(16, 4, Dest, 2), Sigma1(W16));
auto W19 = _Add(_VExtractToGPR(16, 4, Dest, 3), Sigma1(W17));
auto D3 = _VInsGPR(16, 4, 3, Dest, W19);
auto D2 = _VInsGPR(16, 4, 2, D3, W18);
auto D1 = _VInsGPR(16, 4, 1, D2, W17);
auto D0 = _VInsGPR(16, 4, 0, D1, W16);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
const auto Ch = [this](OrderedNode *E, OrderedNode *F, OrderedNode *G) -> OrderedNode* {
return _Xor(_And(E, F), _And(_Not(E), G));
};
const auto Major = [this](OrderedNode *A, OrderedNode *B, OrderedNode *C) -> OrderedNode* {
return _Xor(_Xor(_And(A, B), _And(A, C)), _And(B, C));
};
const auto Sigma0 = [this](OrderedNode *A) -> OrderedNode* {
return _Xor(_Xor(_Ror(A, _Constant(32, 2)), _Ror(A, _Constant(32, 13))), _Ror(A, _Constant(32, 22)));
};
const auto Sigma1 = [this](OrderedNode *E) -> OrderedNode* {
return _Xor(_Xor(_Ror(E, _Constant(32, 6)), _Ror(E, _Constant(32, 11))), _Ror(E, _Constant(32, 25)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *XMM0 = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
auto C0 = _VExtractToGPR(16, 4, Dest, 3);
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto E0 = _VExtractToGPR(16, 4, Src, 1);
auto F0 = _VExtractToGPR(16, 4, Src, 0);
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
auto H0 = _VExtractToGPR(16, 4, Dest, 0);
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*,
OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
const auto Round = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C, OrderedNode *D,
OrderedNode *E, OrderedNode *F, OrderedNode *G, OrderedNode *H,
OrderedNode* WK) -> RoundResult {
auto ANext = _Add(_Add(_Add(_Add(_Add(Ch(E, F, G), Sigma1(E)), WK), H), Major(A, B, C)), Sigma0(A));
auto BNext = A;
auto CNext = B;
auto DNext = C;
auto ENext = _Add(_Add(_Add(_Add(Ch(E, F, G), Sigma1(E)), WK), H), D);
auto FNext = E;
auto GNext = F;
auto HNext = G;
return {ANext, BNext, CNext, DNext, ENext, FNext, GNext, HNext};
};
auto [A1, B1, C1, D1, E1, F1, G1, H1] = Round(A0, B0, C0, D0, E0, F0, G0, H0, WK0);
auto Final = Round(A1, B1, C1, D1, E1, F1, G1, H1, WK1);
auto Res3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Res2 = _VInsGPR(16, 4, 2, Res3, std::get<1>(Final));
auto Res1 = _VInsGPR(16, 4, 1, Res2, std::get<4>(Final));
auto Res0 = _VInsGPR(16, 4, 0, Res1, std::get<5>(Final));
StoreResult(FPRClass, Op, Res0, -1);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESImc(Src);
@@ -60,4 +306,26 @@ void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Selector needs to be literal here");
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
auto Res = _PCLMUL(Dest, Src, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[2].IsLiteral(), "Selector needs to be literal here");
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
auto Res = _PCLMUL(Src1, Src2, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
}
@@ -28,6 +28,11 @@ void OpDispatchBuilder::MOVVectorOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Src, 1);
}
void OpDispatchBuilder::MOVVectorNTOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1, true, false, MemoryAccessType::ACCESS_STREAM);
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::MOVAPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
StoreResult(FPRClass, Op, Src, -1);
@@ -750,7 +755,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs) {
}
else {
// If we are storing to memory then we store the size of the element extracted
StoreResult(GPRClass, Op, Result, -1);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, ElementSize, -1);
}
}
@@ -1194,7 +1199,8 @@ void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
{
auto DestByte = _Bfe(8, 8 * Select, DestElement);
auto MemLocation = _Add(MemDest, _Constant(Element * 8 + Select));
_StoreMemAutoTSO(GPRClass, 1, MemLocation, DestByte, 1);
// MASKMOVDQU/MASKMOVQ is explicitly weakly-ordered on its store
_StoreMem(GPRClass, 1, MemLocation, DestByte, 1);
}
auto Jump = _Jump();
auto NextJumpTarget = CreateNewCodeBlockAfter(StoreBlock);
@@ -1809,7 +1815,6 @@ void OpDispatchBuilder::VPFCMPOp(OpcodeArgs) {
}
StoreResult(FPRClass, Op, Result, -1);
ShouldDump = true;
}
template
@@ -621,8 +621,8 @@ void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037
auto NewFCW = _Constant(16, 0x037);
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
_F80LoadFCW(NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
File diff suppressed because it is too large. Load diff
+5 -4
View File
@@ -95,10 +95,11 @@ namespace FEXCore {
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", FHU::Syscalls::gettid());
}
else {
if (Handler.Handler &&
Handler.Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
for (auto &Handler : Handler.Handlers) {
if (Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
}
}
if (Handler.FrontendHandler &&
@@ -167,7 +167,7 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2, nullptr}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0xCA, 2, X86InstInfo{"RETF", TYPE_PRIV, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_DEBUG, 0, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, FLAGS_DEBUG , 1, nullptr}},
@@ -87,6 +87,14 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC8), 1, X86InstInfo{"SHA1NEXTE", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC9), 1, X86InstInfo{"SHA1MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCA), 1, X86InstInfo{"SHA1MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCB), 1, X86InstInfo{"SHA256RNDS2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCC), 1, X86InstInfo{"SHA256MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCD), 1, X86InstInfo{"SHA256MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDB), 1, X86InstInfo{"AESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDC), 1, X86InstInfo{"AESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDD), 1, X86InstInfo{"AESENCLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -31,7 +31,7 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_8BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -42,13 +42,15 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
};
@@ -343,8 +343,8 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_15, PF_F3, 0), 1, X86InstInfo{"RDFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 1), 1, X86InstInfo{"RDGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"WRFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"WRFSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 5), 1, X86InstInfo{"INCSSPQ", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"CLRSSBSY", TYPE_INST, FLAGS_NONE, 0, nullptr}},
@@ -441,7 +441,7 @@ void InitializeVEXTables() {
{OPD(3, 0b01, 0x40), 1, X86InstInfo{"VDPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x42), 1, X86InstInfo{"VMPSADBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x48), 1, X86InstInfo{"VPERMILzz2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
+127
View File
@@ -0,0 +1,127 @@
#include "GDBJIT.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/Utils/LogManager.h>
#if defined(GDB_SYMBOLS_ENABLED)
#include <FEXCore/Debug/GDBReaderInterface.h>
extern "C" {
enum jit_actions_t { JIT_NOACTION = 0, JIT_REGISTER_FN, JIT_UNREGISTER_FN };
struct jit_code_entry {
jit_code_entry *next_entry;
jit_code_entry *prev_entry;
const char *symfile_addr;
uint64_t symfile_size;
};
struct jit_descriptor {
uint32_t version;
/* This type should be jit_actions_t, but we use uint32_t
to be explicit about the bitwidth. */
uint32_t action_flag;
jit_code_entry *relevant_entry;
jit_code_entry *first_entry;
};
/* Make sure to specify the version statically, because the
debugger may check the version before we can set it. */
constinit jit_descriptor __jit_debug_descriptor = {.version = 1};
/* GDB puts a breakpoint in this function. */
void __attribute__((noinline)) __jit_debug_register_code() {
asm volatile("" ::"r"(&__jit_debug_descriptor));
};
}
namespace FEXCore {
void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart,
uint64_t GuestRIP, uintptr_t HostEntry,
FEXCore::Core::DebugData *DebugData) {
auto map = Entry->SourcecodeMap.get();
if (map) {
auto FileOffset = GuestRIP - VAFileStart;
auto Sym = map->FindSymbolMapping(FileOffset);
std::string SymName = HLE::SourcecodeSymbolMapping::SymName(
Sym, Entry->Filename, HostEntry, FileOffset);
std::vector<gdb_line_mapping> Lines;
for (const auto &GuestOpcode : DebugData->GuestOpcodes) {
auto Line = map->FindLineMapping(GuestRIP + GuestOpcode.GuestEntryOffset -
VAFileStart);
if (Line) {
Lines.push_back(
{Line->LineNumber, HostEntry + GuestOpcode.HostEntryOffset});
}
}
size_t size = sizeof(info_t) + 1 * sizeof(blocks_t) +
Lines.size() * sizeof(gdb_line_mapping);
auto mem = (uint8_t *)malloc(size);
auto base = mem;
info_t *info = (info_t *)mem;
mem += sizeof(info_t);
strncpy(info->filename, map->SourceFile.c_str(), 511);
info->nblocks = 1;
auto blocks = (blocks_t *)mem;
info->blocks_ofs = mem - base;
mem += info->nblocks * sizeof(blocks_t);
for (int i = 0; i < info->nblocks; i++) {
strncpy(blocks[i].name, SymName.c_str(), 511);
blocks[i].start = HostEntry;
blocks[i].end = HostEntry + DebugData->HostCodeSize;
}
info->nlines = Lines.size();
auto lines = (gdb_line_mapping *)mem;
info->lines_ofs = mem - base;
mem += info->nlines * sizeof(gdb_line_mapping);
if (info->nlines) {
memcpy(lines, &Lines.at(0), info->nlines * sizeof(gdb_line_mapping));
}
auto entry = new jit_code_entry{0, 0, 0, 0};
entry->symfile_addr = (const char *)info;
entry->symfile_size = size;
if (__jit_debug_descriptor.first_entry) {
__jit_debug_descriptor.relevant_entry->next_entry = entry;
entry->prev_entry = __jit_debug_descriptor.relevant_entry;
} else {
__jit_debug_descriptor.first_entry = entry;
}
__jit_debug_descriptor.relevant_entry = entry;
__jit_debug_descriptor.action_flag = JIT_REGISTER_FN;
__jit_debug_register_code();
}
}
} // namespace FEXCore
#else
namespace FEXCore {
void GDBJITRegister([[maybe_unused]] FEXCore::IR::AOTIRCacheEntry *Entry,
[[maybe_unused]] uintptr_t VAFileStart,
[[maybe_unused]] uint64_t GuestRIP,
[[maybe_unused]] uintptr_t HostEntry,
[[maybe_unused]] FEXCore::Core::DebugData *DebugData) {
ERROR_AND_DIE_FMT("GDBSymbols support not compiled in");
}
} // namespace FEXCore
#endif
+7
View File
@@ -0,0 +1,7 @@
#include <Interface/IR/AOTIR.h>
namespace FEXCore {
void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry, FEXCore::Core::DebugData *DebugData);
}
+258 -14
View File
@@ -10,14 +10,18 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include "FEXCore/Utils/CompilerDefs.h"
#include "Thunks.h"
#include <cstdint>
#include <dlfcn.h>
#include <Interface/Context/Context.h>
#include "FEXCore/Core/X86Enums.h"
#include <malloc.h>
#include <map>
#include <mutex>
#include <unordered_map>
#include <memory>
#include <shared_mutex>
#include <stdint.h>
@@ -26,26 +30,106 @@ $end_info$
struct LoadlibArgs {
const char *Name;
uintptr_t CallbackThunks;
};
static thread_local FEXCore::Core::InternalThreadState *Thread;
static __attribute__((aligned(16), naked, section("HostToGuestTrampolineTemplate"))) void HostToGuestTrampolineTemplate() {
#if defined(_M_X86_64)
asm(
"lea 0f(%rip), %r11 \n"
"jmpq *0f(%rip) \n"
".align 8 \n"
"0: \n"
".quad 0, 0, 0, 0 \n" // TrampolineInstanceInfo
);
#elif defined(_M_ARM_64)
asm(
"adr x11, 0f \n"
"ldr x16, [x11] \n"
"br x16 \n"
// Manually align to the next 8-byte boundary
// NOTE: GCC over-aligns to a full page when using .align directives on ARM (last tested on GCC 11.2)
"nop \n"
"0: \n"
".quad 0, 0, 0, 0 \n" // TrampolineInstanceInfo
);
#else
#error Unsupported host architecture
#endif
}
extern char __start_HostToGuestTrampolineTemplate[];
extern char __stop_HostToGuestTrampolineTemplate[];
namespace FEXCore {
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
struct TrampolineInstanceInfo {
uintptr_t HostPacker;
uintptr_t CallCallback;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
};
struct GuestcallInfo {
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
bool operator==(const GuestcallInfo&) const noexcept = default;
};
struct GuestcallInfoHash {
size_t operator()(const GuestcallInfo& x) const noexcept {
// Hash only the target address, which is generally unique.
// For the unlikely case of a hash collision, std::unordered_map still picks the correct bucket entry.
return std::hash<uintptr_t>{}(x.GuestTarget);
}
};
// Bits in a SHA256 sum are already randomly distributed, so truncation yields a suitable hash function
struct TruncatingSHA256Hash {
size_t operator()(const FEXCore::IR::SHA256Sum& SHA256Sum) const noexcept {
return (const size_t&)SHA256Sum;
}
};
class ThunkHandler_impl final: public ThunkHandler {
std::shared_mutex ThunksMutex;
std::map<IR::SHA256Sum, ThunkedFunction*> Thunks = {
std::unordered_map<IR::SHA256Sum, ThunkedFunction*, TruncatingSHA256Hash> Thunks = {
{
// sha256(fex:loadlib)
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80},
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80 },
&LoadLib
},
{
// sha256(fex:is_lib_loaded)
{ 0xee, 0x57, 0xba, 0x0c, 0x5f, 0x6e, 0xef, 0x2a, 0x8c, 0xb5, 0x19, 0x81, 0xc9, 0x23, 0xe6, 0x51, 0xae, 0x65, 0x02, 0x8f, 0x2b, 0x5d, 0x59, 0x90, 0x6a, 0x7e, 0xe2, 0xe7, 0x1c, 0x33, 0x8a, 0xff },
&IsLibLoaded
},
{
// sha256(fex:link_address_to_function)
{ 0xe6, 0xa8, 0xec, 0x1c, 0x7b, 0x74, 0x35, 0x27, 0xe9, 0x4f, 0x5b, 0x6e, 0x2d, 0xc9, 0xa0, 0x27, 0xd6, 0x1f, 0x2b, 0x87, 0x8f, 0x2d, 0x35, 0x50, 0xea, 0x16, 0xb8, 0xc4, 0x5e, 0x42, 0xfd, 0x77 },
&LinkAddressToGuestFunction
},
{
// sha256(fex:make_host_trampoline_for_guest_function)
{ 0x1e, 0x51, 0x6b, 0x07, 0x39, 0xeb, 0x50, 0x59, 0xb3, 0xf3, 0x4f, 0xca, 0xdd, 0x58, 0x37, 0xe9, 0xf0, 0x30, 0xe5, 0x89, 0x81, 0xc7, 0x14, 0xfb, 0x24, 0xf9, 0xba, 0xe7, 0x0e, 0x00, 0x1e, 0x86 },
&MakeHostTrampolineForGuestFunction
}
};
// Can't be a string_view. We need to keep a copy of the library name in-case string_view pointer goes away.
// Ideally we track when a library has been unloaded and remove it from this set before the memory backing goes away.
std::set<std::string> Libs;
std::unordered_map<GuestcallInfo, uintptr_t, GuestcallInfoHash> GuestcallToHostTrampoline;
uint8_t *HostTrampolineInstanceDataPtr;
size_t HostTrampolineInstanceDataAvailable = 0;
/*
Set arg0/1 to arg regs, use CTX::HandleCallback to handle the callback
*/
@@ -56,13 +140,160 @@ namespace FEXCore {
Thread->CTX->HandleCallback(Thread, (uintptr_t)callback);
}
/**
* Instructs the Core to redirect calls to functions at the given
* address to another function. The original callee address is passed
* to the target function through an implicit argument stored in r11.
*
* The primary use case of this is ensuring that host function pointers
* returned from thunked APIs can safely be called by the guest.
*/
static void LinkAddressToGuestFunction(void* argsv) {
struct args_t {
uintptr_t original_callee;
uintptr_t target_addr; // Guest function to call when branching to original_callee
};
auto args = reinterpret_cast<args_t*>(argsv);
auto CTX = Thread->CTX;
LOGMAN_THROW_A_FMT(args->original_callee, "Tried to link null pointer address to guest function");
LOGMAN_THROW_A_FMT(args->target_addr, "Tried to link address to null pointer guest function");
if (!CTX->Config.Is64BitMode) {
LOGMAN_THROW_A_FMT((args->original_callee >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_A_FMT((args->target_addr >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
}
LogMan::Msg::DFmt("Thunks: Adding guest trampoline from address {:#x} to guest function {:#x}",
args->original_callee, args->target_addr);
auto Result = Thread->CTX->AddCustomIREntrypoint(
args->original_callee,
[CTX, GuestThunkEntrypoint = args->target_addr](uintptr_t Entrypoint, FEXCore::IR::IREmitter *emit) {
auto IRHeader = emit->_IRHeader(emit->Invalid(), 0);
auto Block = emit->CreateCodeNode();
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
const uint8_t GPRSize = CTX->GetGPRSize();
emit->_StoreContext(GPRSize, IR::GPRClass, emit->_Constant(Entrypoint), offsetof(Core::CPUState, gregs[X86State::REG_R11]));
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
}, CTX->ThunkHandler.get(), (void*)args->target_addr);
if (!Result) {
if (Result.Creator != CTX->ThunkHandler.get()) {
ERROR_AND_DIE_FMT("Input address for LinkAddressToGuestFunction is already linked by another module");
}
if (Result.Data != (void*)args->target_addr) {
// NOTE: This may happen in Vulkan thunks if the Vulkan driver resolves two different symbols
// to the same function (e.g. vkGetPhysicalDeviceFeatures2/vkGetPhysicalDeviceFeatures2KHR)
LogMan::Msg::EFmt("Input address for LinkAddressToGuestFunction is already linked elsewhere");
}
}
}
/**
* Generates a host-callable trampoline to call guest functions via the host ABI.
*
* This trampoline uses the same calling convention as the given HostPacker. Trampolines
* are cached, so it's safe to call this function repeatedly on the same arguments without
* leaking memory.
*
* Invoking the returned trampoline has the effect of:
* - packing the arguments (using the HostPacker identified by its SHA256)
* - performing a host->guest transition
* - unpacking the arguments via GuestUnpacker
* - calling the function at GuestTarget
*
* The primary use case of this is ensuring that guest function pointers ("callbacks")
* passed to thunked APIs can safely be called by the native host library.
*/
static void MakeHostTrampolineForGuestFunction(void* ArgsRV) {
struct ArgsRV_t {
IR::SHA256Sum *HostPackerSha256;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
uintptr_t rv; // Pointer to host trampoline + TrampolineInstanceInfo
} *args = reinterpret_cast<ArgsRV_t*>(ArgsRV);
LOGMAN_THROW_A_FMT(args->GuestTarget, "Tried to create host-trampoline to null pointer guest function");
const auto CTX = Thread->CTX;
const auto ThunkHandler = reinterpret_cast<ThunkHandler_impl *>(CTX->ThunkHandler.get());
const GuestcallInfo gci = { args->GuestUnpacker, args->GuestTarget };
// Try first with shared_lock
{
std::shared_lock lk(ThunkHandler->ThunksMutex);
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
std::lock_guard lk(ThunkHandler->ThunksMutex);
// Retry lookup with full lock before making a new trampoline to avoid double trampolines
{
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
// No entry found => create new trampoline
auto HostPackerEntry = ThunkHandler->Thunks.find(*args->HostPackerSha256);
if (HostPackerEntry == ThunkHandler->Thunks.end()) {
ERROR_AND_DIE_FMT("Unknown host packing function for callback");
}
LogMan::Msg::DFmt("Thunks: Adding host trampoline for guest function {:#x}",
args->GuestTarget);
const auto Length = __stop_HostToGuestTrampolineTemplate - __start_HostToGuestTrampolineTemplate;
const auto InstanceInfoOffset = Length - sizeof(TrampolineInstanceInfo);
if (ThunkHandler->HostTrampolineInstanceDataAvailable < Length) {
const auto allocation_step = 16 * 1024;
ThunkHandler->HostTrampolineInstanceDataAvailable = allocation_step;
ThunkHandler->HostTrampolineInstanceDataPtr = (uint8_t *)mmap(
0, ThunkHandler->HostTrampolineInstanceDataAvailable,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
LOGMAN_THROW_A_FMT(ThunkHandler->HostTrampolineInstanceDataPtr != MAP_FAILED, "Failed to mmap HostTrampolineInstanceDataPtr");
}
const TrampolineInstanceInfo NewTrampolineInfo {
.HostPacker = reinterpret_cast<uintptr_t>(HostPackerEntry->second),
.CallCallback = (uintptr_t)&CallCallback,
.GuestUnpacker = args->GuestUnpacker,
.GuestTarget = args->GuestTarget
};
uint8_t* const HostTrampoline = ThunkHandler->HostTrampolineInstanceDataPtr;
ThunkHandler->HostTrampolineInstanceDataAvailable -= Length;
ThunkHandler->HostTrampolineInstanceDataPtr += Length;
memcpy(HostTrampoline, (void*)&HostToGuestTrampolineTemplate, Length);
memcpy(HostTrampoline + InstanceInfoOffset, &NewTrampolineInfo, sizeof(NewTrampolineInfo));
args->rv = reinterpret_cast<uintptr_t>(HostTrampoline);
ThunkHandler->GuestcallToHostTrampoline[gci] = args->rv;
}
static void LoadLib(void *ArgsV) {
auto CTX = Thread->CTX;
auto Args = reinterpret_cast<LoadlibArgs*>(ArgsV);
auto Name = Args->Name;
auto CallbackThunks = Args->CallbackThunks;
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
@@ -75,13 +306,13 @@ namespace FEXCore {
const auto InitSym = std::string("fexthunks_exports_") + Name;
ExportEntry* (*InitFN)(void *, uintptr_t);
ExportEntry* (*InitFN)();
(void*&)InitFN = dlsym(Handle, InitSym.c_str());
if (!InitFN) {
ERROR_AND_DIE_FMT("LoadLib: Failed to find export {}", InitSym);
}
auto Exports = InitFN((void*)&CallCallback, CallbackThunks);
auto Exports = InitFN();
if (!Exports) {
ERROR_AND_DIE_FMT("LoadLib: Failed to initialize thunk library {}. "
"Check if the corresponding host library is installed "
@@ -91,7 +322,9 @@ namespace FEXCore {
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
{
std::unique_lock lk(That->ThunksMutex);
std::lock_guard lk(That->ThunksMutex);
That->Libs.insert(Name);
int i;
for (i = 0; Exports[i].sha256; i++) {
@@ -102,6 +335,23 @@ namespace FEXCore {
}
}
static void IsLibLoaded(void* ArgsRV) {
struct ArgsRV_t {
const char *Name;
bool rv;
};
auto &[Name, rv] = *reinterpret_cast<ArgsRV_t*>(ArgsRV);
auto CTX = Thread->CTX;
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
{
std::shared_lock lk(That->ThunksMutex);
rv = That->Libs.contains(Name);
}
}
public:
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) {
@@ -120,12 +370,6 @@ namespace FEXCore {
void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) {
::Thread = Thread;
}
ThunkHandler_impl() {
}
~ThunkHandler_impl() {
}
};
ThunkHandler* ThunkHandler::Create() {
+86 -79
View File
@@ -4,6 +4,9 @@
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/Core/LookupCache.h>
#include <Interface/GDBJIT/GDBJIT.h>
#include <cstddef>
#include <cstdint>
@@ -15,6 +18,7 @@
#include <unistd.h>
#include <xxhash.h>
namespace FEXCore::IR {
AOTIRInlineEntry *AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
@@ -77,7 +81,7 @@ namespace FEXCore::IR {
return true;
}
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd) {
static bool LoadAOTIRCache(AOTIRCacheEntry *Entry, int streamfd) {
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != FEXCore::IR::AOTIR_COOKIE)
@@ -99,6 +103,10 @@ namespace FEXCore::IR {
if (!readAll(streamfd, (char*)&Module[0], Module.size()))
return false;
if (Entry->FileId != Module) {
return false;
}
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize)))
@@ -119,19 +127,16 @@ namespace FEXCore::IR {
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
AOTIRCache->insert({Module, {Array, FilePtr, Size}});
LOGMAN_THROW_A_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
Entry->Array = Array;
Entry->FilePtr = FilePtr;
Entry->Size = Size;
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
}
AOTIRCaptureCache::~AOTIRCaptureCache() {
for (auto &Mod: AOTIRCache) {
FEXCore::Allocator::munmap(Mod.second.mapping, Mod.second.size);
}
}
void AOTIRCaptureCache::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
@@ -230,39 +235,27 @@ namespace FEXCore::IR {
void AOTIRCaptureCache::WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
std::shared_lock lk(AOTIRCacheLock);
for( const auto &File: FilesWithCode) {
Writer(File.first, File.second);
for( const auto &Entry: AOTIRCache) {
if (Entry.second.ContainsCode) {
Writer(Entry.second.FileId, Entry.second.Filename);
}
}
}
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
{
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
if (!file->second.ContainsCode) {
file->second.ContainsCode = true;
FilesWithCode[file->second.fileid] = file->second.filename;
}
}
}
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
PreGenerateIRFetchResult Result{};
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
auto Mod = (FEXCore::IR::AOTIRInlineIndex*)file->second.CachedFileEntry;
if (Mod == nullptr) {
file->second.CachedFileEntry = Mod = AOTIRCache[file->second.fileid].Array;
}
if (AOTIRCacheEntry.Entry) {
AOTIRCacheEntry.Entry->ContainsCode = true;
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
auto Mod = AOTIRCacheEntry.Entry->Array;
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - file->second.Start + file->second.Offset);
auto AOTEntry = Mod->Find(GuestRIP - AOTIRCacheEntry.VAFileStart);
if (AOTEntry) {
// verify hash
@@ -272,7 +265,7 @@ namespace FEXCore::IR {
Result.IRList = AOTEntry->GetIRData();
//LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
Result.RAData = AOTEntry->GetRAData();;
Result.RAData = AOTEntry->GetRAData()->CreateCopy();
Result.DebugData = new FEXCore::Core::DebugData();
Result.StartAddr = MappedStart;
Result.Length = AOTEntry->GuestLength;
@@ -296,21 +289,23 @@ namespace FEXCore::IR {
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::RegisterAllocationData::UniquePtr RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount) {
bool GeneratedIR) {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
if (GeneratedIR || CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
auto file = FindAddrForFile(StartAddr, Length);
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (AOTIRCacheEntry.Entry) {
if (DebugData && CTX->Config.LibraryJITNaming()) {
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, AOTIRCacheEntry.Entry->Filename);
}
if (CTX->Config.GDBSymbols()) {
GDBJITRegister(AOTIRCacheEntry.Entry, AOTIRCacheEntry.VAFileStart, GuestRIP, (uintptr_t)CodePtr, DebugData);
}
// Add to AOT cache if aot generation is enabled
@@ -319,29 +314,35 @@ namespace FEXCore::IR {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCacheMap[fileid];
auto LocalRIP = GuestRIP - AOTIRCacheEntry.VAFileStart;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.VAFileStart;
auto FileId = AOTIRCacheEntry.Entry->FileId;
// The underlying pointer and the unique_ptr deleter for RAData must
// be marshalled separately to the lambda below. Otherwise, the
// lambda can't be used as an std::function due to being non-copyable
auto RADataCopy = RAData->CreateCopy();
auto RADataCopyDeleter = RADataCopy.get_deleter();
auto IRListCopy = IRList->CreateCopy();
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy=RADataCopy.release(), RADataCopyDeleter, FileId]() {
// It is guaranteed via AOTIRCaptureCacheWriteoutLock and AOTIRCaptureCacheWriteoutFlusing that this will not run concurrently
// Memory coherency is guaranteed via AOTIRCaptureCacheWriteoutLock
auto *AotFile = &AOTIRCaptureCacheMap[FileId];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
AotFile->Stream = AOTIRWriter(FileId);
uint64_t tag = FEXCore::IR::AOTIR_COOKIE;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy);
RADataCopyDeleter(RADataCopy);
delete IRListCopy;
});
if (CTX->Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount) {
--Thread->CompileBlockReentrantRefCount;
}
Thread->CPUBackend->ClearCache();
return true;
}
}
@@ -349,30 +350,25 @@ namespace FEXCore::IR {
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
if (CTX->GetGdbServerStatus()) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), std::move(RAData), decltype(Entry.DebugData)(DebugData)};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->DebugStore.insert({GuestRIP, std::move(Entry)});
}
else {
// If the IR doesn't need to be retained then we can just delete it now
delete DebugData;
if (IRList->IsCopy()) delete IRList;
}
}
}
return false;
}
AOTIRCaptureCache::AddrToFileMapType::iterator AOTIRCaptureCache::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
void AOTIRCaptureCache::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
AOTIRCacheEntry *AOTIRCaptureCache::LoadAOTIRCacheEntry(const std::string &filename) {
auto base_filename = std::filesystem::path(filename).filename().string();
if (!base_filename.empty()) {
@@ -388,21 +384,32 @@ namespace FEXCore::IR {
std::unique_lock lk(AOTIRCacheLock);
AddrToFile.insert({ Base, { Base, Size, Offset, fileid, filename, nullptr, false} });
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry { .FileId = fileid, .Filename = filename }});
auto Entry = &(Inserted.first->second);
if (CTX->Config.AOTIRLoad && !AOTIRCache.contains(fileid) && AOTIRLoader) {
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
if (CTX->Config.AOTIRLoad && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
FEXCore::IR::LoadAOTIRCache(&AOTIRCache, streamfd);
FEXCore::IR::LoadAOTIRCache(Entry, streamfd);
close(streamfd);
}
}
return Entry;
}
return nullptr;
}
void AOTIRCaptureCache::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
std::unique_lock lk(AOTIRCacheLock);
// TODO: Support partial removing
AddrToFile.erase(Base);
void AOTIRCaptureCache::UnloadAOTIRCacheEntry(AOTIRCacheEntry *Entry) {
LOGMAN_THROW_A_FMT(Entry != nullptr, "Removing not existing entry");
if (Entry->Array) {
FEXCore::Allocator::munmap(Entry->FilePtr, Entry->Size);
Entry->Array = nullptr;
Entry->FilePtr = nullptr;
Entry->Size = 0;
}
}
}
+13 -26
View File
@@ -1,5 +1,6 @@
#pragma once
#include "FEXCore/IR/RegisterAllocationData.h"
#include <FEXCore/Config/Config.h>
#include <atomic>
@@ -11,6 +12,7 @@
#include <unordered_map>
#include <shared_mutex>
#include <queue>
#include <FEXCore/HLE/SourcecodeResolver.h>
namespace FEXCore::Core {
struct DebugData;
@@ -72,18 +74,20 @@ namespace FEXCore::IR {
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
void *FilePtr;
size_t Size;
std::unique_ptr<FEXCore::HLE::SourcecodeMap> SourcecodeMap;
std::string FileId;
std::string Filename;
bool ContainsCode;
};
using AOTCacheType = std::unordered_map<std::string, FEXCore::IR::AOTIRCacheEntry>;
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd);
class AOTIRCaptureCache final {
public:
AOTIRCaptureCache(FEXCore::Context::Context *ctx) : CTX {ctx} {}
~AOTIRCaptureCache();
void FinalizeAOTIRCache();
void AOTIRCaptureCacheWriteoutQueue_Flush();
@@ -92,7 +96,7 @@ namespace FEXCore::IR {
struct PreGenerateIRFetchResult {
FEXCore::IR::IRListView *IRList {};
FEXCore::IR::RegisterAllocationData *RAData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
FEXCore::Core::DebugData *DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
@@ -105,14 +109,13 @@ namespace FEXCore::IR {
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::RegisterAllocationData::UniquePtr RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount);
bool GeneratedIR);
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(AOTIRCacheEntry *Entry);
// Callbacks
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
@@ -136,27 +139,11 @@ namespace FEXCore::IR {
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
std::map<std::string, std::string> FilesWithCode;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
FEXCore::IR::AOTCacheType AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ofstream>(const std::string&)> AOTIRWriter;
std::function<void(const std::string&)> AOTIRRenamer;
std::unordered_map<std::string, FEXCore::IR::AOTIRCaptureCacheEntry> AOTIRCaptureCacheMap;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
};
}
+59 -16
View File
@@ -181,13 +181,18 @@
"RAOverride": "0"
},
"GuestOpcode u32:$GuestEntryOffset": {
"Desc": ["Marks the beginning of a guest opcode"],
"HasSideEffects": true
},
"GPR = ValidateCode u64:$CodeOriginalLow, u64:$CodeOriginalhigh, i64:$Offset, u8:$CodeLength": {
"HasSideEffects": true,
"HasDest": true,
"DestSize": "8"
},
"RemoveCodeEntry": {
"RemoveThreadCodeEntry": {
"HasSideEffects": true
},
@@ -231,6 +236,11 @@
],
"DestSize": "16",
"NumElements": "2"
},
"Yield": {
"HasSideEffects": true,
"Desc": ["This is a hint instruction that the CPU is likely to do a spin so it might want to pause to help out SMP",
"Can be implemented as a NOP if necessary"]
}
},
"Branch": {
@@ -289,12 +299,6 @@
],
"DestSize": "16",
"NumElements": "2"
},
"GuestCallDirect u64:$RIP, u64:$NextRIP": {
"HasSideEffects": true
},
"GuestCallIndirect GPR:$RIP, u64:$NextRIP": {
"HasSideEffects": true
}
},
"Moves": {
@@ -938,7 +942,7 @@
},
"FPR = VBitcast u8:#RegisterSize, u8:#ElementSize, FPR:$Source": {
"Dest": ["Workaround for issue with LLVM breaking when loading scalar elements to vectors"],
"Desc": ["Workaround for issue with LLVM breaking when loading scalar elements to vectors"],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
@@ -1272,12 +1276,12 @@
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VUMull2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Dest": "Multiplies the high elements with size extension",
"Desc": "Multiplies the high elements with size extension",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VSMull2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Dest": "Multiplies the high elements with size extension",
"Desc": "Multiplies the high elements with size extension",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
@@ -1449,35 +1453,74 @@
},
"Crypto": {
"FPR = VAESImc FPR:$Vector": {
"Dest": "Does a stage of the inverse mix column transformation",
"Desc": "Does a stage of the inverse mix column transformation",
"DestSize": "16"
},
"FPR = VAESEnc FPR:$State, FPR:$Key": {
"Dest": "Does a step of AES encryption",
"Desc": "Does a step of AES encryption",
"DestSize": "16"
},
"FPR = VAESEncLast FPR:$State, FPR:$Key": {
"Dest": "Does the last step of AES encryption",
"Desc": "Does the last step of AES encryption",
"DestSize": "16"
},
"FPR = VAESDec FPR:$State, FPR:$Key": {
"Dest": "Does a step of AES decryption",
"Desc": "Does a step of AES decryption",
"DestSize": "16"
},
"FPR = VAESDecLast FPR:$State, FPR:$Key": {
"Dest": "Does the last step of AES decryption",
"Desc": "Does the last step of AES decryption",
"DestSize": "16"
},
"FPR = VAESKeyGenAssist FPR:$Src, u8:$RCON": {
"Dest": "Assists in key generation",
"Desc": "Assists in key generation",
"DestSize": "16"
},
"GPR = CRC32 GPR:$Src1, GPR:$Src2, u8:$SrcSize": {
"Desc": ["CRC32 using polynomial 0x1EDC6F41"
],
"DestSize": "std::max<uint8_t>(4, GetOpSize(_Src1))"
},
"FPR = PCLMUL FPR:$Src1, FPR:$Src2, u8:$Selector": {
"Desc": [
"Performs carryless multiplication of 64-bit elements depending on the selector.",
"Selector = 0b00000000: Uses low 64-bit elements from both input vectors",
"Selector = 0b00000001: Uses high 64-bit element from Src1 and low 64-bit element from Src2",
"Selector = 0b00010000: Uses low 64-bit element from Src1 and high 64-bit element from Src2",
"Selector = 0b00010001: Uses high 64-bit elements from both input vectors"
],
"DestSize": "16"
}
},
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "8"
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "8"
},
"FPR = F64FPREM1 FPR:$Src1, FPR:$Src2": {
"DestSize": "8"
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "8"
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "8"
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "8"
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "8"
},
"FPR = F64SIN FPR:$Src": {
"DestSize": "8"
},
"FPR = F64COS FPR:$Src": {
"DestSize": "8"
}
},
"F80": {
"F80LoadFCW GPR:$Src": {
"HasSideEffects": true
+21
View File
@@ -16,6 +16,27 @@ $end_info$
#include <vector>
namespace FEXCore::IR {
bool IsFragmentExit(FEXCore::IR::IROps Op) {
switch (Op) {
case OP_EXITFUNCTION:
case OP_BREAK:
return true;
default:
return false;
}
}
bool IsBlockExit(FEXCore::IR::IROps Op) {
switch(Op) {
case OP_JUMP:
case OP_CONDJUMP:
return true;
default:
return IsFragmentExit(Op);
}
}
FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(OrderedNode *Node) {
auto Class = GetOpRegClass(Node);
switch (Class) {
+17 -36
View File
@@ -5,6 +5,8 @@ tags: ir|parser
$end_info$
*/
#include "Common/StringUtils.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
@@ -42,28 +44,6 @@ enum class DecodeFailure {
};
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string DecodeErrorToString(DecodeFailure Failure) {
switch (Failure) {
case DecodeFailure::DECODE_OKAY: return "Okay";
@@ -295,7 +275,7 @@ class IRParser: public FEXCore::IR::IREmitter {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
// Strip off the type qualifier from the ssa value
std::string SSAName = trim(Arg);
std::string SSAName = FEXCore::StringUtils::Trim(Arg);
const size_t ArgEnd = SSAName.find_first_of(' ');
if (ArgEnd != std::string::npos) {
@@ -329,7 +309,8 @@ class IRParser: public FEXCore::IR::IREmitter {
LineDefinition *CurrentDef{};
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
IRParser(std::istream *text) {
IRParser(FEXCore::Utils::IntrusivePooledAllocator &ThreadAllocator, std::istream *text)
: IREmitter {ThreadAllocator} {
InitializeNameMap();
std::string TmpLine;
@@ -382,7 +363,7 @@ class IRParser: public FEXCore::IR::IREmitter {
CurrentDef = &Def;
Def.LineNumber = i;
Line = trim(Line);
Line = FEXCore::StringUtils::Trim(Line);
// Skip empty lines
if (Line.empty()) {
@@ -401,7 +382,7 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of('=', CurrentPos)) != std::string::npos) {
Def.Definition = Line.substr(0, DefinitionEnd);
Def.Definition = trim(Def.Definition);
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition);
Def.HasDefinition = true;
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
}
@@ -421,7 +402,7 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t SSAEnd = std::string::npos;
if ((SSAEnd = Line.find_last_of(' ', DefinitionEnd)) != std::string::npos) {
std::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
Type = trim(Type);
Type = FEXCore::StringUtils::Trim(Type);
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) {
@@ -430,7 +411,7 @@ class IRParser: public FEXCore::IR::IREmitter {
Def.Size = DefinitionSize.second;
}
Def.Definition = trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
Def.Definition = FEXCore::StringUtils::Trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
CurrentPos = DefinitionEnd + 1;
}
@@ -447,8 +428,8 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t NameEnd = std::string::npos;
if ((NameEnd = Def.Definition.find_first_of(' ')) != std::string::npos) {
std::string Type = Def.Definition.substr(NameEnd + 1);
Type = trim(Type);
Def.Definition = trim(Def.Definition.substr(0, NameEnd));
Type = FEXCore::StringUtils::Trim(Type);
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition.substr(0, NameEnd));
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
@@ -465,11 +446,11 @@ class IRParser: public FEXCore::IR::IREmitter {
// Let's get the IR op
size_t OpNameEnd = std::string::npos;
std::string RemainingLine = trim(Line.substr(CurrentPos));
std::string RemainingLine = FEXCore::StringUtils::Trim(Line.substr(CurrentPos));
CurrentPos = 0;
if ((OpNameEnd = RemainingLine.find_first_of(" \t\n\r\0", CurrentPos)) != std::string::npos) {
Def.IROp = RemainingLine.substr(CurrentPos, OpNameEnd);
Def.IROp = trim(Def.IROp);
Def.IROp = FEXCore::StringUtils::Trim(Def.IROp);
Def.HasArgs = true;
CurrentPos = OpNameEnd;
}
@@ -486,7 +467,7 @@ class IRParser: public FEXCore::IR::IREmitter {
}
if (Def.HasArgs) {
RemainingLine = trim(RemainingLine.substr(CurrentPos));
RemainingLine = FEXCore::StringUtils::Trim(RemainingLine.substr(CurrentPos));
CurrentPos = 0;
if (RemainingLine.empty()) {
// How did we get here?
@@ -495,7 +476,7 @@ class IRParser: public FEXCore::IR::IREmitter {
else {
while (!RemainingLine.empty()) {
const size_t ArgEnd = RemainingLine.find(',');
std::string Arg = trim(RemainingLine.substr(0, ArgEnd));
std::string Arg = FEXCore::StringUtils::Trim(RemainingLine.substr(0, ArgEnd));
Def.Args.emplace_back(std::move(Arg));
@@ -680,8 +661,8 @@ class IRParser: public FEXCore::IR::IREmitter {
} // anon namespace
std::unique_ptr<IREmitter> Parse(std::istream *in) {
auto parser = std::make_unique<IRParser>(in);
std::unique_ptr<IREmitter> Parse(FEXCore::Utils::IntrusivePooledAllocator &ThreadAllocator, std::istream *in) {
auto parser = std::make_unique<IRParser>(ThreadAllocator, in);
if (parser->Loaded) {
return parser;
+4 -3
View File
@@ -6,6 +6,7 @@ desc: Defines which passes are run, and runs them
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -15,7 +16,7 @@ $end_info$
namespace FEXCore::IR {
class IREmitter;
void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllocation) {
void PassManager::AddDefaultPasses(FEXCore::Context::Context *ctx, bool InlineConstants, bool StaticRegisterAllocation) {
FEX_CONFIG_OPT(DisablePasses, O0);
if (!DisablePasses()) {
@@ -29,7 +30,7 @@ void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllo
InsertPass(CreateDeadStoreElimination());
InsertPass(CreatePassDeadCodeElimination());
InsertPass(CreateConstProp(InlineConstants));
InsertPass(CreateConstProp(InlineConstants, ctx->HostFeatures.SupportsTSOImm9));
////// InsertPass(CreateDeadFlagCalculationEliminination());
@@ -48,7 +49,7 @@ void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllo
// If the IR is compacted post-RA then the node indexing gets messed up and the backend isn't able to find the register assigned to a node
// Compact before IR, don't worry about RA generating spills/fills
InsertPass(CreateIRCompaction(), "Compaction");
InsertPass(CreateIRCompaction(ctx->OpDispatcherAllocator), "Compaction");
}
void PassManager::AddDefaultValidationPasses() {
+2 -1
View File
@@ -7,6 +7,7 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <functional>
#include <memory>
@@ -39,7 +40,7 @@ protected:
class PassManager final {
friend class SyscallOptimization;
public:
void AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllocation);
void AddDefaultPasses(FEXCore::Context::Context *ctx, bool InlineConstants, bool StaticRegisterAllocation);
void AddDefaultValidationPasses();
Pass* InsertPass(std::unique_ptr<Pass> Pass, std::string Name = "") {
Pass->RegisterPassManager(this);
+6 -2
View File
@@ -2,18 +2,22 @@
#include <memory>
namespace FEXCore::Utils {
class IntrusivePooledAllocator;
}
namespace FEXCore::IR {
class Pass;
class RegisterAllocationPass;
class RegisterAllocationData;
std::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool InlineConstants);
std::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool InlineConstants, bool SupportsTSOImm9);
std::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination();
std::unique_ptr<FEXCore::IR::Pass> CreateSyscallOptimization();
std::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
std::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination();
std::unique_ptr<FEXCore::IR::Pass> CreatePassDeadCodeElimination();
std::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction();
std::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
std::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass, bool OptimizeSRA);
std::unique_ptr<FEXCore::IR::Pass> CreateStaticRegisterAllocationPass();
std::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass();
Loaded 100 of 444 files, more files were not shown because too many files have changed in this diff. Show more