Compare commits

..
199 Commits
Author SHA1 Message Date
Ryan Houdek a9d00b3f8d Docs: Update for release FEX-2207 2022-07-07 10:06:29 -07:00
Ryan Houdek fb41ba172d Merge pull request #1835 from wannacu/main
AOTIR: Fix IRList delete
2022-07-07 10:03:33 -07:00
Ryan Houdek aec5b21d2a Merge pull request #1833 from neobrain/feature_generic_callbacks
Thunks: Implement generic callback support
2022-07-07 10:02:02 -07:00
Ryan Houdek 124097d563 Merge pull request #1839 from neobrain/fix_value_dom_validation
ValueDominanceValidation: Avoid stack exhaustion when aggregating predecessors
2022-07-07 08:51:54 -07:00
Ryan Houdek e137c2edff Merge pull request #1837 from neobrain/fix_unknown_vulkan_functions
Vulkan: Handle queries for unknown functions more gracefully
2022-07-07 08:45:48 -07:00
Tony Wasserka 2b35dd829c Thunks: Mark inline assembly in CallHostThunkFromRuntimePointer as volatile 2022-07-07 17:33:31 +02:00
Tony Wasserka 686895802a Thunks/GL: Drop unused manual function implementations
These functions are annotated with fexgen::callback_stub nowadays, so
they don't need custom host endpoints anymore.
2022-07-07 17:33:31 +02:00
Tony Wasserka 72d0228cd7 Thunks/gen: Remove now unneeded callback_structs and callback_typedefs files 2022-07-07 17:33:31 +02:00
Tony Wasserka 298e6cad0c Thunks/Xext: Enable automatic handling of callbacks 2022-07-07 17:33:31 +02:00
Tony Wasserka ad34c228e3 unittests/ThunkLibs: Extend FunctionPointerParameter test 2022-07-07 17:33:31 +02:00
Tony Wasserka 66f74a431f Thunks/gen: Clean up implementation of generic callbacks 2022-07-07 17:33:31 +02:00
Tony Wasserka 57bae90c69 Thunks: Fix generic callbacks on ARM hosts 2022-07-07 17:20:28 +02:00
Tony Wasserka 83c2e52ba5 Vulkan: Handle queries for unknown functions more gracefully 2022-07-07 17:00:15 +02:00
Tony Wasserka 4f8525da4d ValueDominanceValidation: Avoid stack exhaustion when aggregating predecessors
The recursive algorithm used here previously led to deeply nested function
calls, which eventually exhausted the available stack space. The simple
non-recursive algorithm used now avoids this problem at the expense of
small overhead.
2022-07-07 16:46:38 +02:00
Ryan Houdek 3b8491b558 Merge pull request #1836 from neobrain/fix_duplicate_gch_links
Thunks: Soften error condition to be non-fatal
2022-07-07 03:00:41 -07:00
Tony Wasserka f69c53d294 Thunks: Soften error condition to be non-fatal
Dota Underlords hit this when querying Vulkan two different symbols
that resolve to the same function. Ignoring the error in that specific
case is safe, since the linked guest functions have the same implementation.
2022-07-07 11:33:01 +02:00
Ryan Houdek 6b226dd6af Merge pull request #1834 from Sonicadvance1/fix_steam_vulkan_thunks
Thunks: Adds libvulkan steam pinned library thunking support
2022-07-06 23:57:04 -07:00
wannacu 227462e4d8 AOTIR: Fix IRList delete 2022-07-07 14:39:00 +08:00
Ryan Houdek 5f00ba5cae Thunks: Adds libvulkan steam pinned library thunking support
Since we have switched over to thunking the vulkan loader, behaviour has
changed here and we need to thunk this library.

While a bit unsafe to thunk arbitrary libraries, we know this one is
safe to thunk.

Fixes thunking Vulkan on steam games which have been broken since
switching over to vulkan loader thunking.

In the future this may become unnecessary but it is required for now.
2022-07-06 20:45:14 -07:00
Ryan Houdek 0b8a6d9599 Thunks: Add support for a @HOME@ prefix
Currently thunk prefixes are mutally exclusive. No thunks database path
can currently have more than one prefix. If this is necessary then we
can add it in the future.

Most minor of optimizations here, we only scan the string once to find a
prefix, instead of searching for it on all four prefix replacements.
2022-07-06 20:42:00 -07:00
Ryan Houdek 13bf04a81e FEXCore: Expose the Paths namespace publically
We will need the ability to get the home directory from the frontend.
Get it the same way as FEXCore.
2022-07-06 20:40:49 -07:00
Tony Wasserka 7eb8409f71 Thunks: Move definitions for LoadLib and IsLibLoaded together 2022-07-06 18:41:53 +02:00
Tony Wasserka e821072cb9 unittests/ThunkLibs: Fix FunctionPointerParameter test 2022-07-06 18:41:53 +02:00
Stefanos Kornilios Misis Poiitidis 9c01dd9d9f Thunks: PoC Callbacks using sha256 exports from host 2022-07-06 18:41:53 +02:00
Ryan Houdek 6a43db8c8f Merge pull request #1832 from Sonicadvance1/fix_thunk_crash
Thunks: Fix std::set crash
2022-07-06 02:00:04 -07:00
Ryan Houdek 3ffc301dd0 Thunks: Fix std::set crash
std::string_view was sticking around for longer than libraries being
loaded.
This was causing a crash.
Change this to a std::string directly until we support cleanly removing
this data on library unload.
2022-07-05 21:25:40 -07:00
Ryan Houdek 46fcbe2fc0 Merge pull request #1830 from Sonicadvance1/support_erofs
Support EroFS
2022-07-05 07:21:34 -07:00
Ryan Houdek b7806e47e9 Merge pull request #1831 from Sonicadvance1/minor_fexserver_fixes_pt2
FEXServer: Stop leaking FDs to subprocesses
2022-07-05 06:16:34 -07:00
Ryan Houdek a1ed54adc7 FEXRootFSFetcher: Add support for EroFS
Fixes a bug in `ExecAndWaitForResponse` where results > 1024 bytes would
overwrite data.

Switches from a custom format txt file to a json file.
JSON file now has a "Type" field to specify squashfs versus erofs.

JSON is now versioned so we don't need to move the file around, just
append to a new versioned segment.

Only shows erofs files if you have the bleeding edge `erofsfuse`
application.
This application was available starting with erofs-utils v1.5 which was
released on 2022-06-13, so it isn't available pretty much everywhere.
2022-07-05 04:40:13 -07:00
Ryan Houdek 250a4ea4e3 FEXServer: Stop leaking FDs to subprocesses
Attach the CLOEXEC flag on each of the FDs/Sockets we open.
2022-07-05 03:53:39 -07:00
Stefanos Kornilios Mitsis Poiitidis b3e090c8ff Merge pull request #1826 from Sonicadvance1/fix_gdb_install_library
FEXGDBReader: Fix install path
2022-07-05 08:47:42 +00:00
Stefanos Kornilios Mitsis Poiitidis 982518d3a4 Merge pull request #1829 from Sonicadvance1/fix_ioctl32_vblank
Ioctl32: Fix DRM_IOCTL_WAIT_VBLANK
2022-07-05 08:45:46 +00:00
Ryan Houdek 9ad1d5548d FEXServer: Add support for EroFS 2022-07-04 20:51:59 -07:00
Ryan Houdek 0de2558ef7 FEXConfig: Add support for erofs 2022-07-04 20:51:34 -07:00
Ryan Houdek 3e0e601616 FileFormatCheck: Add support for EroFS 2022-07-04 20:51:13 -07:00
Ryan Houdek 8b202b0ebf Ioctl32: Fix DRM_IOCTL_WAIT_VBLANK
Oops. Used the incorrect ioctl for this one. Copy and paste fail.
2022-07-04 18:14:07 -07:00
Ryan Houdek 6d2f98a379 Merge pull request #1822 from Sonicadvance1/check_binfmt_misc_conflict
CMake: Check for binfmt_misc conflicts before install
2022-07-02 01:38:37 -07:00
Ryan Houdek 84379b5fdf CMake: Check for binfmt_misc conflicts before install
Check for qemu and box binfmt_misc file conflicts before the
`binfmt_misc` install command.

This ensures if you're building from source that you won't inadvertently
install conflicting binfmt_misc files, breaking program execution.
2022-07-01 13:42:45 -07:00
Ryan Houdek d005fdcd03 Merge pull request #1823 from Sonicadvance1/classify_arm
unittests: Classify CPU based on CPU features
2022-06-30 23:34:36 -07:00
Ryan Houdek a97fb2f34f Merge pull request #1825 from Sonicadvance1/disable_posix_flake
unittests: Disable known flake in posix tests
2022-06-30 23:34:22 -07:00
Ryan Houdek d48981b6b0 FEXGDBReader: Fix install path 2022-06-30 22:34:25 -07:00
Ryan Houdek af32228e38 unittests: Disable known flake in posix tests
Interpreter is flakey here, likely due to some race, but is periodically
fails and makes CI red.
2022-06-30 14:41:14 -07:00
Ryan Houdek 8b35275ec1 unittests: Classify CPU based on CPU features
Instead of relying on runner features, classify based on CPU features.

This fixes an annoying issue where if running unit tests locally without
it set then you get an unexpected failure.

Fixes #1807
2022-06-30 13:55:38 -07:00
Ryan Houdek 302a6c96ff Merge pull request #1818 from FEX-Emu/skmp/fix-guest-h
ThunkLibs: Fix Guest.h
2022-06-30 10:02:48 -07:00
Stefanos Kornilios Misis Poiitidis f4a4b2c14b ThunkLibs: Fix Guest.h 2022-06-30 14:08:01 +03:00
Stefanos Kornilios Mitsis Poiitidis 88b94bef54 Merge pull request #1812 from FEX-Emu/skmp/add-thunks-islibloaded
Thunks: Add fex:is_lib_loaded
2022-06-30 04:45:37 +00:00
Stefanos Kornilios Mitsis Poiitidis 751b66d45d Merge pull request #1816 from Sonicadvance1/fix_vulkan_debug_report
Thunks/vulkan: Disable debug report callback
2022-06-30 04:43:43 +00:00
Stefanos Kornilios Mitsis Poiitidis ae6a57e667 Merge pull request #1815 from Sonicadvance1/fix_wine_preloader
Config: Fixes AppConfig for wine-preloader
2022-06-30 04:43:37 +00:00
Stefanos Kornilios Mitsis Poiitidis 9110546d34 Merge pull request #1814 from Sonicadvance1/pressure_vessel_fexserver_fix
FEXServerClient: When running under pressure-vessel don't use FEXServer rootfs
2022-06-30 04:43:26 +00:00
Stefanos Kornilios Mitsis Poiitidis 1f1d0706dd Merge pull request #1813 from Sonicadvance1/fexserver_changes
FEXServer: Minor changes
2022-06-30 04:43:16 +00:00
Ryan Houdek a82d41ee22 Thunks/vulkan: Disable debug report callback
This was checking for the wrong debug structure and it wasn't actually
unlinking it correctly from the linked list.
Since it was talking directly to a_0 instead of the current modifying
struct.

Instead of having a stubbed debug report struct, we can just remove the
structure from the linked list. Since we are already abusing a const
cast there anyway.

Fixes Vulkan thunks on Snapdragon.
2022-06-29 20:02:52 -07:00
Ryan Houdek e8e70828d1 Config: Fixes AppConfig for wine-preloader
When wine-preloader is executed it doesn't do an execve to passed in
wine program. It will instead map the executable directly in to memory
and start executing it.

This way we end up with a program executing like `wine-preloader
<absolute wine path> Game.exe`

This now handles the wine-preloader case so we can get the correct
application profile here.
2022-06-29 20:00:32 -07:00
Ryan Houdek b7a58fdb8a FEXServerClient: When running under pressure-vessel don't use FEXServer rootfs
pressure-vessel overrides our rootfs when it does a pivot_root.
Since we are still communicating to the FEXServer we were pulling the
configured rootfs.

Instead check if we are in pressure vessel and avoid doing that.

This fixes FEX running under pressure-vessel.

(There may be some implications to this down the road with code caching
but let's worry about that later)
2022-06-29 19:57:07 -07:00
Ryan Houdek dc9fde8fae FEXServer: Change server lock fifo to regular file
This doesn't need to be a FIFO.
Resolves an issue of trying to run a rootfs from a fex config folder mapped over nfs/sshfs.
2022-06-29 13:46:49 -07:00
Ryan Houdek c0cf4f6a61 FEXServer: Change socket pathname to include euid
Just to ensure the socket path is unique per user.

Noticed this while running independent FEXServers with multiple users.
2022-06-29 13:44:48 -07:00
Stefanos Kornilios Misis Poiitidis 21fc6bedcb Thunks: Add fex:is_lib_loaded 2022-06-29 19:20:08 +03:00
Ryan Houdek ad6fd5ab72 Merge pull request #1804 from neobrain/fix_cmake_thunks_portability
Allow building thunks on a wider range of platforms
2022-06-29 02:08:55 -07:00
Ryan Houdek 0f696c6092 Merge pull request #1811 from FEX-Emu/skmp/fix-vixl-assert
Dispatcher/Arm64: Fix vixl assert
2022-06-29 02:08:08 -07:00
Stefanos Kornilios Misis Poiitidis 9af3bc1864 Dispatcher/Arm64: Fix vixl assert 2022-06-29 10:41:24 +03:00
Ryan Houdek b020e593a5 Merge pull request #1802 from Sonicadvance1/fexserver_wait
FEXServer: Adds -w option for waiting on current FEXServer
2022-06-28 09:33:31 -07:00
Tony Wasserka 4771a340f5 Merge pull request #1803 from neobrain/fix_vulkan_debug_report
ThunkLibs/vulkan: Work around lack of generic callback support in VK_EXT_debug_report
2022-06-28 16:39:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 4449b60459 Merge pull request #1787 from FEX-Emu/skmp/gdb-jit-integration
gdb: jit integration
2022-06-27 11:57:40 +00:00
Stefanos Kornilios Misis Poiitidis 4139332ad9 GDBSymbols: Cleanups 2022-06-27 14:44:07 +03:00
Stefanos Kornilios Mitsis Poiitidis 30a28ff1ad Merge pull request #1801 from FEX-Emu/skmp/remove-thunk-warnings
ThunkLibs: silence warnings
2022-06-27 10:35:43 +00:00
Tony Wasserka 498d0fc145 CMake: Use toolchain files to set up x86 cross compilation 2022-06-25 14:00:36 +02:00
Tony Wasserka 4a547fe95f Thunks/CMake: Link against clang-cpp instead of clangTooling 2022-06-25 13:54:12 +02:00
Tony Wasserka ca906589d4 Thunks/CMake: Automatically discover the clang resource directory 2022-06-25 13:53:48 +02:00
Tony Wasserka 8bafae2262 ThunkLibs/vulkan: Work around lack of generic callback support in VK_EXT_debug_report 2022-06-25 12:44:58 +02:00
Tony Wasserka 4ef82c82b5 Revert "Thunks/vulkan: Disable support for debug extensions as they require callback support"
This reverts commit c04d2409da.
2022-06-25 11:48:39 +02:00
Ryan Houdek e03d253310 FEXServer: Adds -w option for waiting on current FEXServer
It can be useful to know in tooling when the current active FEXServer
has exited.

Two things can happen when this command is run.
No FEXServer is active, returns immediately.
A FEXServer is active, we query for a pidfd from the active server, then
we wait until it exits.

Both instances of this is valid to use.
2022-06-24 22:56:13 -07:00
Stefanos Kornilios Misis Poiitidis 48c45da2de ThunkLibs: silence warnings 2022-06-25 08:27:57 +03:00
Ryan Houdek 04a1ac967c Merge pull request #1760 from neobrain/feature_guest_callable_hostptrs
Thunks: Support returning host function pointers to the guest
2022-06-24 18:22:24 -07:00
Ryan Houdek aa17f64593 Merge pull request #1799 from FEX-Emu/skmp/fix-get_fdpath
FDUtils: Fix get_fdpath
2022-06-24 16:42:30 -07:00
Stefanos Kornilios Misis Poiitidis ae5cfcc249 Symbols: Add SymName function 2022-06-24 19:38:40 +03:00
Stefanos Kornilios Misis Poiitidis 8f578b57f2 GDBSymbols: Cleanups 2022-06-24 17:02:39 +03:00
Stefanos Kornilios Misis Poiitidis 3871646611 GDBSymbols: Add gdb reader for fex, GDBSymbols option to enable 2022-06-24 16:27:16 +03:00
Stefanos Kornilios Mitsis Poiitidis 48a574dfd3 FDUtils: Fix get_fdpath 2022-06-24 16:11:10 +03:00
Tony Wasserka ec49100c63 Thunks/CMake: Remove now unneeded helper functionality
This was needed for libvulkan_device. With libvulkan thunked directly now,
there is no further use of this code.
2022-06-24 11:56:36 +02:00
Tony Wasserka c04d2409da Thunks/vulkan: Disable support for debug extensions as they require callback support 2022-06-24 11:56:36 +02:00
Tony Wasserka b93b713179 Thunks/vulkan: Thunk libvulkan directly instead of libvulkan_device 2022-06-24 11:56:36 +02:00
Tony Wasserka fe2f54fc3d Thunks/GL: Use guest-callable host function pointers to implement glXGetProcAddress 2022-06-24 11:56:36 +02:00
Tony Wasserka 599f8f99ed Thunks/gen: Add support for thunking APIs that return host function pointers
Interface definitions must enable this functionality by enclosing functions
that may be called through function pointers in a namespace annotated with
`fexgen::indirect_guest_calls`. The guest thunk must further link any host
function pointers to a guest-side instance of CallHostThunkFromRuntimePointer.
2022-06-24 11:56:36 +02:00
Tony Wasserka b1b338ed5a Thunks: Simplify guest helper macros 2022-06-24 11:56:36 +02:00
Ryan Houdek e4d659a619 Merge pull request #1797 from FEX-Emu/skmp/remove-used-irs
IR: Remove GuestCallDirect, GuestCallIndirect
2022-06-23 10:45:01 -07:00
Stefanos Kornilios Misis Poiitidis 9ce94266e1 IR: Remove GuestCallDirect, GuestCallIndirect 2022-06-23 19:46:25 +03:00
Ryan Houdek 5a19425b28 Merge pull request #1792 from Sonicadvance1/fexserver
FEXServer: Adds new FEXServer service
2022-06-23 09:25:29 -07:00
Ryan Houdek 1494aac861 Merge pull request #1796 from Sonicadvance1/fix_clone3_stack_again
Linux: Fixes for clone3 stack size
2022-06-23 09:21:39 -07:00
Ryan Houdek 8a21ecabee Merge pull request #1795 from neobrain/fix_radata_asan
Fix inconsistent allocation schemes used for RegisterAllocationData
2022-06-23 09:03:05 -07:00
Ryan Houdek 005b2bc3db Linux: Fixes for clone3 stack size
We weren't adjusting the guest stack size when using clone3.
If the clone comes from CLONE3 then we need to offset the RSP by the
provided stack size.

This also translates to fork/vfork through clone, so make sure to adjust
stack in that case as well.

Fixes Ender Lilies crashing with Ubuntu 22.04 rootfs
2022-06-23 08:41:15 -07:00
Tony Wasserka f27830bf41 Fix inconsistent allocation schemes used for RegisterAllocationData 2022-06-23 17:16:44 +02:00
Ryan Houdek 5fe6afc270 FEXServer: Adds new FEXServer service
This is a relatively invasive change since multiple things needed to
happen at once.

* Socket based logging is removed
  * Logging has been replaced to only support stdout, stderr, and server
  * Server is now default and replaces what FEXLogServer did
  * Server logging now uses a pipe instead of a socket
  * Can be faster than stderr and stdout since the application doesn't
    need to wait on terminal output

* FEXMountDaemon has been removed
  * Functionality has been merged in to FEXServer

* FEXServer is always executed on FEX initialization time
  * Similar in behaviour to Wine's wineserver
  * Can explicitly start this before using FEX for logging purposes
  * Stays around until all instances of FEX exit
  * Will stick around for a short amount of time in case of spurious
    execution

* FEXServer will soon be extended to do more than logging and squashfs
  mounting

* FEX rootfs scripts will need to be updated to support this path
  * Just means rbinding the /tmp folder and forcing a FEXServer instance
    to be alive
* Pressure-vessel works fine in this case since FEXServer will already
  be running
  * It already rbinds the host /tmp folder which is why this works
2022-06-23 07:51:56 -07:00
Stefanos Kornilios Mitsis Poiitidis 4f9bc70562 Merge pull request #1794 from FEX-Emu/skmp/fix-unidispatch-multithread
Dispatchers: Use thread local emitters for backend callbacks
2022-06-23 10:45:12 +00:00
Stefanos Kornilios Misis Poiitidis 835155021d Dispatchers: Use thread local emitters for backend callbacks 2022-06-23 13:33:56 +03:00
Stefanos Kornilios Misis Poiitidis e2c5c2292d Dispatchers: Use thread local emitters for backend callbacks 2022-06-23 13:20:03 +03:00
Stefanos Kornilios Mitsis Poiitidis 072690a191 Merge pull request #1782 from FEX-Emu/skmp/ipr-unified-dispatch
Backends: Unified dispatch, interface rework, cleanups
2022-06-23 07:03:14 +00:00
Ryan Houdek a2c9d5a398 Merge pull request #1790 from neobrain/fix_missing_packfn_define
unittests/ThunkLibs: Fix test failures due to missing FEX_PACKFN_LINKAGE define
2022-06-22 19:08:55 -07:00
Ryan Houdek 5944342604 Externals: update cpp-optparse 2022-06-22 17:39:59 -07:00
Stefanos Kornilios Misis Poiitidis 4dbcd34548 Backends: Rework some of the interfaces, unified backend dispatch 2022-06-22 22:24:42 +03:00
Tony Wasserka f02c73a2c2 unittests/ThunkLibs: Fix test failures due to missing FEX_PACKFN_LINKAGE define 2022-06-22 17:03:01 +02:00
Ryan Houdek 158ba1ae3b Merge pull request #1788 from Sonicadvance1/safe_v3d_csd
Ioctl: Safely access v3d csd ioctl structure
2022-06-21 15:33:50 -07:00
Ryan Houdek a85a77e088 Ioctl: Safely access v3d csd ioctl structure
DRM ioctls are taking advantage of the fact that any ioctl structure
that is passed in to the kernel smaller than the expected value will
zero out the remaining members of the structure.

Some more ioctls coming down the pipe will also abuse this, so we might
as well as get started with the v3d ioctl that requires it.

Due to how this works, the kernel knows how big an ioctl structure
should be and if the userspace passes in a smaller struct, it will
memset the remaining size to zero.
2022-06-19 09:53:50 -07:00
Ryan Houdek 542ab04671 Merge pull request #1786 from Sonicadvance1/optimize_fd_path
Linux: Make `get_fdpath` more optimal
2022-06-18 01:57:55 -07:00
Ryan Houdek 28ee2ca5a2 Linux: Make get_fdpath more optimal
std::filesystem::canonical is very heavyweight and walks the full path
to ensure that each folder in the path is not a symlink.

eg:
```
readlink("/proc", 0x7ffd5646e210, 1023)                                         = -1 EINVAL (Invalid argument)
readlink("/proc/self", "880556", 1023)                                          = 6
readlink("/proc/880556", 0x7ffd5646e210, 1023)                                  = -1 EINVAL (Invalid argument)
readlink("/proc/880556/fd", 0x7ffd5646e210, 1023)                               = -1 EINVAL (Invalid argument)
readlink("/proc/880556/fd/5", "/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-6"..., 1023) = 86
readlink("/home", 0x7ffd5646e210, 1023)                                         = -1 EINVAL (Invalid argument)
readlink("/home/ryanh", 0x7ffd5646e210, 1023)                                   = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu", 0x7ffd5646e210, 1023)                          = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS", 0x7ffd5646e210, 1023)                   = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04", 0x7ffd5646e210, 1023)      = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr", 0x7ffd5646e210, 1023)  = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
readlink("/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2", 0x7ffd5646e210, 1023) = -1 EINVAL (Invalid argument)
```

This is what was occuring for every single mmap that occurs. /really/
adding to the time for the syscall to take.
This also happens on a couple of other syscalls which are using this new
path now.

The primary reason why this works is that we know that every entry in
`/proc/self/fd/` is a symlink. So instead of asking for canonical, we
can just read the symlink and this will redirect us to the canonical
path.
So this previous example goes from 14 syscalls down to 1.

eg:
```
readlinkat(AT_FDCWD, "/proc/self/fd/5", "/home/ryanh/.fex-emu/RootFS/Ubuntu_22_04/usr/lib/x86_64-linux-gnu/ld-linux-x86-6"..., 4096) = 86
```

While this is only a minor improvement in the "typical" operating environment,
this significantly improves performance of FEX under proot or if the
rootfs lives on a network share.
2022-06-18 01:45:37 -07:00
Stefanos Kornilios Mitsis Poiitidis 3913dd6c8b Merge pull request #1785 from Sonicadvance1/wine_app_config
Common: Support application profiles for games launched through wine
2022-06-18 07:42:16 +00:00
Ryan Houdek 1ecf147e3e Common: Support application profiles for games launched through wine
Wine will set the application name later in the boot process but we
can't defer application loading that late.

Once an application is loaded with wine or wine64, then check the next
argument for the application name instead.

This will allow us to have wine application application profiles.

eg: FEXInterpreter `which wine` $HOME/.wine/drive_c/GOG\ Games/Oblivion/Oblivion.exe
This will give us the application name of `Oblivion.exe`

Same with: FEXInterpreter `which wine` C:\\GOG\ Games\\Oblivion\\Oblivion.exe
2022-06-17 23:24:42 -07:00
Ryan Houdek d8fa53a445 Merge pull request #1783 from FEX-Emu/skmp/fix-thunksdb-prefixing
ThunksDB: Fix String.find error check
2022-06-17 15:32:08 -07:00
Ryan Houdek 1c1ad876ca Merge pull request #1784 from lioncash/regsize
CoreState: Add register size constants
2022-06-17 15:31:59 -07:00
lioncash 58ae49372c CoreState: Add register size constants
Gets rid of some magic numbers and reduces the number of things that
need to manually change (e.g. when supporting AVX and needing to
increase the xmm size).
2022-06-17 10:55:07 -04:00
Stefanos Kornilios Misis Poiitidis 917e69b021 ThunksDB: Fix String.find error check 2022-06-17 15:04:17 +03:00
Mai 9242e59841 Merge pull request #1781 from Sonicadvance1/fix_free
AOTIR: Fix RAData free
2022-06-16 20:11:55 -04:00
Ryan Houdek 05d15fa052 AOTIR: Fix RAData free
This thing is allocated with malloc so it needs to be free'd with free.
This was poisoning asan runs.
2022-06-16 17:01:06 -07:00
Stefanos Kornilios Mitsis Poiitidis 0a62a4c571 Merge pull request #1775 from FEX-Emu/skmp/dispatcher-per-context
Make Dispatcher per Context from per Thread, Simplify TestHarnessRunner
2022-06-16 21:08:32 +00:00
Stefanos Kornilios Misis Poiitidis 525266e482 Core: Move common signal handling setup to context 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 503881f5df Jit/Arm64: Fix ubuntu 20.04 include 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 04fd4f2bb4 Arm64Dispatcher: Add missing ExitFunctionLinker Pointer 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 8be1e5e260 TestHarnessRunner: Fix formatting 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis e84e084d56 FEXLogServer: Work around linking issues 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 12ad05e089 HostRunner: Fix typo for arm64 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 472ce7c1e4 Context: Add DipsatcherConfig to class, minor header massaging 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 19fbc10c16 TestHarnessRunner: Don't setup a FEX Context when running host tests 2022-06-16 19:41:42 +00:00
Stefanos Kornilios Misis Poiitidis 39f3406fba Backends+Context: Make Dispatcher per Context from per Thread 2022-06-16 19:41:42 +00:00
Ryan Houdek 19b0a9cd7d Merge pull request #1779 from lioncash/irname
Arm64/JIT: Use IR names in opcode implementations
2022-06-16 12:37:18 -07:00
Stefanos Kornilios Mitsis Poiitidis eac579f714 Merge pull request #1778 from FEX-Emu/skmp/fix-thread-creation-race
Context: Fix CreateThread partial initialization issue
2022-06-16 18:17:34 +00:00
Stefanos Kornilios Mitsis Poiitidis b0a31f7105 Merge pull request #1770 from FEX-Emu/skmp/custom-ir-handlers-2
Context: Decouple from CodeLoader, introduce generic CustomIREntrypoints
2022-06-16 18:04:18 +00:00
Stefanos Kornilios Misis Poiitidis 29163cfd07 Context: Fix CreateThread partial initialization issue 2022-06-16 17:51:49 +00:00
Stefanos Kornilios Mitsis Poiitidis f05b24636f OpDisp: Re add ShouldDump 2022-06-16 17:38:44 +00:00
lioncash 39466a3a8f Arm64/VectorOps: Move arguments over to IR names 2022-06-16 12:03:33 -04:00
lioncash bccf18f46c Arm64/MoveOps: Move arguments over to IR names 2022-06-16 12:03:26 -04:00
lioncash 1109126e35 Arm64/MiscOps: Move arguments over to IR names 2022-06-16 12:03:18 -04:00
lioncash 98dd40c3a8 Arm64/MemoryOps: Move arguments over to IR names 2022-06-16 12:03:12 -04:00
lioncash b1cfb104ff Arm64/FlagOps: Move arguments over to IR names 2022-06-16 12:03:02 -04:00
lioncash c339cfe184 Arm64/EncryptionOps: Move arguments over to IR names 2022-06-16 12:02:53 -04:00
lioncash ca8e6a83e1 Arm64/ConversionOps: Move arguments over to IR names 2022-06-16 12:02:45 -04:00
lioncash 760217ab29 Arm64/BranchOps: Move arguments over to IR names 2022-06-16 12:02:38 -04:00
lioncash ee9361bf41 Arm64/ALUOps: Move arguments over to IR names
Makes it a little easier to read.
2022-06-16 12:02:28 -04:00
Stefanos Kornilios Mitsis Poiitidis e62bc24b3f Merge pull request #1777 from FEX-Emu/skmp/create-directories-during-configuration
CMAKE: Create directories during configuration, fixes endless generation of unittests
2022-06-15 01:58:57 +03:00
Stefanos Kornilios Misis Poiitidis 18074307f6 Core: Add Creator, Data to CustomIRHandlers, return them + lock on Add 2022-06-15 01:48:32 +03:00
Ryan Houdek 63b70ff3d4 Merge pull request #1776 from lioncash/vixl-pcl
Arm64/EncryptionOps: Fix register specifiers in PCLMUL movs
2022-06-14 15:46:40 -07:00
Stefanos Kornilios Misis Poiitidis dacdfd5c02 CMAKE: Create directories during configuration, fixes endless generation of unittests 2022-06-15 01:10:33 +03:00
lioncash d5c039cdc7 Arm64/EncryptionOps: Fix register specifiers in PCLMUL movs
Prevents an assertion from firing in vixl, since the destination needs
to be a scalar
2022-06-14 10:15:27 -04:00
Ryan Houdek e2e6f2a92b Merge pull request #1767 from Sonicadvance1/pressure_vessel_thunks
Thunks: Support pressure-vessel prefixes
2022-06-13 10:26:00 -07:00
Ryan Houdek d7d8244593 Thunks: Support pressure-vessel prefixes
Instead of duplicating prefixes in the ThunksDB json file even more,
just do a prefix replacement when the thunks database is being parsed.

Changes the JSON overlay arrays over to @PREFIX@ and reduce the
duplication.

Adds support for the pressure-vessel prefix `/usr/lib/pressure-vessel/overrides/lib`

This gets thunks ready for running under pressure-vessel, once vulkan
thunking switches from device libraries to the vulkan loader it should
just work.
2022-06-13 09:48:17 -07:00
Ryan Houdek b05e5ce14e Merge pull request #1773 from lioncash/header
JITs: Qualify external includes consistently
2022-06-13 09:46:54 -07:00
lioncash 08730c817a X86Dispatcher: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:55:44 -04:00
lioncash 34e086eec2 Arm64Dispatcher: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:54:49 -04:00
lioncash 83c8a9b675 ArchHelpers/Arm64: Qualify external includes
Qualifies external includes with <> instead of quotes.
2022-06-13 10:52:39 -04:00
lioncash a3db629fd6 Arm64Emitter: Qualify external includes
Qualifies external includes with <> instead of quotes
2022-06-13 10:50:03 -04:00
lioncash 64b3cef126 ARM64/JITClass: include dependencies
Includes dependencies directly used by the class. Also qualifies the
vixl includes with <> instead of quotes.
2022-06-13 10:43:13 -04:00
Stefanos Kornilios Misis Poiitidis 7100a2eaee Context: Decouple from CodeLoader, introduce generic CustomIRHandlers 2022-06-13 11:07:16 +03:00
Stefanos Kornilios Mitsis Poiitidis ffcde1823b Merge pull request #1769 from FEX-Emu/skmp/interlocked-invalidate-2
Invalidations: Move invalidation locks to Context
2022-06-13 11:06:34 +03:00
Stefanos Kornilios Mitsis Poiitidis 4c73c715ad Merge pull request #1771 from Sonicadvance1/fix_32bit_allocator
Linux: Fixes 32-bit allocator range scanning
2022-06-13 11:06:14 +03:00
Ryan Houdek 26c3ebd6a7 Linux: Fixes 32-bit allocator range scanning
32-bit allocations were scanning for available pages by shifting by the
length of allocation. This is incorrect since large allocations would
then only scan through the region in very large chunks. Especially if
MAP_32BIT was present. Instead scan by page size to ensure better fit.

This fixes X-Plane 11.

Also ensure we don't try to munmap a range unless the result is
MAP_FAILED, noticed we were trying to munmap ~0ULL
2022-06-12 17:53:12 -07:00
Ryan Houdek 790bd9747f Merge pull request #1756 from FEX-Emu/skmp/ipr-own-irlists
IPR: Store copy of IRLists, Dispatcher cleanups
2022-06-11 12:15:18 -07:00
Stefanos Kornilios Misis Poiitidis 83aa8731d1 Invalidations: Move invalidation locks to Context, extend InvalidateGuestCodeRange with optional callback 2022-06-11 17:15:41 +03:00
Stefanos Kornilios Mitsis Poiitidis bbd9eb5b9a Merge pull request #1766 from FEX-Emu/skmp/fix-smc-mt-2
SMC: Track code pages before frontend decode
2022-06-11 16:05:02 +03:00
Stefanos Kornilios Misis Poiitidis cc90fa5773 LookupCache: Review feedback & cleanups 2022-06-11 13:17:49 +03:00
Stefanos Kornilios Mitsis PoiitidisandTony Wasserka 30ebc6c938 Update External/FEXCore/Source/Interface/Core/LookupCache.h
Co-authored-by: Tony Wasserka <4840017+neobrain@users.noreply.github.com>
2022-06-11 12:52:37 +03:00
Ryan Houdek c0a8984799 Merge pull request #1764 from Sonicadvance1/fix_xxhash
FEXRootFSFetcher: Update and fix xxhash file hashing
2022-06-10 04:57:09 -07:00
Stefanos Kornilios Misis Poiitidis a6e34d301e Backends: Move and use INITIAL_CODE_SIZE, MAX_CODE_SIZE consts 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 82c88168cc JITPointers: Should be a struct now 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis d169c3ed4a Context: Remove unused param from ClearCodeCache 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 563d11a702 IPR: Actually works now 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 31e2ba4696 IPR: Fix the build 2022-06-10 10:39:43 +03:00
Stefanos Kornilios Misis Poiitidis 2c3baaad1c Context: Split LocalIR to PrecompiledIR and DebugStore 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 01bd43aa2b IPR: Store IR in the code bufffer, use executer function, cleanups 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis f4d095ce0f Jit/dispatch: ExecuteBlocksWithCall -> ExecuteBlocksWithCall 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 8c5bd9a058 IPR: Move dispatcher creation to InterpreterCore ctor 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 335ce1d056 JIT: Merge common CPUFrame::Pointers 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 0afb3caaed Refactors: Move CodeBuffers to CPUBackend, SignalHandlerRefCounter to CpuStateFrame, Arm64 Relocation to ARM64Jit 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 8521aaf65c Core, Frontend, LookupCache: Track code pages before frontend decode 2022-06-10 10:13:47 +03:00
Stefanos Kornilios Misis Poiitidis 0a451cfe0d FEXLT: Run 100 iterations of smc-mt-2 to catch races 2022-06-10 10:01:37 +03:00
Stefanos Kornilios Mitsis Poiitidis cb0935c96b Merge pull request #1737 from FEX-Emu/skmp/fex-linux-tests
unittests: Add FEXLinuxTests with a few tests
2022-06-10 08:51:15 +03:00
Stefanos Kornilios Misis Poiitidis 40ec910108 unittests: Add FEXLinuxTests, with a few signal and smc tests 2022-06-10 08:34:40 +03:00
Stefanos Kornilios Mitsis Poiitidis 2d3c6efae2 Merge pull request #1761 from FEX-Emu/skmp/ir-meta-fns
IR: add IsFragmentExit, IsBlockExit
2022-06-10 08:15:03 +03:00
Stefanos Kornilios Mitsis Poiitidis c027acecf8 Merge pull request #1763 from Sonicadvance1/ci_auto_rootfs
CI: Auto rootfs fetching
2022-06-10 07:34:56 +03:00
Ryan Houdek bc2840e4a7 FEXRootFSFetcher: Update and fix xxhash file hashing
The final tail of the file reading was incorrect, so our hashing was
"correct" but it was using stale data from the previous block size read.

Noticed this while wiring up the CI rootfs fetching since the hashing is
a lot simpler there.

Now instead of reading a tail, just attempt to read the full block size
and use the resulting data size instead. Confirmed it matches expected
results now.

In the process we are going to need to update hyperlinks and hashes
anyway, change the hash to XXH3 so it is faster to run.
2022-06-09 20:20:43 -07:00
Ryan Houdek 869b472d91 xxhash: Update xxhash to 0.8.1 2022-06-09 20:09:49 -07:00
Ryan Houdek 81dfc21700 unittests/gvisor-test: Disable semaphore test for now
New rootfs images cause different data to be returned in permissions
than this test was expecting. Disable this while we are upgrading CI.
2022-06-09 20:08:24 -07:00
Ryan Houdek a3d8fe2362 github: Allow CI to fetch its own rootfs
Just need to set some environment variables and execute our new script
2022-06-09 20:08:24 -07:00
Ryan Houdek d2a57f6231 Scripts: Add CI rootfs fetch script
This will allow our CI system to automatically pull their rootfs.
2022-06-09 20:08:24 -07:00
Ryan Houdek 9e9ceb3894 Merge pull request #1762 from lioncash/pcl
OpcodeDispatcher: Handle CLMUL opcode extension
2022-06-09 12:55:54 -07:00
Stefanos Kornilios Misis Poiitidis bd70088877 IR: add IsFragmentExit, IsBlockExit 2022-06-09 20:44:24 +03:00
lioncash 63c9d99b33 CPUID: Enable PCLMULQDQ CPUID bit
Now that we handle all PCLMULQDQ facilities, we can enable this bit.
2022-06-09 12:42:34 -04:00
lioncash a2469f48e7 OpcodeDispatcher: Handle VPCLMULQDQ 2022-06-09 12:42:34 -04:00
lioncash 62f0f421a9 OpcodeDispatcher: Handle PCLMULQDQ 2022-06-09 11:22:32 -04:00
lioncash 0b24758f41 IR: Add PCLMUL IR opcode 2022-06-09 11:22:29 -04:00
lioncash 147897e13d External: Update vixl submodule
Allows use of the 128-bit variant of PMULL/PMULL2
2022-06-09 11:06:47 -04:00
Ryan Houdek ade3a5275d Merge pull request #1758 from lioncash/warn
Tests/IRLoader: Silence missing override warning
2022-06-08 12:35:29 -07:00
Ryan Houdek 51c5f945b4 Merge pull request #1759 from lioncash/ir
IR.json: Correct 'Dest' key to 'Desc'
2022-06-08 12:35:04 -07:00
lioncash a4e2e36243 IR.json: Correct 'Dest' key to 'Desc'
In a few places the description key was accidentally written as Dest.
2022-06-08 10:58:49 -04:00
lioncash 2c5dc13201 Tests/IRLoader: Silence missing override warning 2022-06-08 10:41:28 -04:00
Ryan Houdek 43234939ca Merge pull request #1757 from Sonicadvance1/fix_open_wrapping
Linux: Fixes `open` syscall emulated path handling
2022-06-08 03:52:16 -07:00
Ryan Houdek 87b3d50899 Linux: Fixes open syscall emulated path handling
open is rarely used compared to openat and openat2, so this has been
just missed but nothing really got upset about it.

Fixes Proton Experimental while running under pressure-vessel.
2022-06-07 17:27:25 -07:00
Ryan Houdek da8dbf1777 Merge pull request #1755 from Sonicadvance1/hypervisor_bit
CPUID: Enable the hypervisor bit
2022-06-06 11:52:21 -07:00
Ryan Houdek 34e1fcccf8 CPUID: Enable the hypervisor bit
This was originally set to zero out of concern for any application that
is doing anti-emulation, anti-cheat, anti-VM checks.

This concern is likely unwarranted, and if any application/game starts
hitting this as a problem then we can throw an application profile at it
instead.

This makes pressure-vessel emulation checking more optimal by it not
having to do a uname dance.
2022-06-06 11:36:42 -07:00
Stefanos Kornilios Mitsis Poiitidis c99d1e48bd Merge pull request #1752 from FEX-Emu/skmp/tso-auto-migration
TSO: Add auto migration optimisation for applications that don't need TSO
2022-06-06 19:42:39 +03:00
Stefanos Kornilios Misis Poiitidis 6a428043f6 TSO: Add auto migration optimisation for applications that don't need TSO 2022-06-06 16:51:32 +03:00
Stefanos Kornilios Mitsis Poiitidis aafe7ff10f Merge pull request #1751 from Sonicadvance1/allow_override
Scripts: Allow user override on tagged version
2022-06-05 12:32:03 +03:00
Ryan Houdek 952e157770 Scripts: Allow user override on tagged version
In the case of a missing month or a minor version, need to allow user
defined overrides.
2022-06-04 18:06:28 -07:00
235 changed files with 8788 additions and 5959 deletions

No files matched your search

+26 -1
View File
@@ -29,6 +29,20 @@ jobs:
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
@@ -51,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -153,6 +167,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+14 -4
View File
@@ -1,7 +1,11 @@
cmake_minimum_required(VERSION 3.14)
project(FEX)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
@@ -12,6 +16,7 @@ option(ENABLE_MOLD "Enable linking with mold" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
@@ -24,8 +29,7 @@ option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
set (X86_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86.cmake" CACHE FILEPATH "Toolchain file for the x86 (cross-)compiler")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
# These options are meant for package management
@@ -43,6 +47,12 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_GDB_SYMBOLS)
message(STATUS "GDBSymbols support enabled")
add_definitions(-DGDB_SYMBOLS_ENABLED=1)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
@@ -75,6 +85,7 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
set (X86_TOOLCHAIN_FILE "")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
@@ -513,8 +524,7 @@ if (BUILD_THUNKS)
BINARY_DIR "Guest"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DX86_C_COMPILER:STRING=${X86_C_COMPILER}"
"-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
+60 -239
View File
@@ -6,18 +6,10 @@
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGL.so",
"/usr/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/lib/x86_64-linux-gnu/libGL.so",
"/lib/x86_64-linux-gnu/libGL.so.1",
"/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/lib/x86_64-linux-gnu/libGL.so.1.7.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.2.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.7.0"
]
},
"GLESv2": {
@@ -26,322 +18,151 @@
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/lib/x86_64-linux-gnu/libGLESv2.so",
"/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libX11.so",
"/usr/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libX11.so",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/lib/x86_64-linux-gnu/libX11.so",
"/lib/x86_64-linux-gnu/libX11.so.6",
"/lib/x86_64-linux-gnu/libX11.so.6.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6.4.0"
]
},
"Vulkan-radeon": {
"Library": "libvulkan_radeon-guest.so",
"Vulkan": {
"Library": "libvulkan-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/lib/x86_64-linux-gnu/libvulkan_radeon.so"
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so.1",
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
],
"Comment": [
"Vulkan library relies on xcb, otherwise it crashes with jemalloc"
]
},
"Vulkan-lavapipe": {
"Library": "libvulkan_lvp-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/lib/x86_64-linux-gnu/libvulkan_lvp.so"
]
},
"Vulkan-freedreno": {
"Library": "libvulkan_freedreno-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/lib/x86_64-linux-gnu/libvulkan_freedreno.so"
]
},
"Vulkan-intel": {
"Library": "libvulkan_intel-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/lib/x86_64-linux-gnu/libvulkan_intel.so"
]
},
"Vulkan-panfrost": {
"Library": "libvulkan_panfrost-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/lib/x86_64-linux-gnu/libvulkan_panfrost.so"
]
},
"Vulkan-nvidia": {
"Library": "libvulkan_nvidia-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/usr/local/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/lib/x86_64-linux-gnu/libGLX_nvidia.so.0"
],
"Comment": [
"Not currently wired up"
]
},
"Vulkan-virtio": {
"Library": "libvulkan_virtio-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/lib/x86_64-linux-gnu/libvulkan_virtio.so"
]
},
"xcb": {
"Library": "libxcb-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb.so",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/lib/x86_64-linux-gnu/libxcb.so",
"/lib/x86_64-linux-gnu/libxcb.so.1",
"/lib/x86_64-linux-gnu/libxcb.so.1.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb_dri2-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb_dri3-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb_xfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb_shm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb_sync-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/lib/x86_64-linux-gnu/libxcb-sync.so",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb_randr-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb_present-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-present.so",
"/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb_glx-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libshmfence-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/lib/x86_64-linux-gnu/libxshmfence.so",
"/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libdrm.so",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/lib/x86_64-linux-gnu/libdrm.so",
"/lib/x86_64-linux-gnu/libdrm.so.2",
"/lib/x86_64-linux-gnu/libdrm.so.2.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2.4.0"
]
},
"asound": {
"Library": "libasound-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libasound.so",
"/usr/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libasound.so",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/lib/x86_64-linux-gnu/libasound.so",
"/lib/x86_64-linux-gnu/libasound.so.2",
"/lib/x86_64-linux-gnu/libasound.so.2.0.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2.0.0"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXrender.so",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/lib/x86_64-linux-gnu/libXrender.so",
"/lib/x86_64-linux-gnu/libXrender.so.1",
"/lib/x86_64-linux-gnu/libXrender.so.1.3.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXext.so",
"/usr/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libXext.so",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/lib/x86_64-linux-gnu/libXext.so",
"/lib/x86_64-linux-gnu/libXext.so.6",
"/lib/x86_64-linux-gnu/libXext.so.6.4.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/lib/x86_64-linux-gnu/libXfixes.so",
"/lib/x86_64-linux-gnu/libXfixes.so.3",
"/lib/x86_64-linux-gnu/libXfixes.so.3.1.0"
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3.1.0"
]
},
"":{}
+7 -8
View File
@@ -81,6 +81,7 @@ set (SRCS
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
Interface/Core/Core.cpp
Interface/Core/CPUBackend.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/GdbServer.cpp
@@ -117,6 +118,7 @@ set (SRCS
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/HLE/Thunks/Thunks.cpp
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
@@ -206,7 +208,9 @@ if (ENABLE_JIT_ARM64)
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
Interface/Core/JIT/Arm64/VectorOps.cpp
Interface/Core/JIT/Arm64/Arm64Relocations.cpp
)
endif()
set (LIBS vixl dl xxhash tiny-json)
@@ -224,13 +228,11 @@ set(OUTPUT_IR_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/IR")
set(OUTPUT_NAME "${OUTPUT_IR_FOLDER}/IRDefines.inc")
set(INPUT_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/IR/IR.json")
add_custom_target(CREATE_IR_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_IR_FOLDER}")
file(MAKE_DIRECTORY "${OUTPUT_IR_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_NAME}"
DEPENDS "${INPUT_NAME}"
DEPENDS CREATE_IR_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py" "${INPUT_NAME}" "${OUTPUT_NAME}"
)
@@ -244,7 +246,6 @@ set(OUTPUT_IR_DOC "${CMAKE_BINARY_DIR}/IR.md")
add_custom_command(
OUTPUT "${OUTPUT_IR_DOC}"
DEPENDS "${INPUT_NAME}"
DEPENDS CREATE_IR_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py" "${INPUT_NAME}" "${OUTPUT_IR_DOC}"
)
@@ -265,15 +266,13 @@ set(INPUT_CONFIG_NAME "${CMAKE_BINARY_DIR}/generated/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
set(OUTPUT_MAN_NAME_COMPRESS "${CMAKE_BINARY_DIR}/generated/FEX.1.gz")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
file(MAKE_DIRECTORY "${OUTPUT_CONFIG_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_CONFIG_NAME}"
OUTPUT "${OUTPUT_CONFIG_OPTION_NAME}"
OUTPUT "${OUTPUT_MAN_NAME}"
DEPENDS "${INPUT_CONFIG_NAME}"
DEPENDS CREATE_CONFIG_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py" "${INPUT_CONFIG_NAME}" "${OUTPUT_CONFIG_NAME}" "${OUTPUT_MAN_NAME}"
"${OUTPUT_CONFIG_OPTION_NAME}"
+16
View File
@@ -371,6 +371,22 @@ namespace JSON {
return {};
}
std::string FindContainer() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
}
}
return {};
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
+20 -11
View File
@@ -208,6 +208,16 @@
"Useful for determining hot blocks of code",
"Has some file writing overhead per JIT block"
]
},
"GDBSymbols": {
"Type": "bool",
"Default": "false",
"Desc": [
"Integrates with GDB using the JIT interface.",
"Needs the fex jit loader in GDB, which can be loaded via `jit-reader-load libFEXGDBReader.so.`",
"Also needs x86_64-linux-gnu-objdump in PATH.",
"Can be very slow."
]
}
},
"Logging": {
@@ -219,22 +229,13 @@
"Disables logging"
]
},
"OutputSocket": {
"Type": "str",
"Default": "",
"Desc": [
"Socket to connect to",
"eg: localhost:8087",
"If set will override the OutputLog location"
]
},
"OutputLog": {
"Type": "str",
"Default": "stderr",
"Default": "server",
"ShortArg": "o",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, <Filename>]"
"[stdout, stderr, server, <Filename>]"
]
}
},
@@ -260,6 +261,14 @@
"Highly likely to break any multithreaded application if disabled."
]
},
"TSOAutoMigration": {
"Type": "bool",
"Default": "true",
"Desc": [
"Automatically enables TSO when shared memory is used.",
"Should work without issues in most cases."
]
},
"X87ReducedPrecision": {
"Type": "bool",
"Default": "false",
+7 -2
View File
@@ -43,8 +43,8 @@ namespace FEXCore::Context {
delete CTX;
}
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader) {
return CTX->InitCore(Loader);
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, uint64_t InitialRIP, uint64_t StackPointer) {
return CTX->InitCore(InitialRIP, StackPointer);
}
void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler) {
@@ -156,6 +156,7 @@ namespace FEXCore::Context {
void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler) {
CTX->SyscallHandler = Handler;
CTX->SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
@@ -193,6 +194,10 @@ namespace FEXCore::Context {
return CTX->UnloadAOTIRCacheEntry(Entry);
}
CustomIRResult AddCustomIREntrypoint(FEXCore::Context::Context *CTX, uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
return CTX->AddCustomIREntrypoint(Entrypoint, Handler, Creator, Data);
}
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->CompileRIP(CTX->ParentThread, RIP);
+36 -17
View File
@@ -5,6 +5,7 @@
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -42,10 +43,14 @@ namespace CodeSerialize {
namespace CPU {
class Arm64JITCore;
class X86JITCore;
class InterpreterCore;
class Dispatcher;
}
namespace HLE {
struct SyscallArguments;
class SyscallHandler;
class SourcecodeResolver;
struct SourcecodeMap;
}
}
@@ -72,6 +77,7 @@ namespace FEXCore::Context {
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::CPU::InterpreterCore;
friend class FEXCore::IR::Validation::IRValidation;
struct {
@@ -86,6 +92,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(TSOAutoMigration, TSOAUTOMIGRATION);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(ABINoPF, ABINOPF);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
@@ -102,18 +109,15 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
} Config;
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
IntCallbackReturn InterpreterCallbackReturn;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
uint64_t ThreadID{};
FEXCore::Core::InternalThreadState* ParentThread;
std::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
@@ -125,9 +129,13 @@ namespace FEXCore::Context {
Event PauseWait;
bool Running{};
std::shared_mutex CodeInvalidationMutex;
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
std::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
@@ -142,7 +150,7 @@ namespace FEXCore::Context {
Context();
~Context();
FEXCore::Core::InternalThreadState* InitCore(FEXCore::CodeLoader *Loader);
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
@@ -171,6 +179,11 @@ namespace FEXCore::Context {
RemoveThreadCodeEntry(Frame->Thread, GuestRIP);
}
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
@@ -180,21 +193,19 @@ namespace FEXCore::Context {
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
// User's responsibility to deallocate this.
FEXCore::IR::RegisterAllocationData* RAData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo);
struct CompileCodeResult {
void* CompiledCode;
FEXCore::IR::IRListView* IRData;
FEXCore::Core::DebugData* DebugData;
// User's responsibility to deallocate this.
FEXCore::IR::RegisterAllocationData* RAData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
bool GeneratedIR;
uint64_t StartAddr;
uint64_t Length;
@@ -207,7 +218,7 @@ namespace FEXCore::Context {
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
@@ -230,7 +241,7 @@ namespace FEXCore::Context {
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
/**
* @brief Initializes the TLS data for a thread
* @brief Initializes TID, PID and TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
@@ -297,8 +308,12 @@ namespace FEXCore::Context {
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
void MarkMemoryShared();
bool IsTSOEnabled() { return (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled; }
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread);
private:
/**
@@ -317,17 +332,15 @@ namespace FEXCore::Context {
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* State);
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void WaitForIdleWithTimeout();
void NotifyPause();
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length);
FEXCore::CodeLoader *LocalLoader{};
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
// Entry Cache
uint64_t StartingRIP;
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
@@ -335,7 +348,13 @@ namespace FEXCore::Context {
std::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::unordered_map<uint64_t, std::tuple<std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)>, void *, void *>> CustomIRHandlers;
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
FEXCore::CPU::DispatcherConfig DispatcherConfig;
};
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args);
@@ -1,14 +1,14 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <aarch64/cpu-aarch64.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <atomic>
#include <stdint.h>
#include <signal.h>
#include "aarch64/cpu-aarch64.h"
#include <csignal>
#include <cstdint>
namespace FEXCore::ArchHelpers::Arm64 {
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock, TYPE_HAS_SPLIT_LOCKS);
@@ -7,12 +7,14 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include "aarch64/cpu-aarch64.h"
#include "cpu-features.h"
#include "aarch64/instructions-aarch64.h"
#include "utils-vixl.h"
#include <aarch64/cpu-aarch64.h>
#include <aarch64/instructions-aarch64.h>
#include <cpu-features.h>
#include <utils-vixl.h>
#include <array>
#include <tuple>
#include <utility>
namespace FEXCore::CPU {
#define STATE x28
@@ -309,19 +311,6 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
add(sp, sp, SPOffset);
}
void Arm64Emitter::ResetStack() {
if (SpillSlots == 0)
return;
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
@@ -329,122 +318,4 @@ void Arm64Emitter::Align16B() {
}
}
uint64_t Arm64Emitter::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return Dispatcher->ExitFunctionLinkerAddress;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void Arm64Emitter::InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.GetCode();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
Arm64Emitter::NamedSymbolLiteralPair Arm64Emitter::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
Arm64Emitter::NamedSymbolLiteralPair Lit {
.Lit = Literal(Pointer),
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64Emitter::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
place(&Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
void Arm64Emitter::InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.GetCode();
LoadConstant(Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
bool Arm64Emitter::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
Literal<uint64_t> Lit(Pointer);
place(&Lit);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
}
return true;
}
}
@@ -3,16 +3,17 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "platform-vixl.h"
#include "FEXCore/Config/Config.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/constants-aarch64.h>
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <platform-vixl.h>
#include <FEXCore/Config/Config.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <cstddef>
#include <cstdint>
#include <utility>
namespace FEXCore::CPU {
@@ -63,8 +64,6 @@ class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
FEXCore::Context::Context *EmitterCTX;
vixl::aarch64::CPU CPU;
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
@@ -83,68 +82,7 @@ protected:
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void ResetStack();
void Align16B();
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Literal<uint64_t> Lit;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint64_t GuestEntry{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
+87
View File
@@ -0,0 +1,87 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
namespace FEXCore {
namespace CPU {
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
}
return CurrentCodeBuffer;
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::mmap(nullptr, Buffer.Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (ThreadState->CTX->Config.GlobalJITNaming()) {
ThreadState->CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto &Buffer: CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
}
}
return false;
}
}
}
+2 -2
View File
@@ -416,7 +416,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
Res.ecx =
(1 << 0) | // SSE3
(0 << 1) | // PCLMULQDQ
(1 << 1) | // PCLMULQDQ
(1 << 2) | // DS area supports 64bit layout
(1 << 3) | // MWait
(0 << 4) | // DS-CPL
@@ -446,7 +446,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(SUPPORTS_AVX << 28) | // AVX
(0 << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(0 << 31); // Hypervisor always returns zero
(1 << 31); // Hypervisor always returns one
Res.edx =
(1 << 0) | // FPU
+312 -235
View File
@@ -7,6 +7,7 @@ desc: Glues Frontend, OpDispatcher and IR Opts & Compilation, LookupCache, Dispa
$end_info$
*/
#include <cstdint>
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Core.h"
@@ -17,6 +18,7 @@ $end_info$
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/Interpreter/InterpreterCore.h"
#include "Interface/Core/JIT/JITCore.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
@@ -32,6 +34,7 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
@@ -49,7 +52,6 @@ $end_info$
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <cstdint>
#include <filesystem>
#include <functional>
#include <fstream>
@@ -74,6 +76,7 @@ $end_info$
#include <vector>
#include <xxhash.h>
namespace FEXCore::CPU {
bool CreateCPUCore(FEXCore::Context::Context *CTX) {
// This should be used for generating things that are shared between threads
@@ -177,15 +180,15 @@ namespace FEXCore::Context {
// Initialize default CPU state
NewThreadState.rip = ~0ULL;
for (int i = 0; i < 16; ++i) {
NewThreadState.gregs[i] = 0;
for (auto& greg : NewThreadState.gregs) {
greg = 0;
}
for (int i = 0; i < 16; ++i) {
NewThreadState.xmm[i][0] = 0xDEADBEEFULL;
NewThreadState.xmm[i][1] = 0xBAD0DAD1ULL;
for (auto& xmm : NewThreadState.xmm) {
xmm[0] = 0xDEADBEEFULL;
xmm[1] = 0xBAD0DAD1ULL;
}
memset(NewThreadState.flags, 0, 32);
memset(NewThreadState.flags, 0, Core::CPUState::NUM_EFLAG_BITS);
NewThreadState.flags[1] = 1;
NewThreadState.flags[9] = 1;
NewThreadState.FCW = 0x37F;
@@ -193,19 +196,22 @@ namespace FEXCore::Context {
return NewThreadState;
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
// Initialize the CPU core signal handlers
FEXCore::Core::InternalThreadState* Context::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
FEXCore::CPU::InitializeInterpreterSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetInterpreterBackendFeatures();
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetX86JITBackendFeatures();
#elif (_M_ARM_64 && JIT_ARM64)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetArm64JITBackendFeatures();
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
@@ -218,6 +224,32 @@ namespace FEXCore::Context {
break;
}
DispatcherConfig.StaticRegisterAllocation = Config.StaticRegisterAllocation && BackendFeatures.SupportsStaticRegisterAllocation;
#if (_M_X86_64)
Dispatcher = FEXCore::CPU::Dispatcher::CreateX86(this, DispatcherConfig);
#elif (_M_ARM_64)
Dispatcher = FEXCore::CPU::Dispatcher::CreateArm64(this, DispatcherConfig);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled with an unknown target");
#endif
// Initialize common signal handlers
auto PauseHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSignalPause(Thread, Signal, info, ucontext);
};
SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, PauseHandler, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
return Thread->CTX->Dispatcher->HandleGuestSignal(Thread, Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
// Initialize GDBServer after the signal handlers are installed
// It may install its own handlers that need to be executed AFTER the CPU cores
if (Config.GdbServer) {
@@ -229,7 +261,6 @@ namespace FEXCore::Context {
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
LocalLoader = Loader;
using namespace FEXCore::Core;
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
@@ -238,9 +269,9 @@ namespace FEXCore::Context {
// We are the parent thread
ParentThread = Thread;
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = Loader->GetStackPointer();
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = StartingRIP = Loader->DefaultRIP();
Thread->CurrentFrame->State.rip = InitialRIP;
InitializeThreadData(Thread);
return Thread;
@@ -258,7 +289,7 @@ namespace FEXCore::Context {
}
void Context::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
Thread->CPUBackend->CallbackPtr(Thread->CurrentFrame, RIP);
Thread->CTX->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
}
void Context::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
@@ -352,10 +383,7 @@ namespace FEXCore::Context {
// Walk the threads and tell them to clear their caches
// Useful when our block size is set to a large number and we need to step a single instruction
for (auto &Thread : Threads) {
// Wait for thread to be fully constructed
// XXX: Look into thread partial construction issues
while(Thread->RunningEvents.WaitingToStart.load()) ;
ClearCodeCache(Thread, true);
ClearCodeCache(Thread);
}
}
CoreRunningMode PreviousRunningMode = this->Config.RunningMode;
@@ -448,22 +476,6 @@ namespace FEXCore::Context {
void Context::InitializeThreadData(FEXCore::Core::InternalThreadState *Thread) {
Thread->CPUBackend->Initialize();
auto IRHandler = [Thread](uint64_t Addr, IR::IREmitter *IR) -> void {
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(IR);
Core::LocalIREntry Entry = {Addr, 0ULL,
decltype(Entry.IR)(IR->CreateIRCopy()),
decltype(Entry.RAData)(Thread->PassManager->HasPass("RA")
? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData()
: nullptr),
decltype(Entry.DebugData)(new Core::DebugData())
};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.insert({Addr, std::move(Entry)});
};
LocalLoader->AddIR(IRHandler);
}
struct ExecutionThreadHandler {
@@ -502,50 +514,51 @@ namespace FEXCore::Context {
Thread->StartRunning.NotifyAll();
}
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* State) {
State->OpDispatcher = std::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
State->OpDispatcher->SetMultiblock(Config.Multiblock);
State->LookupCache = std::make_unique<FEXCore::LookupCache>(this);
State->FrontendDecoder = std::make_unique<FEXCore::Frontend::Decoder>(this);
State->PassManager = std::make_unique<FEXCore::IR::PassManager>();
State->PassManager->RegisterExitHandler([this]() {
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = std::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = std::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = std::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->PassManager = std::make_unique<FEXCore::IR::PassManager>();
Thread->PassManager->RegisterExitHandler([this]() {
Stop(false /* Ignore current thread */);
});
State->CTX = this;
Thread->CurrentFrame->Pointers.Common.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->Pointers.Common.L2Pointer = Thread->LookupCache->GetPagePointer();
#if _M_ARM_64
bool DoSRA = State->CTX->Config.StaticRegisterAllocation;
#else
bool DoSRA = false;
#endif
Dispatcher->InitThreadPointers(Thread);
State->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT, DoSRA);
State->PassManager->AddDefaultValidationPasses();
Thread->CTX = this;
State->PassManager->RegisterSyscallHandler(SyscallHandler);
bool DoSRA = DispatcherConfig.StaticRegisterAllocation;
Thread->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT, DoSRA);
Thread->PassManager->AddDefaultValidationPasses();
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
// Create CPU backend
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
State->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, State);
Thread->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, Thread);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
State->PassManager->InsertRegisterAllocationPass(DoSRA);
Thread->PassManager->InsertRegisterAllocationPass(DoSRA);
#if (_M_X86_64 && JIT_X86_64)
State->CPUBackend = FEXCore::CPU::CreateX86JITCore(this, State);
Thread->CPUBackend = FEXCore::CPU::CreateX86JITCore(this, Thread);
#elif (_M_ARM_64 && JIT_ARM64)
State->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, State);
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
State->CPUBackend = CustomCPUFactory(this, State);
Thread->CPUBackend = CustomCPUFactory(this, Thread);
break;
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
@@ -554,14 +567,7 @@ namespace FEXCore::Context {
}
FEXCore::Core::InternalThreadState* Context::CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState *Thread{};
// Grab the new thread object
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
Thread = Threads.emplace_back(new FEXCore::Core::InternalThreadState{});
Thread->ThreadManager.TID = ++ThreadID;
}
FEXCore::Core::InternalThreadState *Thread = new FEXCore::Core::InternalThreadState{};
// Copy over the new thread state to the new object
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
@@ -573,13 +579,19 @@ namespace FEXCore::Context {
InitializeCompiler(Thread);
InitializeThreadData(Thread);
// Insert after the Thread object has been fully initialized
{
std::lock_guard lk(ThreadCreationMutex);
Threads.push_back(Thread);
}
return Thread;
}
void Context::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
// remove new thread object
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
std::lock_guard lk(ThreadCreationMutex);
auto It = std::find(Threads.begin(), Threads.end(), Thread);
LOGMAN_THROW_A_FMT(It != Threads.end(), "Thread wasn't in Threads");
@@ -634,14 +646,11 @@ namespace FEXCore::Context {
FEXCore::Threads::Thread::CleanupAfterFork();
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length) {
// Only call MarkGuestExecutableRange if new pages are marked as containing code
if (Thread->LookupCache->AddBlockMapping(Address, Ptr, Start, Length)) {
Thread->CTX->SyscallHandler->MarkGuestExecutableRange(Start, Length);
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr) {
Thread->LookupCache->AddBlockMapping(Address, Ptr);
}
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache) {
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
{
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
@@ -651,18 +660,16 @@ namespace FEXCore::Context {
Thread->LookupCache->ClearCache();
Thread->CPUBackend->ClearCache();
if (AlsoClearIRCache) {
Thread->LocalIRCache.clear();
}
Thread->DebugStore.clear();
}
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, IR::IREmitter *IREmitter, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
if (DumpIRStr =="stderr") {
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpIRStr =="stderr" || DumpIRStr =="no") {
f = stderr;
}
else if (DumpIRStr =="stdout") {
@@ -676,7 +683,7 @@ namespace FEXCore::Context {
if (f) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
@@ -686,12 +693,12 @@ namespace FEXCore::Context {
}
};
static void ValidateIR(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
static void ValidateIR(FEXCore::Context::Context *ctx, IR::IREmitter *IREmitter) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
compaction->Run(IREmitter);
auto NewIR = IREmitter->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
FEXCore::Utils::PooledAllocatorMalloc Allocator;
@@ -710,151 +717,169 @@ namespace FEXCore::Context {
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
uint8_t const *GuestCode{};
GuestCode = reinterpret_cast<uint8_t const*>(GuestRIP);
bool HadDispatchError {false};
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
uint64_t TotalInstructions {0};
uint64_t TotalInstructionsLength {0};
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP);
auto CodeBlocks = Thread->FrontendDecoder->GetDecodedBlocks();
std::shared_lock lk(CustomIRMutex);
auto Handler = CustomIRHandlers.find(GuestRIP);
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
TotalInstructionsLength = 1;
std::get<0>(Handler->second)(GuestRIP, Thread->OpDispatcher.get());
lk.unlock();
} else {
lk.unlock();
uint8_t const *GuestCode{};
GuestCode = reinterpret_cast<uint8_t const*>(GuestRIP);
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks);
bool HadDispatchError {false};
const uint8_t GPRSize = GetGPRSize();
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
FEXCore::Frontend::Decoder::DecodedBlocks const &Block = CodeBlocks->at(j);
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
uint64_t InstsInBlock = Block.NumInstructions;
for (size_t i = 0; i < InstsInBlock; ++i) {
FEXCore::X86Tables::X86InstInfo const* TableInfo {nullptr};
FEXCore::X86Tables::DecodedInst const* DecodedInfo {nullptr};
TableInfo = Block.DecodedInstructions[i].TableInfo;
DecodedInfo = &Block.DecodedInstructions[i];
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(ExistingCodePtr[0], ExistingCodePtr[1], (uintptr_t)ExistingCodePtr - GuestRIP, DecodedInfo->InstSize);
auto InvalidateCodeCond = Thread->OpDispatcher->_CondJump(CodeChanged);
auto CurrentBlock = Thread->OpDispatcher->GetCurrentBlock();
auto CodeWasChangedBlock = Thread->OpDispatcher->CreateNewCodeBlockAtEnd();
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_RemoveThreadCodeEntry();
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, [Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
Thread->CTX->SyscallHandler->MarkGuestExecutableRange(Start, Length);
}
});
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
Thread->OpDispatcher->HandledLock = false;
Thread->OpDispatcher->ResetDecodeFailure();
std::invoke(Fn, Thread->OpDispatcher, DecodedInfo);
if (Thread->OpDispatcher->HadDecodeFailure()) {
HadDispatchError = true;
}
else {
if (Thread->OpDispatcher->HandledLock != IsLocked) {
HadDispatchError = true;
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name ?: "UND");
}
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
++TotalInstructions;
}
}
else {
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
}
auto CodeBlocks = Thread->FrontendDecoder->GetDecodedBlocks();
// If we had a dispatch error then leave early
if (HadDispatchError) {
if (TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return { nullptr, nullptr, 0, 0, 0, 0 };
}
else {
const uint8_t GPRSize = GetGPRSize();
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks);
// We had some instructions. Early exit
const uint8_t GPRSize = GetGPRSize();
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
FEXCore::Frontend::Decoder::DecodedBlocks const &Block = CodeBlocks->at(j);
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
uint64_t InstsInBlock = Block.NumInstructions;
for (size_t i = 0; i < InstsInBlock; ++i) {
FEXCore::X86Tables::X86InstInfo const* TableInfo {nullptr};
FEXCore::X86Tables::DecodedInst const* DecodedInfo {nullptr};
TableInfo = Block.DecodedInstructions[i].TableInfo;
DecodedInfo = &Block.DecodedInstructions[i];
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
if (ExtendedDebugInfo) {
Thread->OpDispatcher->_GuestOpcode(Block.Entry + BlockInstructionsLength - GuestRIP);
}
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(ExistingCodePtr[0], ExistingCodePtr[1], (uintptr_t)ExistingCodePtr - GuestRIP, DecodedInfo->InstSize);
auto InvalidateCodeCond = Thread->OpDispatcher->_CondJump(CodeChanged);
auto CurrentBlock = Thread->OpDispatcher->GetCurrentBlock();
auto CodeWasChangedBlock = Thread->OpDispatcher->CreateNewCodeBlockAtEnd();
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_RemoveThreadCodeEntry();
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
}
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
Thread->OpDispatcher->HandledLock = false;
Thread->OpDispatcher->ResetDecodeFailure();
std::invoke(Fn, Thread->OpDispatcher, DecodedInfo);
if (Thread->OpDispatcher->HadDecodeFailure()) {
HadDispatchError = true;
}
else {
if (Thread->OpDispatcher->HandledLock != IsLocked) {
HadDispatchError = true;
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name ?: "UND");
}
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
++TotalInstructions;
}
}
else {
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
}
// If we had a dispatch error then leave early
if (HadDispatchError) {
if (TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return { nullptr, nullptr, 0, 0, 0, 0 };
}
else {
const uint8_t GPRSize = GetGPRSize();
// We had some instructions. Early exit
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
break;
}
}
if (Thread->OpDispatcher->FinishOp(DecodedInfo->PC + DecodedInfo->InstSize, i + 1 == InstsInBlock)) {
break;
}
}
if (Thread->OpDispatcher->FinishOp(DecodedInfo->PC + DecodedInfo->InstSize, i + 1 == InstsInBlock)) {
break;
}
}
Thread->OpDispatcher->Finalize();
Thread->FrontendDecoder->DelayedDisownBuffer();
}
Thread->OpDispatcher->Finalize();
IR::IREmitter *IREmitter = Thread->OpDispatcher.get();
auto ShouldDump = Thread->CTX->Config.DumpIR() != "no" || Thread->OpDispatcher->ShouldDump;
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, nullptr);
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
}
if (Thread->CTX->Config.ValidateIRarser) {
ValidateIR(this, Thread);
ValidateIR(this, IREmitter);
}
}
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(Thread->OpDispatcher.get());
Thread->PassManager->Run(IREmitter);
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::IFmt("IR 0x{:x}:\n{}\n@@@@@\n", GuestRIP, out.str());
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
auto IRList = Thread->OpDispatcher->CreateIRCopy();
auto IRList = IREmitter->CreateIRCopy();
Thread->OpDispatcher->DelayedDisownBuffer();
Thread->FrontendDecoder->DelayedDisownBuffer();
IREmitter->DelayedDisownBuffer();
return {
.IRList = IRList,
.RAData = RAData.release(),
.RAData = std::move(RAData),
.TotalInstructions = TotalInstructions,
.TotalInstructionsLength = TotalInstructionsLength,
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
@@ -865,27 +890,11 @@ namespace FEXCore::Context {
Context::CompileCodeResult Context::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
FEXCore::IR::RegisterAllocationData *RAData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
bool GeneratedIR {};
uint64_t StartAddr {};
uint64_t Length {};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
// Do we already have this in the IR cache?
auto LocalEntry = Thread->LocalIRCache.find(GuestRIP);
if (LocalEntry != Thread->LocalIRCache.end()) {
// Entry already exists
// pull in the data
IRList = LocalEntry->second.IR.get();
DebugData = LocalEntry->second.DebugData.get();
RAData = LocalEntry->second.RAData.get();
StartAddr = LocalEntry->second.StartAddr;
Length = LocalEntry->second.Length;
GeneratedIR = false;
}
// JIT Code object cache lookup
if (CodeObjectCacheService) {
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
@@ -893,25 +902,33 @@ namespace FEXCore::Context {
auto CompiledCode = Thread->CPUBackend->RelocateJITObjectCode(GuestRIP, CodeCacheEntry);
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.IRData = nullptr, // No IR data generated
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.RAData = nullptr, // No RA data generated
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
.StartAddr = 0, // Unused
.Length = 0, // Unused
.CompiledCode = CompiledCode,
.IRData = nullptr, // No IR data generated
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.RAData = nullptr, // No RA data generated
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
.StartAddr = 0, // Unused
.Length = 0, // Unused
};
}
}
}
if (SourcecodeResolver && Config.GDBSymbols()) {
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
if (AOTIRCacheEntry.Entry && !AOTIRCacheEntry.Entry->ContainsCode) {
AOTIRCacheEntry.Entry->SourcecodeMap =
SourcecodeResolver->GenerateMap(AOTIRCacheEntry.Entry->Filename, AOTIRCacheEntry.Entry->FileId);
}
}
// AOT IR bookkeeping and cache
{
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
if (_GeneratedIR) {
// Setup pointers to internal structures
IRList = IRCopy;
RAData = RACopy;
RAData = std::move(RACopy);
DebugData = DebugDataCopy;
StartAddr = _StartAddr;
Length = _Length;
@@ -921,11 +938,11 @@ namespace FEXCore::Context {
if (IRList == nullptr) {
// Generate IR + Meta Info
auto [IRCopy, RACopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] = GenerateIR(Thread, GuestRIP);
auto [IRCopy, RACopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] = GenerateIR(Thread, GuestRIP, Config.GDBSymbols());
// Setup pointers to internal structures
IRList = IRCopy;
RAData = RACopy;
RAData = std::move(RACopy);
DebugData = new FEXCore::Core::DebugData();
StartAddr = _StartAddr;
Length = _Length;
@@ -942,10 +959,10 @@ namespace FEXCore::Context {
}
// Attempt to get the CPU backend to compile this code
return {
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData),
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get(), GetGdbServerStatus()),
.IRData = IRList,
.DebugData = DebugData,
.RAData = RAData,
.RAData = std::move(RAData),
.GeneratedIR = GeneratedIR,
.StartAddr = StartAddr,
.Length = Length,
@@ -966,8 +983,8 @@ namespace FEXCore::Context {
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
// Needs to be held for SMC interactions around concurrent compile and invalidation hazards
auto InvalidationLk = Thread->CTX->SyscallHandler->CompileCodeLock(GuestRIP);
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
std::shared_lock lk(CodeInvalidationMutex);
// Is the code in the cache?
// The backends only check L1 and L2, not L3
@@ -978,16 +995,14 @@ namespace FEXCore::Context {
void *CodePtr {};
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
FEXCore::IR::RegisterAllocationData *RAData {};
bool GeneratedIR {};
uint64_t StartAddr {}, Length {};
auto [Code, IR, Data, RA, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP);
auto [Code, IR, Data, RAData, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP);
CodePtr = Code;
IRList = IR;
DebugData = Data;
RAData = RA;
GeneratedIR = Generated;
StartAddr = _StartAddr;
Length = _Length;
@@ -998,22 +1013,25 @@ namespace FEXCore::Context {
// The core managed to compile the code.
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = reinterpret_cast<uint8_t *>(CodePtr);
if (DebugData) {
auto GuestRIPLookup = this->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
auto BlockBasePtr = FragmentBasePtr + Subblock.HostCodeOffset;
if (GuestRIPLookup.Entry) {
Symbols.Register(CodePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.Offset);
Symbols.Register(BlockBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register((void*)Subblock.HostCodeStart, GuestRIP, Subblock.HostCodeSize);
}
Symbols.Register(BlockBasePtr, GuestRIP, Subblock.HostCodeSize);
}
}
} else {
if (GuestRIPLookup.Entry) {
Symbols.Register(CodePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.Offset);
Symbols.Register(FragmentBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(CodePtr, GuestRIP, DebugData->HostCodeSize);
Symbols.Register(FragmentBasePtr, GuestRIP, DebugData->HostCodeSize);
}
}
}
@@ -1046,7 +1064,7 @@ namespace FEXCore::Context {
GuestRIP,
StartAddr,
Length,
RAData,
std::move(RAData),
IRList,
DebugData,
GeneratedIR)) {
@@ -1055,7 +1073,8 @@ namespace FEXCore::Context {
}
// Insert to lookup cache
AddBlockMapping(Thread, GuestRIP, CodePtr, StartAddr, Length);
// Pages containing this block are added via AddBlockExecutableRange before each page gets accessed in the frontend
AddBlockMapping(Thread, GuestRIP, CodePtr);
return (uintptr_t)CodePtr;
}
@@ -1083,7 +1102,7 @@ namespace FEXCore::Context {
Thread->RunningEvents.Running = true;
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->CTX->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
@@ -1132,21 +1151,79 @@ namespace FEXCore::Context {
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard<std::mutex> lk(CTX->ThreadCreationMutex);
std::lock_guard lk(CTX->ThreadCreationMutex);
for (auto &Thread : CTX->Threads) {
if (Thread->RunningEvents.Running.load()) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
std::unique_lock CodeInvalidationLock(CTX->CodeInvalidationMutex);
InvalidateGuestCodeRange(CTX, Start, Length);
CallAfter(Start, Length);
}
void Context::MarkMemoryShared() {
if (!IsMemoryShared) {
IsMemoryShared = true;
if (Config.TSOAutoMigration) {
LogMan::Msg::IFmt("Migrating to shared memory mode");
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
LogMan::Throw::AFmt(Threads.size() == 1, "First MarkMemoryShared called must be before creating any threads");
auto Thread = Threads[0];
// Only the lookup cache is cleared here, so that old code can keep running until next compilation
std::lock_guard<std::recursive_mutex> lkLookupCache(Thread->LookupCache->WriteLock);
Thread->LookupCache->ClearCache();
// DebugStore also needs to be cleared
Thread->DebugStore.clear();
}
}
}
void MarkMemoryShared(FEXCore::Context::Context *CTX) {
CTX->MarkMemoryShared();
}
void Context::RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.erase(GuestRIP);
Thread->DebugStore.erase(GuestRIP);
Thread->LookupCache->Erase(GuestRIP);
}
CustomIRResult Context::AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::unique_lock lk(CustomIRMutex);
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, std::tuple(Handler, Creator, Data));
if (!InsertedIterator.second) {
const auto &[fn, Creator, Data] = InsertedIterator.first->second;
return CustomIRResult(std::move(lk), Creator, Data);
} else {
lk.unlock();
return CustomIRResult(std::move(lk), 0, 0);
}
}
void Context::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(this, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
CustomIRHandlers.erase(Entrypoint);
});
}
// Debug interface
void Context::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
uint64_t RIPBackup = Thread->CurrentFrame->State.rip;
@@ -1171,8 +1248,8 @@ namespace FEXCore::Context {
bool Context::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
std::lock_guard<std::recursive_mutex> lk(ParentThread->LookupCache->WriteLock);
auto it = ParentThread->LocalIRCache.find(RIP);
if (it == ParentThread->LocalIRCache.end()) {
auto it = ParentThread->DebugStore.find(RIP);
if (it == ParentThread->DebugStore.end()) {
return false;
}
@@ -1215,4 +1292,4 @@ namespace FEXCore::Context {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
Thread->FrontendDecoder->SetSectionMaxAddress(SectionMaxAddress);
}
}
}
@@ -16,20 +16,23 @@
#include <array>
#include <bit>
#include <cmath>
#include <cstddef>
#include <cstdint>
#include <memory>
#include <stddef.h>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/constants-aarch64.h>
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <code-buffer-vixl.h>
#include <platform-vixl.h>
#include <sys/syscall.h>
#include <unistd.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
MemOperand(STATE, offsetof(FEXCore::Core::STATE_TYPE, FIELD))
namespace FEXCore::CPU {
using namespace vixl;
@@ -38,12 +41,12 @@ using namespace vixl::aarch64;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
, config(config) {
SetAllowAssembler(true);
DispatchPtr = GetCursorAddress<CPUBackend::AsmDispatch>();
DispatchPtr = GetCursorAddress<AsmDispatch>();
// while (true) {
// Ptr = FindBlock(RIP)
@@ -53,12 +56,9 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Ptr();
// }
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
Literal l_ExitFunctionLink {config.ExitFunctionLink};
Literal l_ExitFunctionLinkThis {config.ExitFunctionLinkThis};
// Push all the register we need to save
PushCalleeSavedRegisters();
@@ -71,11 +71,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
str(x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled) {
if (config.StaticRegisterAllocation) {
FillStaticRegs();
}
@@ -92,11 +92,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
ldr(x2, STATE_PTR(CpuStateFrame, State.rip));
auto RipReg = x2;
// L1 Cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -104,21 +104,17 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
b(&CallBlock);
}
br(x3);
// L1C check failed, do a full lookup
bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, &l_PagePtr);
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, VirtualMemorySize - 1);
}
@@ -157,44 +153,21 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
// Jump to the block
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
bind(&CallBlock);
mov(x0, STATE);
blr(x3);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, &l_CTX);
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
} else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
br(x3);
}
}
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
@@ -209,7 +182,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
constexpr bool SignalSafeCompile = true;
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
if (SignalSafeCompile) {
@@ -231,11 +204,10 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
svc(0);
}
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
mov(x0, STATE);
mov(x1, lr);
ldr(x3, &l_ExitFunctionLink);
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
blr(x3);
if (SignalSafeCompile) {
@@ -256,7 +228,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
mov(x0, x4);
}
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
br(x0);
}
@@ -265,7 +237,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
bind(&NoBlock);
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
if (SignalSafeCompile) {
@@ -312,7 +284,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
add(sp, sp, 16);
}
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
b(&LoopTop);
@@ -331,7 +303,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
hlt(0);
@@ -342,18 +314,17 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
LoadConstant(x0, reinterpret_cast<uint64_t>(&SynchronousFaultData));
LoadConstant(w1, 1);
strb(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)));
strb(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException));
LoadConstant(w1, X86State::X86_TRAPNO_OF);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo));
LoadConstant(w1, 0x80);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code));
LoadConstant(x1, 0);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)));
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code));
// hlt/udf = SIGILL
// brk = SIGTRAP
@@ -364,7 +335,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
if (config.StaticRegisterAllocation)
SpillStaticRegs();
bind(&ThreadPauseHandler);
@@ -399,7 +370,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// On return to the thunk, the thunk can get whatever its return value is from the thread context depending on ABI handling on its end
// When the thunk itself returns, it'll do its regular return logic there
// void ReentrantCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
CallbackPtr = GetCursorAddress<CPUBackend::JITCallback>();
CallbackPtr = GetCursorAddress<JITCallback>();
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
@@ -408,46 +379,39 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
mov(STATE, x0);
// Make sure to adjust the refcounter so we don't clear the cache now
LoadConstant(x0, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
ldr(w2, MemOperand(x0));
ldr(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
add(w2, w2, 1);
str(w2, MemOperand(x0));
str(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
ldr(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
str(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
str(x1, STATE_PTR(CpuStateFrame, State.rip));
// load static regs
if (SRAEnabled)
if (config.StaticRegisterAllocation)
FillStaticRegs();
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
// Long division helpers
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
{
LUDIVHandler = GetCursorAddress<uint64_t>();
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIV)));
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
SpillStaticRegs();
blr(x3);
@@ -462,11 +426,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
LDIVHandler = GetCursorAddress<uint64_t>();
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIV)));
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
SpillStaticRegs();
blr(x3);
@@ -481,11 +445,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
LUREMHandler = GetCursorAddress<uint64_t>();
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREM)));
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
SpillStaticRegs();
blr(x3);
@@ -500,11 +464,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
LREMHandler = GetCursorAddress<uint64_t>();
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREM)));
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
SpillStaticRegs();
blr(x3);
@@ -518,12 +482,9 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ret();
}
place(&l_PagePtr);
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
place(&l_ExitFunctionLink);
place(&l_ExitFunctionLinkThis);
FinalizeCode();
@@ -539,55 +500,114 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Pointers.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
Pointers.LUDIVHandler = LUDIVHandler;
Pointers.LDIVHandler = LDIVHandler;
Pointers.LUREMHandler = LUREMHandler;
Pointers.LREMHandler = LREMHandler;
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext, uint32_t IgnoreMask) {
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline, destination buffer is set before use
static thread_local vixl::aarch64::Assembler emit((uint8_t*)&emit, 1);
size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxGDBPauseCheckSize);
vixl::CodeBufferCheckScope scope(&emit, MaxGDBPauseCheckSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(FEXCore::Context::Context::Config.RunningMode) == 4, "This is expected to be size of 4");
emit.ldr(x0, STATE_PTR(CpuStateFrame, Thread)); // Get thread
emit.ldr(x0, MemOperand(x0, offsetof(FEXCore::Core::InternalThreadState, CTX))); // Get Context
emit.ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then we don't need to stop
emit.cbz(w0, &RunBlock);
{
Literal l_GuestRIP {GuestRIP};
// Make sure RIP is syncronized to the context
emit.ldr(x0, &l_GuestRIP);
emit.str(x0, STATE_PTR(CpuStateFrame, State.rip));
// Stop the thread
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.br(x0);
emit.place(&l_GuestRIP);
}
emit.bind(&RunBlock);
emit.FinalizeCode();
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
return UsedBytes;
}
size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "GenerateInterpreterTrampoline dispatcher does not support SRA");
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
vixl::CodeBufferCheckScope scope(&emit, MaxInterpreterTrampolineSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label InlineIRData;
emit.mov(x0, STATE);
emit.adr(x1, &InlineIRData);
emit.ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.blr(x3);
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.br(x0);
emit.bind(&InlineIRData);
emit.FinalizeCode();
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
return UsedBytes;
}
void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {
for(int i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
// Skip this one, it's already spilled
continue;
}
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
for(int i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&ThreadState->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
memcpy(&Thread->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
}
}
#ifdef _M_ARM_64
void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Dispatcher = std::make_unique<Arm64Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
auto &AArch64 = Thread->CurrentFrame->Pointers.AArch64;
AArch64.LUDIVHandler = LUDIVHandlerAddress;
AArch64.LDIVHandler = LDIVHandlerAddress;
AArch64.LUREMHandler = LUREMHandlerAddress;
AArch64.LREMHandler = LREMHandlerAddress;
}
}
#endif
std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<Arm64Dispatcher>(CTX, Config);
}
}
@@ -15,10 +15,21 @@ namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
protected:
void SpillSRA(void *ucontext, uint32_t IgnoreMask) override;
void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) override;
private:
// Long division helpers
uint64_t LUDIVHandlerAddress{};
uint64_t LDIVHandlerAddress{};
uint64_t LUREMHandlerAddress{};
uint64_t LREMHandlerAddress{};
DispatcherConfig config;
};
}
@@ -1,4 +1,5 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
@@ -39,7 +40,7 @@ void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuS
ctx->IdleWaitCV.notify_all();
}
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, void *ucontext) {
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -64,7 +65,7 @@ ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, vo
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, ThreadState->CurrentFrame, sizeof(FEXCore::Core::CPUState));
memcpy(&Context->GuestState, Thread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
@@ -81,13 +82,13 @@ ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, vo
Context->SigInfoLocation = 0;
// Store fault to top status and then reset it
Context->FaultToTopAndGeneratedException = SynchronousFaultData.FaultToTopAndGeneratedException;
SynchronousFaultData.FaultToTopAndGeneratedException = false;
Context->FaultToTopAndGeneratedException = Thread->CurrentFrame->SynchronousFaultData.FaultToTopAndGeneratedException;
Thread->CurrentFrame->SynchronousFaultData.FaultToTopAndGeneratedException = false;
return Context;
}
void Dispatcher::RestoreThreadState(void *ucontext) {
void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext) {
uint64_t OldSP{};
if (CTX->Config.Core() == FEXCore::Config::CONFIG_IRJIT) {
OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -102,13 +103,13 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(ThreadState->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
memcpy(Thread->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
if (Context->UContextLocation) {
auto Frame = ThreadState->CurrentFrame;
auto Frame = Thread->CurrentFrame;
if (Context->Flags &ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT) {
// XXX: Unsupported since it needs state reconstruction
@@ -132,7 +133,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP];
// XXX: Full context setting
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
@@ -190,7 +191,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
// XXX: Full context setting
// First 32-bytes of flags is EFLAGS broken out
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
@@ -218,7 +219,7 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&Frame->State.mm[i], &fpstate->_st[i], 10);
}
@@ -273,14 +274,14 @@ static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
return 0;
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
auto ContextBackup = StoreThreadState(Signal, ucontext);
bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
auto ContextBackup = StoreThreadState(Thread, Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
auto Frame = Thread->CurrentFrame;
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
++Thread->CurrentFrame->SignalHandlerRefCounter;
uint64_t OldPC = ArchHelpers::Context::GetPc(ucontext);
// Set the new PC
@@ -299,7 +300,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// We are going to be returning to the top of the dispatcher which will fill again
// Otherwise we might load garbage
if (SRAEnabled) {
if (IsAddressInJITCode(OldPC, false)) {
if (Thread->CPUBackend->IsAddressInCodeBuffer(OldPC)) {
uint32_t IgnoreMask{};
#ifdef _M_ARM_64
if (Frame->InSyscallInfo != 0) {
@@ -323,11 +324,11 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#endif
// We are in jit, SRA must be spilled
SpillSRA(ucontext, IgnoreMask);
SpillSRA(Thread, ucontext, IgnoreMask);
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT;
} else {
if (!IsAddressInJITCode(OldPC, true)) {
if (!IsAddressInDispatcher(OldPC)) {
// This is likely to cause issues but in some cases it isn't fatal
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
@@ -409,11 +410,11 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
*guest_siginfo = *HostSigInfo;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = SynchronousFaultData.err_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = Frame->SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = Frame->SynchronousFaultData.err_code;
// Overwrite si_code
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_siginfo->si_code = Thread->CurrentFrame->SynchronousFaultData.si_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
@@ -500,9 +501,9 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = SynchronousFaultData.err_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = Frame->SynchronousFaultData.TrapNo;
guest_siginfo->si_code = Frame->SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = Frame->SynchronousFaultData.err_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
@@ -528,7 +529,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#undef COPY_REG
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&fpstate->_st[i], &Frame->State.mm[i], 10);
}
@@ -633,44 +634,44 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
return true;
}
bool Dispatcher::HandleSIGILL(int Signal, void *info, void *ucontext) {
bool Dispatcher::HandleSIGILL(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) {
if (ArchHelpers::Context::GetPc(ucontext) == SignalHandlerReturnAddress) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
return true;
}
if (ArchHelpers::Context::GetPc(ucontext) == PauseReturnInstruction) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
return true;
}
return false;
}
bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = ThreadState->SignalReason.load();
auto Frame = ThreadState->CurrentFrame;
bool Dispatcher::HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = Thread->SignalReason.load();
auto Frame = Thread->CurrentFrame;
if (SignalReason == FEXCore::Core::SignalEvent::Pause) {
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
StoreThreadState(Thread, Signal, ucontext);
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
@@ -681,9 +682,9 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
++Thread->CurrentFrame->SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -694,16 +695,16 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
ArchHelpers::Context::SetSp(ucontext, Frame->ReturningStackLocation);
// Our ref counting doesn't matter anymore
SignalHandlerRefCounter = 0;
Thread->CurrentFrame->SignalHandlerRefCounter = 0;
// Set the new PC
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
@@ -712,24 +713,24 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
// We need to be a little bit careful here
// If we were already paused (due to GDB) and we are immediately stopping (due to gdb kill)
// Then we need to ensure we don't double decrement our idle thread counter
if (ThreadState->RunningEvents.ThreadSleeping) {
if (Thread->RunningEvents.ThreadSleeping) {
// If the thread was sleeping then its idle counter was decremented
// Reincrement it here to not break logic
++ThreadState->CTX->IdleWaitRefCount;
++Thread->CTX->IdleWaitRefCount;
}
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::Return) {
RestoreThreadState(ucontext);
RestoreThreadState(Thread, ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
--Thread->CurrentFrame->SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -748,28 +749,4 @@ uint64_t Dispatcher::GetCompileBlockPtr() {
return CompileBlockPtr.Data;
}
void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
for (auto iter = CodeBuffers.begin(); iter != CodeBuffers.end(); ++iter) {
auto [start, end] = *iter;
if (start == reinterpret_cast<uint64_t>(start_to_remove)) {
CodeBuffers.erase(iter);
return;
}
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher && IsAddressInDispatcher(Address)) {
return true;
}
return false;
}
}
@@ -1,8 +1,6 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <cstdint>
@@ -21,22 +19,20 @@ struct CpuStateFrame;
struct InternalThreadState;
}
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::CPU {
struct DispatcherConfig {
bool ExecuteBlocksWithCall = false;
uintptr_t ExitFunctionLink = 0;
uintptr_t ExitFunctionLinkThis = 0;
bool StaticRegisterAssignment = false;
bool StaticRegisterAllocation = false;
};
class Dispatcher {
public:
virtual ~Dispatcher() = default;
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
/**
* @name Dispatch Helper functions
* @{ */
@@ -50,59 +46,66 @@ public:
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint64_t IntCallbackReturnAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
} SynchronousFaultData;
uint64_t Start{};
uint64_t End{};
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
bool HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
void RegisterCodeBuffer(uint8_t* start, size_t size) {
CodeBuffers.emplace_back(reinterpret_cast<uint64_t>(start),
reinterpret_cast<uint64_t>(start + size));
}
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
protected:
Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ThreadState {Thread} {}
virtual void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) = 0;
ArchHelpers::Context::ContextBackup* StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
// These are across all arches for now
static constexpr size_t MaxGDBPauseCheckSize = 128;
static constexpr size_t MaxInterpreterTrampolineSize = 128;
virtual size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) = 0;
virtual size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) = 0;
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
protected:
Dispatcher(FEXCore::Context::Context *ctx)
: CTX {ctx}
{}
ArchHelpers::Context::ContextBackup* StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext);
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext);
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext, uint32_t IgnoreMask) {}
virtual void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
private:
std::vector<std::tuple<uint64_t, uint64_t>> CodeBuffers; // Start, End
using AsmDispatch = void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
};
}
@@ -18,21 +18,26 @@
#include <stddef.h>
#include <stdint.h>
#include <sys/mman.h>
#include "xbyak/xbyak.h"
#include <xbyak/xbyak.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
[STATE + offsetof(FEXCore::Core::STATE_TYPE, FIELD)]
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread)
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: Dispatcher(ctx)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE,
FEXCore::Allocator::mmap(nullptr, MAX_DISPATCHER_CODE_SIZE, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0),
nullptr) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "X86 dispatcher does not support SRA");
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
DispatchPtr = getCurr<AsmDispatch>();
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
@@ -78,11 +83,10 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)], rsp);
mov(qword STATE_PTR(CpuStateFrame, ReturningStackLocation), rsp);
Label LoopTop;
Label FullLookup;
Label CallBlock;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
@@ -92,29 +96,24 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
{
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
mov(rdx, qword STATE_PTR(CPUState, rip));
// L1 Cache
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
cmp(qword[r13 + rax + offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)], rdx);
jne(FullLookup);
if (!config.ExecuteBlocksWithCall) {
jmp(qword[r13 + rax + 0]);
} else {
mov(rax, qword[r13 + rax + 0]);
jmp(CallBlock);
}
jmp(qword[r13 + rax + offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)]);
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Full lookup
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
mov(rax, rdx);
mov(rbx, VirtualMemorySize - 1);
and_(rax, rbx);
@@ -143,7 +142,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
je(NoBlock);
// Update L1
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(r13, qword STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
@@ -151,30 +150,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(qword[r13 + rcx*8 + 0], rax);
// Real block if we made it here
if (!config.ExecuteBlocksWithCall) {
jmp(rax);
} else {
L(CallBlock);
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
jmp(rax);
}
{
@@ -282,13 +258,11 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(rax, r9);
}
// {rdi, rsi, rdx}
mov(rdi, config.ExitFunctionLinkThis);
mov(rsi, STATE);
mov(rdx, rax); // rax is set at the block end
// {rdi, rsi}
mov(rdi, STATE);
mov(rsi, rax); // rax is set at the block end
mov(rax, config.ExitFunctionLink);
call(rax);
call(qword STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (SignalSafeCompile) {
// Now restore the signal mask
@@ -330,7 +304,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
CallbackPtr = getCurr<JITCallback>();
push(rbx);
push(rbp);
@@ -345,7 +319,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
add(qword STATE_PTR(CpuStateFrame, SignalHandlerRefCounter), 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
@@ -353,12 +327,12 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])]);
sub(qword STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]), 16);
mov(rbx, qword STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], rsi);
mov(qword STATE_PTR(CpuStateFrame, State.rip), rsi);
// Back to the loop top now
jmp(LoopTop);
@@ -385,18 +359,17 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
mov(rax, reinterpret_cast<uint64_t>(&SynchronousFaultData));
add(byte [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)], 1);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)], X86State::X86_TRAPNO_OF);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)], 0);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)], 0x80);
add(byte STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException), 1);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo), X86State::X86_TRAPNO_OF);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code), 0);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code), 0x80);
hlt();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
IntCallbackReturnAddress = getCurr<uint64_t>();
// using CallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
@@ -430,39 +403,91 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(Start), End-Start);
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
}
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandler = ThreadStopHandlerAddress;
Pointers.ThreadPauseHandler = ThreadPauseHandlerAddress;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline
static thread_local Xbyak::CodeGenerator emit(1, &emit); // actual emit target set with setNewBuffer
size_t X86Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
using namespace Xbyak;
using namespace Xbyak::util;
emit.setNewBuffer(CodeBuffer, MaxGDBPauseCheckSize);
Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
emit.mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then we don't need to stop
emit.cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
emit.je(RunBlock);
{
// Make sure RIP is syncronized to the context
emit.mov(rax, GuestRIP);
emit.mov(qword STATE_PTR(CpuStateFrame, State.rip), rax);
// Stop the thread
emit.mov(rax, qword STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.jmp(rax);
}
emit.L(RunBlock);
emit.ready();
return emit.getSize();
}
size_t X86Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
using namespace Xbyak;
using namespace Xbyak::util;
emit.setNewBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
Label InlineIRData;
emit.mov(rdi, STATE);
emit.lea(rsi, ptr[rip + InlineIRData]);
emit.call(qword STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.jmp(qword STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.L(InlineIRData);
emit.ready();
return emit.getSize();
}
X86Dispatcher::~X86Dispatcher() {
FEXCore::Allocator::munmap(top_, MAX_DISPATCHER_CODE_SIZE);
}
#ifdef _M_X86_64
void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddress;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddress;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Dispatcher = std::make_unique<X86Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
(uintptr_t&)Interpreter.CallbackReturn = IntCallbackReturnAddress;
}
}
#endif
std::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<X86Dispatcher>(CTX, Config);
}
}
@@ -17,7 +17,10 @@ namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
virtual ~X86Dispatcher() override;
};
+33 -2
View File
@@ -19,6 +19,7 @@ $end_info$
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <set>
#include <sys/mman.h>
@@ -1135,7 +1136,7 @@ const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage) {
Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
@@ -1165,6 +1166,13 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
// Entry is a jump target
BlocksToDecode.emplace(PC);
uint64_t CurrentCodePage = PC & FHU::FEX_PAGE_MASK;
std::set<uint64_t> CodePages = { CurrentCodePage };
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
while (!BlocksToDecode.empty()) {
auto BlockDecodeIt = BlocksToDecode.begin();
uint64_t RIPToDecode = *BlockDecodeIt;
@@ -1181,10 +1189,33 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
while (1) {
// MAX_INST_SIZE assumes worst case
auto OpMinAddress = RIPToDecode + PCOffset;
auto OpMaxAddress = OpMinAddress + MAX_INST_SIZE;
auto OpMinPage = OpMinAddress & FHU::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FHU::FEX_PAGE_MASK;
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
}
if (OpMaxPage != CurrentCodePage) {
CurrentCodePage = OpMaxPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
}
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", PC + PCOffset, PC);
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", RIPToDecode + PCOffset, PC);
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
// Error while decoding instruction. We don't know the table or instruction size
+1 -1
View File
@@ -27,7 +27,7 @@ public:
Decoder(FEXCore::Context::Context *ctx);
~Decoder();
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
+12 -10
View File
@@ -125,7 +125,7 @@ static std::string hexstring(std::istringstream &ss, int delm) {
return ret;
}
static std::string encodeHex(unsigned char *data, size_t length) {
static std::string encodeHex(const unsigned char *data, size_t length) {
std::ostringstream ss;
for (size_t i=0; i < length; i++) {
@@ -251,15 +251,15 @@ void GdbServer::SendACK(std::ostream &stream, bool NACK) {
}
struct FEX_PACKED GDBContextDefinition {
uint64_t gregs[16];
uint64_t gregs[Core::CPUState::NUM_GPRS];
uint64_t rip;
uint32_t eflags;
uint32_t cs, ss, ds, es, fs, gs;
X80SoftFloat mm[8];
X80SoftFloat mm[Core::CPUState::NUM_MMS];
uint32_t fctrl;
uint32_t fstat;
uint32_t dummies[6];
uint64_t xmm[16][2];
uint64_t xmm[Core::CPUState::NUM_XMMS][2];
uint32_t mxcsr;
};
@@ -288,12 +288,12 @@ std::string GdbServer::readRegs() {
memcpy(&GDB.gregs[0], &state.gregs[0], sizeof(GDB.gregs));
memcpy(&GDB.rip, &state.rip, sizeof(GDB.rip));
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
uint64_t Flag = state.flags[i];
GDB.eflags |= (Flag << i);
}
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
memcpy(&GDB.mm[i], &state.mm[i], sizeof(GDB.mm));
}
@@ -346,7 +346,7 @@ GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
}
else if (addr == offsetof(GDBContextDefinition, eflags)) {
uint32_t eflags{};
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < Core::CPUState::NUM_EFLAG_BITS; ++i) {
uint64_t Flag = state.flags[i];
eflags |= (Flag << i);
}
@@ -382,7 +382,9 @@ GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
}
else if (addr >= offsetof(GDBContextDefinition, xmm[0][0]) &&
addr < offsetof(GDBContextDefinition, xmm[16][0])) {
return {encodeHex((unsigned char *)(&state.xmm[(addr - offsetof(GDBContextDefinition, xmm[0][0])) / 16][0]), 16), HandledPacketType::TYPE_ACK};
const auto XmmIndex = (addr - offsetof(GDBContextDefinition, xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
const auto *Data = (unsigned char *)&state.xmm[XmmIndex][0];
return {encodeHex(Data, Core::CPUState::XMM_REG_SIZE), HandledPacketType::TYPE_ACK};
}
else if (addr == offsetof(GDBContextDefinition, mxcsr)) {
uint32_t Empty{};
@@ -424,7 +426,7 @@ std::string buildTargetXML() {
// We want to just memcpy our x86 state to gdb, so we tell it the ordering.
// GPRs
for (int i=0; i < 16; i++) {
for (uint32_t i = 0; i < Core::CPUState::NUM_GPRS; i++) {
reg(FEXCore::Core::GetGRegName(i), "int64", 64);
}
@@ -481,7 +483,7 @@ std::string buildTargetXML() {
)";
// SSE regs
for (int i=0; i < 16; i++) {
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
reg("xmm" + std::to_string(i), "vec128", 128);
}
@@ -4,6 +4,7 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
@@ -25,20 +26,13 @@ static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
DEF_OP(CallbackReturn) {
Data->State->CTX->InterpreterCallbackReturn(Data->State, Data->StackEntry);
Data->State->CurrentFrame->Pointers.Interpreter.CallbackReturn(Data->State, Data->StackEntry);
}
DEF_OP(ExitFunction) {
@@ -513,6 +513,44 @@ DEF_OP(CRC32) {
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Selector = Op->Selector;
auto* Dst = GetDest<uint64_t*>(Data->SSAData, Node);
auto* Src1 = GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
auto* Src2 = GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const uint64_t TMP1 = (Selector & 0x01) == 0 ? Src1[0] : Src1[1];
const uint64_t TMP2 = (Selector & 0x10) == 0 ? Src2[0] : Src2[1];
const auto make_lo = [](uint64_t lhs, uint64_t rhs) {
uint64_t result = 0;
for (size_t i = 0; i < 64; i++) {
if ((lhs & (1ULL << i)) != 0) {
result ^= rhs << i;
}
}
return result;
};
const auto make_hi = [](uint64_t lhs, uint64_t rhs) {
uint64_t result = 0;
for (size_t i = 1; i < 64; i++) {
if ((lhs & (1ULL << i)) != 0) {
result ^= rhs >> (64 - i);
}
}
return result;
};
Dst[0] = make_lo(TMP1, TMP2);
Dst[1] = make_hi(TMP1, TMP2);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -8,6 +8,7 @@
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::CPU {
class Dispatcher;
class X86DispatchGenerator;
class Arm64DispatchGenerator;
@@ -20,7 +21,7 @@ using DestMapType = std::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(FEXCore::Context::Context *ctx,
explicit InterpreterCore(Dispatcher *Dispatch,
FEXCore::Core::InternalThreadState *Thread);
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
@@ -28,23 +29,19 @@ public:
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
bool NeedsRetainedIRCopy() const override { return true; }
void ClearCache() override;
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
std::unique_ptr<Dispatcher> Dispatcher{};
size_t BufferUsed;
Dispatcher *Dispatch;
};
template<typename T>
@@ -9,70 +9,101 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <memory>
#include <signal.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include "InterpreterOps.h"
#if defined(_M_X86_64)
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#elif defined(_M_ARM_64)
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#else
#error missing arch
#endif
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
namespace FEXCore::IR {
class IRListView;
class RegisterAllocationData;
}
namespace FEXCore::CPU {
class CPUBackend;
static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
InterpreterCore::InterpreterCore(Dispatcher *Dispatcher, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Dispatch(Dispatcher)
{
auto LocalEntry = Thread->LocalIRCache.find(Thread->CurrentFrame->State.rip);
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
Interpreter.FragmentExecuter = reinterpret_cast<uint64_t>(&InterpreterOps::InterpretIR);
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, State {Thread} {
CreateAsmDispatch(ctx, Thread);
ClearCache();
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
return reinterpret_cast<void*>(InterpreterExecution);
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
const auto IRSize = AlignUp(IR->GetInlineSize(), 16);
const auto MaxSize = IRSize + Dispatcher::MaxInterpreterTrampolineSize + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((BufferUsed + MaxSize) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState);
}
const auto BufferStart = CurrentCodeBuffer->Ptr + BufferUsed;
auto DestBuffer = BufferStart;
if (GDBEnabled) {
const auto GDBSize = Dispatch->GenerateGDBPauseCheck(DestBuffer, Entry);
DestBuffer += GDBSize;
BufferUsed += GDBSize;
}
const auto TrampolineSize = Dispatch->GenerateInterpreterTrampoline(DestBuffer);
DestBuffer += TrampolineSize;
BufferUsed += TrampolineSize;
IR->Serialize(DestBuffer);
DestBuffer += IRSize;
BufferUsed += IRSize;
return BufferStart;
}
void InterpreterCore::ClearCache() {
// Calling this one is needed to setup the initial CurrentCodeBuffer
[[maybe_unused]] auto CodeBuffer = GetEmptyCodeBuffer();
BufferUsed = 0;
}
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<InterpreterCore>(ctx, Thread);
return std::make_unique<InterpreterCore>(ctx->Dispatcher.get(), Thread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
CPUBackendFeatures GetInterpreterBackendFeatures() {
return CPUBackendFeatures { };
}
}
@@ -12,10 +12,11 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
struct DispatcherConfig;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetInterpreterBackendFeatures();
} // namespace FEXCore::CPU
@@ -112,8 +112,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -166,6 +164,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -284,6 +283,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
@@ -333,7 +333,7 @@ void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
void InterpreterOps::InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *CurrentIR) {
volatile void *StackEntry = alloca(0);
uintptr_t ListSize = CurrentIR->GetSSACount();
@@ -344,9 +344,9 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
auto BlockEnd = CurrentIR->GetBlocks().end();
InterpreterOps::IROpData OpData{};
OpData.State = Thread;
OpData.State = Frame->Thread;
OpData.SSAData = alloca(ListSize * 16);
OpData.CurrentEntry = Entry;
OpData.CurrentEntry = Frame->State.rip;
OpData.CurrentIR = CurrentIR;
OpData.StackEntry = StackEntry;
OpData.BlockIterator = CurrentIR->GetBlocks().begin();
@@ -47,14 +47,14 @@ namespace FEXCore::CPU {
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static void InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *IR);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
uint64_t CurrentEntry{};
FEXCore::IR::IRListView *CurrentIR{};
FEXCore::IR::IRListView const *CurrentIR{};
volatile void *StackEntry{};
void *SSAData{};
struct {
@@ -142,8 +142,6 @@ namespace FEXCore::CPU {
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -304,6 +302,7 @@ namespace FEXCore::CPU {
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
///< F80 ops
DEF_OP(F80LOADFCW);
@@ -408,7 +407,7 @@ namespace FEXCore::CPU {
return CompResult;
}
static uint8_t GetOpSize(FEXCore::IR::IRListView *CurrentIR, IR::OrderedNodeWrapper Node) {
static uint8_t GetOpSize(FEXCore::IR::IRListView const *CurrentIR, IR::OrderedNodeWrapper Node) {
auto IROp = CurrentIR->GetOp<FEXCore::IR::IROp_Header>(Node);
return IROp->Size;
}
@@ -4,6 +4,7 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
@@ -4,6 +4,8 @@ tags: backend|interpreter
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,131 @@
/*
$info$
tags: backend|arm64
desc: relocation logic of the arm64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.GetCode();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
.Lit = Literal(Pointer),
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
place(&Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
void Arm64JITCore::InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.GetCode();
LoadConstant(Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
Literal<uint64_t> Lit(Pointer);
place(&Lit);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
}
return true;
}
}
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Core/LookupCache.h"
@@ -19,13 +20,6 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// First we must reset the stack
@@ -33,7 +27,7 @@ DEF_OP(SignalReturn) {
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalReturnHandler)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)));
br(x0);
}
@@ -46,10 +40,10 @@ DEF_OP(CallbackReturn) {
ResetStack();
// We can now lower the ref counter again
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalHandlerRefCountPointer)));
ldr(w2, MemOperand(x0));
ldr(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
sub(w2, w2, 1);
str(w2, MemOperand(x0));
str(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
@@ -73,7 +67,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchHost{ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker};
Literal l_BranchGuest{NewRIP};
ldr(x0, &l_BranchHost);
@@ -82,10 +76,10 @@ DEF_OP(ExitFunction) {
place(&l_BranchHost);
place(&l_BranchGuest);
} else {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
RipReg = GetReg<RA_64>(Op->NewRIP.ID());
// L1 Cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -96,7 +90,7 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.DispatcherLoopTop)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop)));
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
@@ -104,9 +98,9 @@ DEF_OP(ExitFunction) {
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
const auto Target = Op->TargetBlock.ID();
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
@@ -199,8 +193,8 @@ DEF_OP(Syscall) {
str(GetReg<RA_64>(Op->Header.Args[i].ID()), MemOperand(sp, i * 8));
}
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerFunc)));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
blr(x3);
@@ -383,7 +377,7 @@ DEF_OP(Thunk) {
PushDynamicRegsAndLR();
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x0, GetReg<RA_64>(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
@@ -452,7 +446,7 @@ DEF_OP(RemoveThreadCodeEntry) {
mov(x0, STATE);
LoadConstant(x1, Entry);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.RemoveThreadCodeEntryFromJIT)));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -469,10 +463,10 @@ DEF_OP(CPUID) {
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Function.ID()));
mov(x2, GetReg<RA_64>(Op->Leaf.ID()));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
@@ -489,8 +483,6 @@ DEF_OP(CPUID) {
#undef DEF_OP
void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -13,22 +13,22 @@ using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
mov(GetDst(Node), GetSrc(Op->DestVector.ID()));
switch (Op->Header.ElementSize) {
case 1: {
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 2: {
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 4: {
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 8: {
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Src.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -39,18 +39,18 @@ DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
switch (Op->Header.ElementSize) {
case 1:
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
uxtb(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 2:
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
uxth(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 4:
fmov(GetDst(Node).S(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
fmov(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()).W());
break;
case 8:
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Header.Args[0].ID()).X());
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()).X());
break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
}
@@ -58,22 +58,22 @@ DEF_OP(VCastFromGPR) {
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()));
break;
}
case 0x0408: { // Float <- int64_t
scvtf(GetDst(Node).S(), GetReg<RA_64>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).S(), GetReg<RA_64>(Op->Src.ID()));
break;
}
case 0x0804: { // Double <- int32_t
scvtf(GetDst(Node).D(), GetReg<RA_32>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).D(), GetReg<RA_32>(Op->Src.ID()));
break;
}
case 0x0808: { // Double <- int64_t
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Header.Args[0].ID()));
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()));
break;
}
}
@@ -81,14 +81,14 @@ DEF_OP(Float_FromGPR_S) {
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // Double <- Float
fcvt(GetDst(Node).D(), GetSrc(Op->Header.Args[0].ID()).S());
fcvt(GetDst(Node).D(), GetSrc(Op->Scalar.ID()).S());
break;
}
case 0x0408: { // Float <- Double
fcvt(GetDst(Node).S(), GetSrc(Op->Header.Args[0].ID()).D());
fcvt(GetDst(Node).S(), GetSrc(Op->Scalar.ID()).D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
@@ -99,10 +99,10 @@ DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
switch (Op->Header.ElementSize) {
case 4:
scvtf(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
scvtf(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
scvtf(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
scvtf(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", Op->Header.ElementSize);
}
@@ -112,10 +112,10 @@ DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
switch (Op->Header.ElementSize) {
case 4:
fcvtzs(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
fcvtzs(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", Op->Header.ElementSize);
}
@@ -125,11 +125,11 @@ DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
fcvtzs(GetDst(Node).V4S(), GetDst(Node).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetDst(Node).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", Op->Header.ElementSize);
@@ -142,11 +142,11 @@ DEF_OP(Vector_FToF) {
switch (Conv) {
case 0x0804: { // Double <- Float
fcvtl(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2S());
fcvtl(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2S());
break;
}
case 0x0408: { // Float <- Double
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Vector.ID()).V2D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
@@ -159,50 +159,50 @@ DEF_OP(Vector_FToI) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
case 4:
frintn(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintn(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintn(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintn(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintm(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintm(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintm(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintm(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintp(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintp(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintp(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintp(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
case 4:
frintz(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frintz(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frintz(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frintz(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
}
break;
@@ -14,41 +14,41 @@ using namespace vixl::aarch64;
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
aesimc(GetDst(Node).V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
aesimc(GetDst(Node).V16B(), GetSrc(Op->Vector.ID()).V16B());
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
aesmc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
aesimc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
}
DEF_OP(AESKeyGenAssist) {
@@ -59,7 +59,7 @@ DEF_OP(AESKeyGenAssist) {
// Do a "regular" AESE step
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
mov(VTMP1.V16B(), GetSrc(Op->Src.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
// Do a table shuffle to undo ShiftRows
@@ -102,16 +102,45 @@ DEF_OP(CRC32) {
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
auto Dst = GetDst(Node).Q();
auto Src1 = GetSrc(Op->Src1.ID()).V2D();
auto Src2 = GetSrc(Op->Src2.ID()).V2D();
switch (Op->Selector) {
case 0b00000000:
pmull(Dst, Src1, Src2);
break;
case 0b00000001:
mov(VTMP1.V1D(), Src1, 1);
pmull(Dst, VTMP1.V2D(), Src2);
break;
case 0b00010000:
mov(VTMP1.V1D(), Src2, 1);
pmull(Dst, VTMP1.V2D(), Src1);
break;
case 0b00010001:
pmull2(Dst, Src1, Src2);
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
#undef REGISTER_OP
}
}
@@ -13,7 +13,7 @@ using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
+133 -212
View File
@@ -35,6 +35,10 @@ $end_info$
#include <unistd.h>
#include <string.h>
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
namespace {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
@@ -88,7 +92,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -103,7 +107,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -122,7 +126,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -147,7 +151,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
else {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -168,7 +172,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -187,7 +191,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -204,7 +208,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -223,7 +227,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
mov(v1.D(), GetSrc(IROp->Args[1].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -243,7 +247,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -261,7 +265,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -279,7 +283,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -300,7 +304,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -318,7 +322,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -341,7 +345,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -364,35 +368,64 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
//fmt::print("ExitFunctionLink: Aborting, {:X} not in cache\n", GuestRip);
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
auto offset = HostCode/4 - branch/4;
if (IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.b(offset);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
emit.ldr(x0, &l_BranchHost);
emit.blr(x0);
emit.place(&l_BranchHost);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
});
} else {
// fallback case - do a soft-er link by patching the pointer
record[0] = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
}
return HostCode;
}
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
Arm64JITCore::CodeBuffer Arm64JITCore::AllocateNewCodeBuffer(size_t Size) {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(
FEXCore::Allocator::mmap(nullptr,
Buffer.Size,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS,
-1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
Dispatcher->RegisterCodeBuffer(Buffer.Ptr, Buffer.Size);
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
Dispatcher->RemoveCodeBuffer(Buffer.Ptr);
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx, 0)
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
#if DEBUG
@@ -432,84 +465,57 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RegisterVectorHandlers();
RegisterEncryptionHandlers();
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
// Common
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore_ExitFunctionLink);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
// Platform Specific
auto &AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
AArch64.LUREM = reinterpret_cast<uint64_t>(LUREM);
AArch64.LREM = reinterpret_cast<uint64_t>(LREM);
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
// Must be done after Dispatcher init
SetAllowAssembler(true);
EmitDetectionString();
CurrentCodeBuffer = &InitialCodeBuffer;
ClearCache();
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
if (!Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Thread->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
void Arm64JITCore::EmitDetectionString() {
@@ -521,54 +527,14 @@ void Arm64JITCore::EmitDetectionString() {
void Arm64JITCore::ClearCache() {
// Get the backing code buffer
auto Buffer = GetBuffer();
if (Dispatcher->SignalHandlerRefCounter == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
// Set the current code buffer to the initial
*Buffer = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
CurrentCodeBuffer = &InitialCodeBuffer;
}
if (CurrentCodeBuffer->Size == MAX_CODE_SIZE) {
// Rewind to the start of the code cache start
Buffer->Reset();
}
else {
FreeCodeBuffer(InitialCodeBuffer);
// Resize the code buffer and reallocate our code size
InitialCodeBuffer.Size *= 1.5;
InitialCodeBuffer.Size = std::min(InitialCodeBuffer.Size, MAX_CODE_SIZE);
InitialCodeBuffer = AllocateNewCodeBuffer(InitialCodeBuffer.Size);
*Buffer = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
}
else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = Arm64JITCore::AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
*Buffer = vixl::CodeBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
auto CodeBuffer = GetEmptyCodeBuffer();
*GetBuffer() = vixl::CodeBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
FreeCodeBuffer(InitialCodeBuffer);
}
IR::PhysicalRegister Arm64JITCore::GetPhys(IR::NodeID Node) const {
@@ -699,13 +665,14 @@ bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
using namespace aarch64;
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
#ifndef NDEBUG
LoadConstant(x0, Entry);
@@ -714,9 +681,9 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
this->IR = IR;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((GetCursorOffset() + BufferRange) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState, false);
CTX->ClearCodeCache(ThreadState);
}
// AAPCS64
@@ -739,31 +706,11 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
// X1-X3 = Temp
// X4-r18 = RA
GuestEntry = GetCursorAddress<uint64_t>();
GuestEntry = GetCursorAddress<uint8_t *>();
if (CTX->GetGdbServerStatus()) {
aarch64::Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Thread))); // Get thread
ldr(x0, MemOperand(x0, offsetof(FEXCore::Core::InternalThreadState, CTX))); // Get Context
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then we don't need to stop
cbz(w0, &RunBlock);
{
// Make sure RIP is syncronized to the context
LoadConstant(x0, Entry);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// Stop the thread
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(x0);
}
bind(&RunBlock);
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
GetBuffer()->CursorForward(GDBSize);
}
//LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
@@ -788,7 +735,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uintptr_t BlockStartHostCode = GetCursorAddress<uintptr_t>();
auto BlockStartHostCode = GetCursorAddress<uint8_t *>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -812,7 +759,10 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.push_back({BlockStartHostCode, static_cast<uint32_t>(GetCursorAddress<uintptr_t>() - BlockStartHostCode)});
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t *>() - BlockStartHostCode)
});
}
}
@@ -825,66 +775,30 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
FinalizeCode();
auto CodeEnd = GetCursorAddress<uint64_t>();
CPU.EnsureIAndDCacheCoherency(reinterpret_cast<void*>(GuestEntry), CodeEnd - reinterpret_cast<uint64_t>(GuestEntry));
auto CodeEnd = GetCursorAddress<uint8_t *>();
CPU.EnsureIAndDCacheCoherency(GuestEntry, CodeEnd - GuestEntry);
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(CodeEnd) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->HostCodeSize = CodeEnd - GuestEntry;
DebugData->Relocations = &Relocations;
}
this->IR = nullptr;
return reinterpret_cast<void*>(GuestEntry);
return GuestEntry;
}
uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
void Arm64JITCore::ResetStack() {
if (SpillSlots == 0)
return;
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
//fmt::print("ExitFunctionLink: Aborting, {:X} not in cache\n", GuestRip);
Frame->State.rip = GuestRip;
return core->Dispatcher->AbsoluteLoopTopAddress;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = core->Dispatcher->ExitFunctionLinkerAddress;
auto offset = HostCode/4 - branch/4;
if (IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.b(offset);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
emit.ldr(x0, &l_BranchHost);
emit.blr(x0);
emit.place(&l_BranchHost);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
});
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// fallback case - do a soft-er link by patching the pointer
record[0] = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
return HostCode;
}
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
@@ -894,4 +808,11 @@ std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, F
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
}
CPUBackendFeatures GetArm64JITBackendFeatures() {
return CPUBackendFeatures {
.SupportsStaticRegisterAllocation = true
};
}
}
+76 -48
View File
@@ -6,17 +6,23 @@ $end_info$
#pragma once
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <array>
#include <cstdint>
#include <map>
#include <utility>
#include <vector>
#define STATE x28
#define TMP1 x0
#define TMP2 x1
@@ -37,11 +43,6 @@ using namespace vixl::aarch64;
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread);
~Arm64JITCore() override;
@@ -51,7 +52,7 @@ public:
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -59,13 +60,6 @@ public:
void ClearCache() override;
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
@@ -73,10 +67,8 @@ public:
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
@@ -147,38 +139,77 @@ private:
vixl::aarch64::Decoder Decoder;
#endif
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
void FreeCodeBuffer(CodeBuffer Buffer);
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
CodeBuffer InitialCodeBuffer{};
// This is the array of /additional/ code buffers that we may need to allocate
// Allocation only occurs when we've hit signals and need to clear code cache
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
// We don't want to mvoe above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
#if DEBUG
vixl::aarch64::Disassembler Disasm;
#endif
static uint64_t ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
FEXCore::Core::DebugData *DebugData;
void ResetStack();
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Literal<uint64_t> Lit;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
@@ -269,8 +300,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -320,7 +349,7 @@ private:
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
DEF_OP(GuestOpcode);
DEF_OP(Fence);
DEF_OP(Break);
DEF_OP(Phi);
@@ -437,9 +466,8 @@ private:
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
#undef DEF_OP
};
}
} // namespace FEXCore::CPU
@@ -4,6 +4,8 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include <FEXCore/Utils/CompilerDefs.h>
@@ -58,26 +60,26 @@ DEF_OP(LoadContext) {
DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1:
strb(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
strb(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 2:
strh(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
strh(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 4:
str(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
str(GetReg<RA_32>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
case 8:
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
str(GetReg<RA_64>(Op->Value.ID()), MemOperand(STATE, Op->Offset));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
}
}
else {
auto Src = GetSrc(Op->Header.Args[0].ID());
auto Src = GetSrc(Op->Value.ID());
switch (OpSize) {
case 1:
str(Src.B(), MemOperand(STATE, Op->Offset));
@@ -104,7 +106,7 @@ DEF_OP(LoadRegister) {
auto Op = IROp->C<IR::IROp_LoadRegister>();
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0])) / 8;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
@@ -135,7 +137,7 @@ DEF_OP(LoadRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "out of range regId");
@@ -189,7 +191,7 @@ DEF_OP(StoreRegister) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
if (Op->Class == IR::GPRClass) {
auto regId = Op->Offset / 8 - 1;
auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
@@ -219,7 +221,7 @@ DEF_OP(StoreRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "regId out of range");
@@ -261,8 +263,8 @@ DEF_OP(StoreRegister) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[0].ID());
const size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Index.ID());
if (Op->Class == FEXCore::IR::GPRClass) {
switch (Op->Stride) {
@@ -349,11 +351,11 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[1].ID());
const size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Index.ID());
if (Op->Class == FEXCore::IR::GPRClass) {
auto value = GetReg<RA_64>(Op->Header.Args[0].ID());
auto value = GetReg<RA_64>(Op->Value.ID());
switch (Op->Stride) {
case 1:
@@ -392,7 +394,7 @@ DEF_OP(StoreContextIndexed) {
}
}
else {
auto value = GetSrc(Op->Header.Args[0].ID());
auto value = GetSrc(Op->Value.ID());
switch (Op->Stride) {
case 1:
@@ -441,25 +443,25 @@ DEF_OP(StoreContextIndexed) {
DEF_OP(SpillRegister) {
auto Op = IROp->C<IR::IROp_SpillRegister>();
uint8_t OpSize = IROp->Size;
uint32_t SlotOffset = Op->Slot * 16;
const uint8_t OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * 16;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
strb(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 2: {
strh(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
strh(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 4: {
str(GetReg<RA_32>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetReg<RA_32>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
case 8: {
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetReg<RA_64>(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -467,15 +469,15 @@ DEF_OP(SpillRegister) {
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
case 4: {
str(GetSrc(Op->Header.Args[0].ID()).S(), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()).S(), MemOperand(sp, SlotOffset));
break;
}
case 8: {
str(GetSrc(Op->Header.Args[0].ID()).D(), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()).D(), MemOperand(sp, SlotOffset));
break;
}
case 16: {
str(GetSrc(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
str(GetSrc(Op->Value.ID()), MemOperand(sp, SlotOffset));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -539,7 +541,7 @@ DEF_OP(LoadFlag) {
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
strb(GetReg<RA_64>(Op->Value.ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
}
MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
@@ -570,7 +572,7 @@ MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Registe
DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -992,7 +994,7 @@ DEF_OP(VStoreMemElement) {
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
@@ -1007,7 +1009,7 @@ DEF_OP(CacheLineClear) {
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use this instruction directly
+20 -11
View File
@@ -4,13 +4,21 @@ tags: backend|arm64
$end_info$
*/
#include <syscall.h>
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - GuestEntry});
}
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -36,7 +44,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.OverflowExceptionHandler)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)));
br(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
@@ -46,13 +54,13 @@ DEF_OP(Break) {
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadStopHandlerSpillSRA)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)));
br(TMP1);
break;
}
@@ -60,7 +68,7 @@ DEF_OP(Break) {
{
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.UnimplementedInstructionHandler)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)));
br(TMP1);
break;
@@ -97,7 +105,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Src = GetReg<RA_64>(Op->RoundMode.ID());
// Setup the rounding flags correctly
and_(TMP1, Src, 0b11);
@@ -132,15 +140,15 @@ DEF_OP(Print) {
PushDynamicRegsAndLR();
if (IsGPR(Op->Header.Args[0].ID())) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintValue)));
if (IsGPR(Op->Value.ID())) {
mov(x0, GetReg<RA_64>(Op->Value.ID()));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)));
}
else {
fmov(x0, GetSrc(Op->Header.Args[0].ID()).V1D());
fmov(x0, GetSrc(Op->Value.ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Header.Args[0].ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintVectorValue)));
fmov(x1, GetSrc(Op->Value.ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)));
}
SpillStaticRegs();
blr(x3);
@@ -231,6 +239,7 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, GuestOpcode);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -15,13 +15,13 @@ DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
case 4: {
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_32>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_32>(Node), Regs[Op->Element]);
break;
}
case 8: {
auto Src = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_64>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_64>(Node), Regs[Op->Element]);
break;
@@ -40,15 +40,15 @@ DEF_OP(CreateElementPair) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_32>(Op->Header.Args[1].ID());
RegFirst = GetReg<RA_32>(Op->Lower.ID());
RegSecond = GetReg<RA_32>(Op->Upper.ID());
RegTmp = w0;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetReg<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_64>(Op->Header.Args[1].ID());
RegFirst = GetReg<RA_64>(Op->Lower.ID());
RegSecond = GetReg<RA_64>(Op->Upper.ID());
RegTmp = x0;
break;
}
@@ -70,7 +70,7 @@ DEF_OP(CreateElementPair) {
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
mov(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()));
}
#undef DEF_OP
File diff suppressed because it is too large. Load diff
+2
View File
@@ -16,9 +16,11 @@ class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetX86JITBackendFeatures();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -29,13 +29,6 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
@@ -43,7 +36,7 @@ DEF_OP(SignalReturn) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
}
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalReturnHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)]);
}
DEF_OP(CallbackReturn) {
@@ -53,7 +46,7 @@ DEF_OP(CallbackReturn) {
}
// Make sure to adjust the refcounter so we don't clear the cache now
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
@@ -91,14 +84,15 @@ DEF_OP(ExitFunction) {
jmp(qword[rax]);
L(l_BranchHost);
dq(Dispatcher->ExitFunctionLinkerAddress);
//FEX_TODO(this is not per thread)
dq(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
L(l_BranchGuest);
dq(NewRIP);
} else {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer)]);
mov(rax, RipReg);
@@ -113,7 +107,7 @@ DEF_OP(ExitFunction) {
L(FullLookup);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.DispatcherLoopTop)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop)]);
}
#ifdef BLOCKSTATS
@@ -181,13 +175,13 @@ DEF_OP(Syscall) {
}
mov(rsi, STATE); // Move thread in to rsi
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerObj)]);
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj)]);
mov(rdx, rsp);
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerFunc)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -272,7 +266,7 @@ DEF_OP(RemoveThreadCodeEntry) {
mov(rax, Entry); // imm64 move
mov(rsi, rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.RemoveThreadCodeEntryFromJIT)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -296,7 +290,7 @@ DEF_OP(CPUID) {
// rsi can be in the source registers, so copy argument to edx first
mov (edx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov (esi, GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDObj)]);
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)]);
auto NumPush = RA64.size();
@@ -304,7 +298,7 @@ DEF_OP(CPUID) {
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDFunction)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -320,8 +314,6 @@ DEF_OP(CPUID) {
#undef DEF_OP
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -75,16 +75,37 @@ DEF_OP(CRC32) {
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
auto Dst = GetDst(Node);
auto Src1 = GetSrc(Op->Src1.ID());
auto Src2 = GetSrc(Op->Src2.ID());
switch (Op->Selector) {
case 0b00000000:
case 0b00000001:
case 0b00010000:
case 0b00010001:
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
#undef DEF_OP
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
#undef REGISTER_OP
}
}
+82 -170
View File
@@ -43,6 +43,9 @@ $end_info$
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
namespace {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
@@ -55,26 +58,6 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
namespace FEXCore::CPU {
CodeBuffer AllocateNewCodeBuffer(FEXCore::Context::Context *CTX, size_t Size) {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(
FEXCore::Allocator::mmap(nullptr,
Buffer.Size,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS,
-1, 0));
LOGMAN_THROW_A_FMT(Buffer.Ptr != reinterpret_cast<uint8_t*>(~0ULL), "Couldn't allocate code buffer");
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
void X86JITCore::PushRegs() {
sub(rsp, 16 * RAXMM_x.size());
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
@@ -115,7 +98,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_VOID_U16: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
break;
@@ -124,7 +107,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movss(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -138,7 +121,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -153,7 +136,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -169,7 +152,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -183,7 +166,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -196,7 +179,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -210,7 +193,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
movsd(xmm1, GetSrc(IROp->Args[1].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -224,7 +207,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -237,7 +220,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -250,7 +233,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -266,7 +249,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -279,7 +262,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -297,7 +280,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -317,17 +300,34 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
}
static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->CurrentFrame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
}
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
record[0] = HostCode;
return HostCode;
}
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer)
: CodeGenerator(Buffer.Size, Buffer.Ptr, nullptr)
, CTX {ctx}
, ThreadState {Thread}
, InitialCodeBuffer {Buffer}
{
CurrentCodeBuffer = &InitialCodeBuffer;
EmitDetectionString();
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, CodeGenerator(0, this, nullptr) // this is not used here
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -356,69 +356,39 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
RegisterVectorHandlers();
RegisterEncryptionHandlers();
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&X86JITCore_ExitFunctionLink);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
}
// Must be done after Dispatcher init
ClearCache();
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
X86JITCore::~X86JITCore() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
FreeCodeBuffer(InitialCodeBuffer);
}
void X86JITCore::EmitDetectionString() {
@@ -429,44 +399,8 @@ void X86JITCore::EmitDetectionString() {
}
void X86JITCore::ClearCache() {
if (Dispatcher->SignalHandlerRefCounter == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
// Set the current code buffer to the initial
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
CurrentCodeBuffer = &InitialCodeBuffer;
}
if (CurrentCodeBuffer->Size == MAX_CODE_SIZE) {
// Rewind to the start of the code cache start
reset();
}
else {
FreeCodeBuffer(InitialCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MAX_CODE_SIZE);
InitialCodeBuffer = AllocateNewCodeBuffer(CTX, CurrentCodeBuffer->Size);
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
}
else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(CTX, X86JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
auto CodeBuffer = GetEmptyCodeBuffer();
setNewBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
}
@@ -635,40 +569,27 @@ std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::G
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((getSize() + BufferRange) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState, false);
CTX->ClearCodeCache(ThreadState);
}
void *GuestEntry = getCurr<void*>();
GuestEntry = getCurr<uint8_t*>();
CursorEntry = getSize();
this->IR = IR;
if (CTX->GetGdbServerStatus()) {
Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(RunBlock);
// Else we need to pause now
mov(rax, Dispatcher->ThreadPauseHandlerAddress);
jmp(rax);
ud2();
L(RunBlock);
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
setSize(getSize() + GDBSize);
}
LOGMAN_THROW_A_FMT(RAData != nullptr, "Needs RA");
@@ -726,12 +647,13 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
{
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = getCurr<uint8_t *>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -786,6 +708,13 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(getCurr<uint8_t *>() - BlockStartHostCode)
});
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
@@ -808,29 +737,12 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
return GuestEntry;
}
uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->CurrentFrame->State.rip = GuestRip;
return core->Dispatcher->AbsoluteLoopTopAddress;
}
auto LinkerAddress = core->Dispatcher->ExitFunctionLinkerAddress;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
record[0] = HostCode;
return HostCode;
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread);
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, X86JITCore::INITIAL_CODE_SIZE));
CPUBackendFeatures GetX86JITBackendFeatures() {
return CPUBackendFeatures { };
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
+12 -39
View File
@@ -6,6 +6,7 @@ $end_info$
#pragma once
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
@@ -25,13 +26,6 @@ using namespace Xbyak;
#include <tuple>
namespace FEXCore::CPU {
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void FreeCodeBuffer(CodeBuffer Buffer);
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
@@ -58,8 +52,7 @@ const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
CodeBuffer Buffer);
FEXCore::Core::InternalThreadState *Thread);
~X86JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
@@ -67,7 +60,7 @@ public:
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -75,13 +68,6 @@ public:
void ClearCache() override;
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
@@ -150,9 +136,7 @@ private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
uint64_t Entry;
std::unordered_map<IR::NodeID, Label> JumpTargets;
@@ -205,33 +189,23 @@ private:
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
FEXCore::Core::DebugData *DebugData;
#ifdef BLOCKSTATS
bool GetSamplingData {true};
#endif
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
CodeBuffer InitialCodeBuffer{};
// This is the array of /additional/ code buffers that we may need to allocate
// Allocation only occurs when we've hit signals and need to clear code cache
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (X86JITCore::*)(const Label& label, LabelType type);
@@ -330,8 +304,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -375,7 +347,7 @@ private:
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
DEF_OP(GuestOpcode);
DEF_OP(Fence);
DEF_OP(Break);
DEF_OP(Phi);
@@ -491,6 +463,7 @@ private:
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
DEF_OP(PCLMUL);
#undef DEF_OP
};
@@ -4,6 +4,7 @@ tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include <FEXCore/Core/CoreState.h>
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
@@ -20,6 +21,12 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - GuestEntry});
}
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -45,7 +52,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.OverflowExceptionHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)]);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
@@ -53,7 +60,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
break;
}
case FEXCore::IR::Break_Interrupt3: // INT3
@@ -65,7 +72,7 @@ DEF_OP(Break) {
}
// This jump target needs to be a constant offset here
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadPauseHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)]);
}
else {
// If we don't have a gdb server attached then....crash?
@@ -73,7 +80,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
}
break;
}
@@ -84,7 +91,7 @@ DEF_OP(Break) {
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.UnimplementedInstructionHandler)]);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
@@ -130,13 +137,13 @@ DEF_OP(Print) {
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintValue)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)]);
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintVectorValue)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)]);
}
PopRegs();
@@ -178,6 +185,7 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, GuestOpcode);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
@@ -4,6 +4,7 @@ tags: backend|x86-64
desc: relocation logic of the x86-64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -11,7 +12,7 @@ namespace FEXCore::CPU {
uint64_t X86JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return Dispatcher->ExitFunctionLinkerAddress;
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
@@ -52,14 +52,6 @@ LookupCache::~LookupCache() {
FEXCore::Allocator::munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
}
void LookupCache::HintUsedRange(uint64_t Address, uint64_t Size) {
// Tell the kernel we will definitely need [Address, Address+Size) mapped for the page pointer
// Page Pointer is allocated per page, so shift by page size
Address >>= 12;
Size >>= 12;
madvise(reinterpret_cast<void*>(PagePointer + Address), Size, MADV_WILLNEED);
}
void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear out the page memory
+17 -18
View File
@@ -25,9 +25,6 @@ public:
LookupCache(FEXCore::Context::Context *CTX);
~LookupCache();
using LookupCacheIter = uintptr_t;
uintptr_t End() { return 0; }
uintptr_t FindBlock(uint64_t Address) {
// Try L1, no lock needed
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
@@ -72,30 +69,34 @@ public:
std::map<uint64_t, std::vector<uint64_t>> CodePages;
// Appends Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto InsertPoint =
#endif
BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A_FMT(InsertPoint.second == true, "Dupplicate block mapping added");
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
rv |= CodePages[CurrentPage].size() == 0;
CodePages[CurrentPage].push_back(Address);
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length -1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto &CodePage = CodePages[CurrentPage];
rv |= CodePage.size() == 0;
CodePage.push_back(Address);
}
return rv;
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void *HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_A_FMT(Inserted, "Duplicate block mapping added");
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
return rv;
}
void Erase(uint64_t Address) {
@@ -149,8 +150,6 @@ public:
void ClearCache();
void ClearL2Cache();
void HintUsedRange(uint64_t Address, uint64_t Size);
uintptr_t GetL1Pointer() const { return L1Pointer; }
uintptr_t GetPagePointer() const { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
@@ -158,7 +157,7 @@ public:
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::LocalIRCache,
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::DebugStore,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
@@ -6566,6 +6566,7 @@ constexpr uint16_t PF_F2 = 3;
{OPD(0, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<4>},
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(0, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
};
@@ -6649,6 +6650,8 @@ constexpr uint16_t PF_F2 = 3;
{OPD(2, 0b10, 0xF7), 1, &OpDispatchBuilder::BMI2Shift},
{OPD(2, 0b11, 0xF7), 1, &OpDispatchBuilder::BMI2Shift},
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::VPCLMULQDQOp},
{OPD(3, 0b11, 0xF0), 1, &OpDispatchBuilder::RORX},
};
#undef OPD
+7 -3
View File
@@ -76,8 +76,10 @@ public:
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
// Used during new op bringup
bool ShouldDump {false};
struct JumpTargetInfo {
OrderedNode* BlockEntry;
bool HaveEmitted;
@@ -621,6 +623,8 @@ public:
void DPPOp(OpcodeArgs);
void MPSADBWOp(OpcodeArgs);
void PCLMULQDQOp(OpcodeArgs);
void VPCLMULQDQOp(OpcodeArgs);
void CRC32(OpcodeArgs);
@@ -1236,14 +1240,14 @@ private:
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
if (CTX->IsTSOEnabled())
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
if (CTX->IsTSOEnabled())
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
@@ -306,4 +306,26 @@ void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Selector needs to be literal here");
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
auto Res = _PCLMUL(Dest, Src, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[2].IsLiteral(), "Selector needs to be literal here");
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
auto Res = _PCLMUL(Src1, Src2, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
}
@@ -42,7 +42,7 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -441,7 +441,7 @@ void InitializeVEXTables() {
{OPD(3, 0b01, 0x40), 1, X86InstInfo{"VDPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x42), 1, X86InstInfo{"VMPSADBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x48), 1, X86InstInfo{"VPERMILzz2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
+127
View File
@@ -0,0 +1,127 @@
#include "GDBJIT.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/Utils/LogManager.h>
#if defined(GDB_SYMBOLS_ENABLED)
#include <FEXCore/Debug/GDBReaderInterface.h>
extern "C" {
enum jit_actions_t { JIT_NOACTION = 0, JIT_REGISTER_FN, JIT_UNREGISTER_FN };
struct jit_code_entry {
jit_code_entry *next_entry;
jit_code_entry *prev_entry;
const char *symfile_addr;
uint64_t symfile_size;
};
struct jit_descriptor {
uint32_t version;
/* This type should be jit_actions_t, but we use uint32_t
to be explicit about the bitwidth. */
uint32_t action_flag;
jit_code_entry *relevant_entry;
jit_code_entry *first_entry;
};
/* Make sure to specify the version statically, because the
debugger may check the version before we can set it. */
constinit jit_descriptor __jit_debug_descriptor = {.version = 1};
/* GDB puts a breakpoint in this function. */
void __attribute__((noinline)) __jit_debug_register_code() {
asm volatile("" ::"r"(&__jit_debug_descriptor));
};
}
namespace FEXCore {
void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart,
uint64_t GuestRIP, uintptr_t HostEntry,
FEXCore::Core::DebugData *DebugData) {
auto map = Entry->SourcecodeMap.get();
if (map) {
auto FileOffset = GuestRIP - VAFileStart;
auto Sym = map->FindSymbolMapping(FileOffset);
std::string SymName = HLE::SourcecodeSymbolMapping::SymName(
Sym, Entry->Filename, HostEntry, FileOffset);
std::vector<gdb_line_mapping> Lines;
for (const auto &GuestOpcode : DebugData->GuestOpcodes) {
auto Line = map->FindLineMapping(GuestRIP + GuestOpcode.GuestEntryOffset -
VAFileStart);
if (Line) {
Lines.push_back(
{Line->LineNumber, HostEntry + GuestOpcode.HostEntryOffset});
}
}
size_t size = sizeof(info_t) + 1 * sizeof(blocks_t) +
Lines.size() * sizeof(gdb_line_mapping);
auto mem = (uint8_t *)malloc(size);
auto base = mem;
info_t *info = (info_t *)mem;
mem += sizeof(info_t);
strncpy(info->filename, map->SourceFile.c_str(), 511);
info->nblocks = 1;
auto blocks = (blocks_t *)mem;
info->blocks_ofs = mem - base;
mem += info->nblocks * sizeof(blocks_t);
for (int i = 0; i < info->nblocks; i++) {
strncpy(blocks[i].name, SymName.c_str(), 511);
blocks[i].start = HostEntry;
blocks[i].end = HostEntry + DebugData->HostCodeSize;
}
info->nlines = Lines.size();
auto lines = (gdb_line_mapping *)mem;
info->lines_ofs = mem - base;
mem += info->nlines * sizeof(gdb_line_mapping);
if (info->nlines) {
memcpy(lines, &Lines.at(0), info->nlines * sizeof(gdb_line_mapping));
}
auto entry = new jit_code_entry{0, 0, 0, 0};
entry->symfile_addr = (const char *)info;
entry->symfile_size = size;
if (__jit_debug_descriptor.first_entry) {
__jit_debug_descriptor.relevant_entry->next_entry = entry;
entry->prev_entry = __jit_debug_descriptor.relevant_entry;
} else {
__jit_debug_descriptor.first_entry = entry;
}
__jit_debug_descriptor.relevant_entry = entry;
__jit_debug_descriptor.action_flag = JIT_REGISTER_FN;
__jit_debug_register_code();
}
}
} // namespace FEXCore
#else
namespace FEXCore {
void GDBJITRegister([[maybe_unused]] FEXCore::IR::AOTIRCacheEntry *Entry,
[[maybe_unused]] uintptr_t VAFileStart,
[[maybe_unused]] uint64_t GuestRIP,
[[maybe_unused]] uintptr_t HostEntry,
[[maybe_unused]] FEXCore::Core::DebugData *DebugData) {
ERROR_AND_DIE_FMT("GDBSymbols support not compiled in");
}
} // namespace FEXCore
#endif
+7
View File
@@ -0,0 +1,7 @@
#include <Interface/IR/AOTIR.h>
namespace FEXCore {
void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry, FEXCore::Core::DebugData *DebugData);
}
+258 -14
View File
@@ -10,14 +10,18 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include "FEXCore/Utils/CompilerDefs.h"
#include "Thunks.h"
#include <cstdint>
#include <dlfcn.h>
#include <Interface/Context/Context.h>
#include "FEXCore/Core/X86Enums.h"
#include <malloc.h>
#include <map>
#include <mutex>
#include <unordered_map>
#include <memory>
#include <shared_mutex>
#include <stdint.h>
@@ -26,26 +30,106 @@ $end_info$
struct LoadlibArgs {
const char *Name;
uintptr_t CallbackThunks;
};
static thread_local FEXCore::Core::InternalThreadState *Thread;
static __attribute__((aligned(16), naked, section("HostToGuestTrampolineTemplate"))) void HostToGuestTrampolineTemplate() {
#if defined(_M_X86_64)
asm(
"lea 0f(%rip), %r11 \n"
"jmpq *0f(%rip) \n"
".align 8 \n"
"0: \n"
".quad 0, 0, 0, 0 \n" // TrampolineInstanceInfo
);
#elif defined(_M_ARM_64)
asm(
"adr x11, 0f \n"
"ldr x16, [x11] \n"
"br x16 \n"
// Manually align to the next 8-byte boundary
// NOTE: GCC over-aligns to a full page when using .align directives on ARM (last tested on GCC 11.2)
"nop \n"
"0: \n"
".quad 0, 0, 0, 0 \n" // TrampolineInstanceInfo
);
#else
#error Unsupported host architecture
#endif
}
extern char __start_HostToGuestTrampolineTemplate[];
extern char __stop_HostToGuestTrampolineTemplate[];
namespace FEXCore {
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
struct TrampolineInstanceInfo {
uintptr_t HostPacker;
uintptr_t CallCallback;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
};
struct GuestcallInfo {
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
bool operator==(const GuestcallInfo&) const noexcept = default;
};
struct GuestcallInfoHash {
size_t operator()(const GuestcallInfo& x) const noexcept {
// Hash only the target address, which is generally unique.
// For the unlikely case of a hash collision, std::unordered_map still picks the correct bucket entry.
return std::hash<uintptr_t>{}(x.GuestTarget);
}
};
// Bits in a SHA256 sum are already randomly distributed, so truncation yields a suitable hash function
struct TruncatingSHA256Hash {
size_t operator()(const FEXCore::IR::SHA256Sum& SHA256Sum) const noexcept {
return (const size_t&)SHA256Sum;
}
};
class ThunkHandler_impl final: public ThunkHandler {
std::shared_mutex ThunksMutex;
std::map<IR::SHA256Sum, ThunkedFunction*> Thunks = {
std::unordered_map<IR::SHA256Sum, ThunkedFunction*, TruncatingSHA256Hash> Thunks = {
{
// sha256(fex:loadlib)
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80},
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80 },
&LoadLib
},
{
// sha256(fex:is_lib_loaded)
{ 0xee, 0x57, 0xba, 0x0c, 0x5f, 0x6e, 0xef, 0x2a, 0x8c, 0xb5, 0x19, 0x81, 0xc9, 0x23, 0xe6, 0x51, 0xae, 0x65, 0x02, 0x8f, 0x2b, 0x5d, 0x59, 0x90, 0x6a, 0x7e, 0xe2, 0xe7, 0x1c, 0x33, 0x8a, 0xff },
&IsLibLoaded
},
{
// sha256(fex:link_address_to_function)
{ 0xe6, 0xa8, 0xec, 0x1c, 0x7b, 0x74, 0x35, 0x27, 0xe9, 0x4f, 0x5b, 0x6e, 0x2d, 0xc9, 0xa0, 0x27, 0xd6, 0x1f, 0x2b, 0x87, 0x8f, 0x2d, 0x35, 0x50, 0xea, 0x16, 0xb8, 0xc4, 0x5e, 0x42, 0xfd, 0x77 },
&LinkAddressToGuestFunction
},
{
// sha256(fex:make_host_trampoline_for_guest_function)
{ 0x1e, 0x51, 0x6b, 0x07, 0x39, 0xeb, 0x50, 0x59, 0xb3, 0xf3, 0x4f, 0xca, 0xdd, 0x58, 0x37, 0xe9, 0xf0, 0x30, 0xe5, 0x89, 0x81, 0xc7, 0x14, 0xfb, 0x24, 0xf9, 0xba, 0xe7, 0x0e, 0x00, 0x1e, 0x86 },
&MakeHostTrampolineForGuestFunction
}
};
// Can't be a string_view. We need to keep a copy of the library name in-case string_view pointer goes away.
// Ideally we track when a library has been unloaded and remove it from this set before the memory backing goes away.
std::set<std::string> Libs;
std::unordered_map<GuestcallInfo, uintptr_t, GuestcallInfoHash> GuestcallToHostTrampoline;
uint8_t *HostTrampolineInstanceDataPtr;
size_t HostTrampolineInstanceDataAvailable = 0;
/*
Set arg0/1 to arg regs, use CTX::HandleCallback to handle the callback
*/
@@ -56,13 +140,160 @@ namespace FEXCore {
Thread->CTX->HandleCallback(Thread, (uintptr_t)callback);
}
/**
* Instructs the Core to redirect calls to functions at the given
* address to another function. The original callee address is passed
* to the target function through an implicit argument stored in r11.
*
* The primary use case of this is ensuring that host function pointers
* returned from thunked APIs can safely be called by the guest.
*/
static void LinkAddressToGuestFunction(void* argsv) {
struct args_t {
uintptr_t original_callee;
uintptr_t target_addr; // Guest function to call when branching to original_callee
};
auto args = reinterpret_cast<args_t*>(argsv);
auto CTX = Thread->CTX;
LOGMAN_THROW_A_FMT(args->original_callee, "Tried to link null pointer address to guest function");
LOGMAN_THROW_A_FMT(args->target_addr, "Tried to link address to null pointer guest function");
if (!CTX->Config.Is64BitMode) {
LOGMAN_THROW_A_FMT((args->original_callee >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_A_FMT((args->target_addr >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
}
LogMan::Msg::DFmt("Thunks: Adding guest trampoline from address {:#x} to guest function {:#x}",
args->original_callee, args->target_addr);
auto Result = Thread->CTX->AddCustomIREntrypoint(
args->original_callee,
[CTX, GuestThunkEntrypoint = args->target_addr](uintptr_t Entrypoint, FEXCore::IR::IREmitter *emit) {
auto IRHeader = emit->_IRHeader(emit->Invalid(), 0);
auto Block = emit->CreateCodeNode();
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
const uint8_t GPRSize = CTX->GetGPRSize();
emit->_StoreContext(GPRSize, IR::GPRClass, emit->_Constant(Entrypoint), offsetof(Core::CPUState, gregs[X86State::REG_R11]));
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
}, CTX->ThunkHandler.get(), (void*)args->target_addr);
if (!Result) {
if (Result.Creator != CTX->ThunkHandler.get()) {
ERROR_AND_DIE_FMT("Input address for LinkAddressToGuestFunction is already linked by another module");
}
if (Result.Data != (void*)args->target_addr) {
// NOTE: This may happen in Vulkan thunks if the Vulkan driver resolves two different symbols
// to the same function (e.g. vkGetPhysicalDeviceFeatures2/vkGetPhysicalDeviceFeatures2KHR)
LogMan::Msg::EFmt("Input address for LinkAddressToGuestFunction is already linked elsewhere");
}
}
}
/**
* Generates a host-callable trampoline to call guest functions via the host ABI.
*
* This trampoline uses the same calling convention as the given HostPacker. Trampolines
* are cached, so it's safe to call this function repeatedly on the same arguments without
* leaking memory.
*
* Invoking the returned trampoline has the effect of:
* - packing the arguments (using the HostPacker identified by its SHA256)
* - performing a host->guest transition
* - unpacking the arguments via GuestUnpacker
* - calling the function at GuestTarget
*
* The primary use case of this is ensuring that guest function pointers ("callbacks")
* passed to thunked APIs can safely be called by the native host library.
*/
static void MakeHostTrampolineForGuestFunction(void* ArgsRV) {
struct ArgsRV_t {
IR::SHA256Sum *HostPackerSha256;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
uintptr_t rv; // Pointer to host trampoline + TrampolineInstanceInfo
} *args = reinterpret_cast<ArgsRV_t*>(ArgsRV);
LOGMAN_THROW_A_FMT(args->GuestTarget, "Tried to create host-trampoline to null pointer guest function");
const auto CTX = Thread->CTX;
const auto ThunkHandler = reinterpret_cast<ThunkHandler_impl *>(CTX->ThunkHandler.get());
const GuestcallInfo gci = { args->GuestUnpacker, args->GuestTarget };
// Try first with shared_lock
{
std::shared_lock lk(ThunkHandler->ThunksMutex);
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
std::lock_guard lk(ThunkHandler->ThunksMutex);
// Retry lookup with full lock before making a new trampoline to avoid double trampolines
{
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
// No entry found => create new trampoline
auto HostPackerEntry = ThunkHandler->Thunks.find(*args->HostPackerSha256);
if (HostPackerEntry == ThunkHandler->Thunks.end()) {
ERROR_AND_DIE_FMT("Unknown host packing function for callback");
}
LogMan::Msg::DFmt("Thunks: Adding host trampoline for guest function {:#x}",
args->GuestTarget);
const auto Length = __stop_HostToGuestTrampolineTemplate - __start_HostToGuestTrampolineTemplate;
const auto InstanceInfoOffset = Length - sizeof(TrampolineInstanceInfo);
if (ThunkHandler->HostTrampolineInstanceDataAvailable < Length) {
const auto allocation_step = 16 * 1024;
ThunkHandler->HostTrampolineInstanceDataAvailable = allocation_step;
ThunkHandler->HostTrampolineInstanceDataPtr = (uint8_t *)mmap(
0, ThunkHandler->HostTrampolineInstanceDataAvailable,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
LOGMAN_THROW_A_FMT(ThunkHandler->HostTrampolineInstanceDataPtr != MAP_FAILED, "Failed to mmap HostTrampolineInstanceDataPtr");
}
const TrampolineInstanceInfo NewTrampolineInfo {
.HostPacker = reinterpret_cast<uintptr_t>(HostPackerEntry->second),
.CallCallback = (uintptr_t)&CallCallback,
.GuestUnpacker = args->GuestUnpacker,
.GuestTarget = args->GuestTarget
};
uint8_t* const HostTrampoline = ThunkHandler->HostTrampolineInstanceDataPtr;
ThunkHandler->HostTrampolineInstanceDataAvailable -= Length;
ThunkHandler->HostTrampolineInstanceDataPtr += Length;
memcpy(HostTrampoline, (void*)&HostToGuestTrampolineTemplate, Length);
memcpy(HostTrampoline + InstanceInfoOffset, &NewTrampolineInfo, sizeof(NewTrampolineInfo));
args->rv = reinterpret_cast<uintptr_t>(HostTrampoline);
ThunkHandler->GuestcallToHostTrampoline[gci] = args->rv;
}
static void LoadLib(void *ArgsV) {
auto CTX = Thread->CTX;
auto Args = reinterpret_cast<LoadlibArgs*>(ArgsV);
auto Name = Args->Name;
auto CallbackThunks = Args->CallbackThunks;
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
@@ -75,13 +306,13 @@ namespace FEXCore {
const auto InitSym = std::string("fexthunks_exports_") + Name;
ExportEntry* (*InitFN)(void *, uintptr_t);
ExportEntry* (*InitFN)();
(void*&)InitFN = dlsym(Handle, InitSym.c_str());
if (!InitFN) {
ERROR_AND_DIE_FMT("LoadLib: Failed to find export {}", InitSym);
}
auto Exports = InitFN((void*)&CallCallback, CallbackThunks);
auto Exports = InitFN();
if (!Exports) {
ERROR_AND_DIE_FMT("LoadLib: Failed to initialize thunk library {}. "
"Check if the corresponding host library is installed "
@@ -91,7 +322,9 @@ namespace FEXCore {
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
{
std::unique_lock lk(That->ThunksMutex);
std::lock_guard lk(That->ThunksMutex);
That->Libs.insert(Name);
int i;
for (i = 0; Exports[i].sha256; i++) {
@@ -102,6 +335,23 @@ namespace FEXCore {
}
}
static void IsLibLoaded(void* ArgsRV) {
struct ArgsRV_t {
const char *Name;
bool rv;
};
auto &[Name, rv] = *reinterpret_cast<ArgsRV_t*>(ArgsRV);
auto CTX = Thread->CTX;
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
{
std::shared_lock lk(That->ThunksMutex);
rv = That->Libs.contains(Name);
}
}
public:
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) {
@@ -120,12 +370,6 @@ namespace FEXCore {
void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) {
::Thread = Thread;
}
ThunkHandler_impl() {
}
~ThunkHandler_impl() {
}
};
ThunkHandler* ThunkHandler::Create() {
+26 -16
View File
@@ -6,6 +6,7 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/Core/LookupCache.h>
#include <Interface/GDBJIT/GDBJIT.h>
#include <cstddef>
#include <cstdint>
@@ -17,6 +18,7 @@
#include <unistd.h>
#include <xxhash.h>
namespace FEXCore::IR {
AOTIRInlineEntry *AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
@@ -242,9 +244,9 @@ namespace FEXCore::IR {
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
PreGenerateIRFetchResult Result{};
if (AOTIRCacheEntry.Entry) {
AOTIRCacheEntry.Entry->ContainsCode = true;
@@ -253,7 +255,7 @@ namespace FEXCore::IR {
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - AOTIRCacheEntry.Offset);
auto AOTEntry = Mod->Find(GuestRIP - AOTIRCacheEntry.VAFileStart);
if (AOTEntry) {
// verify hash
@@ -263,7 +265,7 @@ namespace FEXCore::IR {
Result.IRList = AOTEntry->GetIRData();
//LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
Result.RAData = AOTEntry->GetRAData();;
Result.RAData = AOTEntry->GetRAData()->CreateCopy();
Result.DebugData = new FEXCore::Core::DebugData();
Result.StartAddr = MappedStart;
Result.Length = AOTEntry->GuestLength;
@@ -287,12 +289,13 @@ namespace FEXCore::IR {
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::RegisterAllocationData::UniquePtr RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR) {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming()) {
if (GeneratedIR || CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
@@ -301,18 +304,26 @@ namespace FEXCore::IR {
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, AOTIRCacheEntry.Entry->Filename);
}
if (CTX->Config.GDBSymbols()) {
GDBJITRegister(AOTIRCacheEntry.Entry, AOTIRCacheEntry.VAFileStart, GuestRIP, (uintptr_t)CodePtr, DebugData);
}
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && RAData &&
(CTX->Config.AOTIRCapture() || CTX->Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - AOTIRCacheEntry.Offset;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.Offset;
auto LocalRIP = GuestRIP - AOTIRCacheEntry.VAFileStart;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.VAFileStart;
auto FileId = AOTIRCacheEntry.Entry->FileId;
// The underlying pointer and the unique_ptr deleter for RAData must
// be marshalled separately to the lambda below. Otherwise, the
// lambda can't be used as an std::function due to being non-copyable
auto RADataCopy = RAData->CreateCopy();
auto RADataCopyDeleter = RADataCopy.get_deleter();
auto IRListCopy = IRList->CreateCopy();
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy, FileId]() {
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy=RADataCopy.release(), RADataCopyDeleter, FileId]() {
// It is guaranteed via AOTIRCaptureCacheWriteoutLock and AOTIRCaptureCacheWriteoutFlusing that this will not run concurrently
// Memory coherency is guaranteed via AOTIRCaptureCacheWriteoutLock
@@ -325,7 +336,7 @@ namespace FEXCore::IR {
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy);
FEXCore::Allocator::free(RADataCopy);
RADataCopyDeleter(RADataCopy);
delete IRListCopy;
});
@@ -339,18 +350,17 @@ namespace FEXCore::IR {
// Insert to caches if we generated IR
if (GeneratedIR) {
if (Thread->CPUBackend->NeedsRetainedIRCopy()) {
if (CTX->GetGdbServerStatus()) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), std::move(RAData), decltype(Entry.DebugData)(DebugData)};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
Thread->DebugStore.insert({GuestRIP, std::move(Entry)});
}
else {
// If the IR doesn't need to be retained then we can just delete it now
delete DebugData;
delete RAData;
delete IRList;
if (IRList->IsCopy()) delete IRList;
}
}
}
@@ -374,7 +384,7 @@ namespace FEXCore::IR {
std::unique_lock lk(AOTIRCacheLock);
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry{0, 0, 0, fileid, filename, false}});
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry { .FileId = fileid, .Filename = filename }});
auto Entry = &(Inserted.first->second);
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
+5 -2
View File
@@ -1,5 +1,6 @@
#pragma once
#include "FEXCore/IR/RegisterAllocationData.h"
#include <FEXCore/Config/Config.h>
#include <atomic>
@@ -11,6 +12,7 @@
#include <unordered_map>
#include <shared_mutex>
#include <queue>
#include <FEXCore/HLE/SourcecodeResolver.h>
namespace FEXCore::Core {
struct DebugData;
@@ -74,6 +76,7 @@ namespace FEXCore::IR {
AOTIRInlineIndex *Array;
void *FilePtr;
size_t Size;
std::unique_ptr<FEXCore::HLE::SourcecodeMap> SourcecodeMap;
std::string FileId;
std::string Filename;
bool ContainsCode;
@@ -93,7 +96,7 @@ namespace FEXCore::IR {
struct PreGenerateIRFetchResult {
FEXCore::IR::IRListView *IRList {};
FEXCore::IR::RegisterAllocationData *RAData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
FEXCore::Core::DebugData *DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
@@ -106,7 +109,7 @@ namespace FEXCore::IR {
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::RegisterAllocationData::UniquePtr RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR);
+24 -15
View File
@@ -181,6 +181,11 @@
"RAOverride": "0"
},
"GuestOpcode u32:$GuestEntryOffset": {
"Desc": ["Marks the beginning of a guest opcode"],
"HasSideEffects": true
},
"GPR = ValidateCode u64:$CodeOriginalLow, u64:$CodeOriginalhigh, i64:$Offset, u8:$CodeLength": {
"HasSideEffects": true,
"HasDest": true,
@@ -294,12 +299,6 @@
],
"DestSize": "16",
"NumElements": "2"
},
"GuestCallDirect u64:$RIP, u64:$NextRIP": {
"HasSideEffects": true
},
"GuestCallIndirect GPR:$RIP, u64:$NextRIP": {
"HasSideEffects": true
}
},
"Moves": {
@@ -943,7 +942,7 @@
},
"FPR = VBitcast u8:#RegisterSize, u8:#ElementSize, FPR:$Source": {
"Dest": ["Workaround for issue with LLVM breaking when loading scalar elements to vectors"],
"Desc": ["Workaround for issue with LLVM breaking when loading scalar elements to vectors"],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
@@ -1277,12 +1276,12 @@
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VUMull2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Dest": "Multiplies the high elements with size extension",
"Desc": "Multiplies the high elements with size extension",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VSMull2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Dest": "Multiplies the high elements with size extension",
"Desc": "Multiplies the high elements with size extension",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
@@ -1454,33 +1453,43 @@
},
"Crypto": {
"FPR = VAESImc FPR:$Vector": {
"Dest": "Does a stage of the inverse mix column transformation",
"Desc": "Does a stage of the inverse mix column transformation",
"DestSize": "16"
},
"FPR = VAESEnc FPR:$State, FPR:$Key": {
"Dest": "Does a step of AES encryption",
"Desc": "Does a step of AES encryption",
"DestSize": "16"
},
"FPR = VAESEncLast FPR:$State, FPR:$Key": {
"Dest": "Does the last step of AES encryption",
"Desc": "Does the last step of AES encryption",
"DestSize": "16"
},
"FPR = VAESDec FPR:$State, FPR:$Key": {
"Dest": "Does a step of AES decryption",
"Desc": "Does a step of AES decryption",
"DestSize": "16"
},
"FPR = VAESDecLast FPR:$State, FPR:$Key": {
"Dest": "Does the last step of AES decryption",
"Desc": "Does the last step of AES decryption",
"DestSize": "16"
},
"FPR = VAESKeyGenAssist FPR:$Src, u8:$RCON": {
"Dest": "Assists in key generation",
"Desc": "Assists in key generation",
"DestSize": "16"
},
"GPR = CRC32 GPR:$Src1, GPR:$Src2, u8:$SrcSize": {
"Desc": ["CRC32 using polynomial 0x1EDC6F41"
],
"DestSize": "std::max<uint8_t>(4, GetOpSize(_Src1))"
},
"FPR = PCLMUL FPR:$Src1, FPR:$Src2, u8:$Selector": {
"Desc": [
"Performs carryless multiplication of 64-bit elements depending on the selector.",
"Selector = 0b00000000: Uses low 64-bit elements from both input vectors",
"Selector = 0b00000001: Uses high 64-bit element from Src1 and low 64-bit element from Src2",
"Selector = 0b00010000: Uses low 64-bit element from Src1 and high 64-bit element from Src2",
"Selector = 0b00010001: Uses high 64-bit elements from both input vectors"
],
"DestSize": "16"
}
},
"F64": {
+21
View File
@@ -16,6 +16,27 @@ $end_info$
#include <vector>
namespace FEXCore::IR {
bool IsFragmentExit(FEXCore::IR::IROps Op) {
switch (Op) {
case OP_EXITFUNCTION:
case OP_BREAK:
return true;
default:
return false;
}
}
bool IsBlockExit(FEXCore::IR::IROps Op) {
switch(Op) {
case OP_JUMP:
case OP_CONDJUMP:
return true;
default:
return IsFragmentExit(Op);
}
}
FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(OrderedNode *Node) {
auto Class = GetOpRegClass(Node);
switch (Class) {
@@ -107,11 +107,11 @@ namespace {
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < 16; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GPRS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gregs[0]) + sizeof(FEXCore::Core::CPUState::gregs[0]) * i,
sizeof(FEXCore::Core::CPUState::gregs[0]),
FEXCore::Core::CPUState::GPR_REG_SIZE,
},
DefaultAccess[1],
FEXCore::IR::InvalidClass,
@@ -127,11 +127,11 @@ namespace {
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < 16; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, xmm[0][0]) + sizeof(FEXCore::Core::CPUState::xmm[0]) * i,
sizeof(FEXCore::Core::CPUState::xmm[0]),
FEXCore::Core::CPUState::XMM_REG_SIZE,
},
DefaultAccess[3],
FEXCore::IR::InvalidClass,
@@ -192,11 +192,11 @@ namespace {
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < (sizeof(FEXCore::Core::CPUState::flags) / sizeof(FEXCore::Core::CPUState::flags[0])); ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, flags[0]) + sizeof(FEXCore::Core::CPUState::flags[0]) * i,
sizeof(FEXCore::Core::CPUState::flags[0]),
FEXCore::Core::CPUState::FLAG_SIZE,
},
DefaultAccess[10],
FEXCore::IR::InvalidClass,
@@ -212,11 +212,11 @@ namespace {
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_MMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, mm[0][0]) + sizeof(FEXCore::Core::CPUState::mm[0]) * i,
sizeof(FEXCore::Core::CPUState::mm[0]),
FEXCore::Core::CPUState::MM_REG_SIZE
},
DefaultAccess[12],
FEXCore::IR::InvalidClass,
@@ -224,7 +224,7 @@ namespace {
}
// GDTs
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GDTS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gdt[0]) + sizeof(FEXCore::Core::CPUState::gdt[0]) * i,
@@ -286,12 +286,12 @@ namespace {
};
size_t Offset = 0;
SetAccess(Offset++, DefaultAccess[0]);
for (size_t i = 0; i < 16; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GPRS; ++i) {
SetAccess(Offset++, DefaultAccess[1]);
}
SetAccess(Offset++, DefaultAccess[2]);
for (size_t i = 0; i < 16; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
SetAccess(Offset++, DefaultAccess[3]);
}
@@ -303,17 +303,17 @@ namespace {
SetAccess(Offset++, DefaultAccess[9]);
for (size_t i = 0; i < (sizeof(FEXCore::Core::CPUState::flags) / sizeof(FEXCore::Core::CPUState::flags[0])); ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
SetAccess(Offset++, DefaultAccess[10]);
}
SetAccess(Offset++, DefaultAccess[11]);
for (size_t i = 0; i < 8; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_MMS; ++i) {
SetAccess(Offset++, DefaultAccess[12]);
}
for (size_t i = 0; i < 32; ++i) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GDTS; ++i) {
SetAccess(Offset++, DefaultAccess[13]);
}
@@ -623,7 +623,7 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
// Loop through non-reserved flag stores and eliminate unused ones.
for (unsigned F = 0; F < 32; F++) {
for (size_t F = 0; F < Core::CPUState::NUM_EFLAG_BITS; F++) {
if (!(Op->Flags & (1ULL << F))) {
continue;
}
@@ -101,9 +101,9 @@ uint64_t FPRBit(uint32_t Offset, uint32_t Size) {
return 0;
}
auto begin = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto begin = offsetof(Core::CpuStateFrame, State.xmm[0][0]);
auto regn = (Offset - begin)/16;
auto regn = (Offset - begin) / Core::CPUState::XMM_REG_SIZE;
auto bitn = regn * 3;
if (!IsTrackedWriteFPR(Offset, Size))
@@ -244,16 +244,10 @@ bool IRValidation::Run(IREmitter *IREmit) {
// Blocks need to have an instruction that leaves the block in some way before the EndBlock instruction
{
auto Op = GetOp(CodeCurrent);
switch (Op) {
case OP_EXITFUNCTION:
case OP_JUMP:
case OP_CONDJUMP:
case OP_BREAK:
break;
default:
HadError |= true;
Errors << "%ssa" << BlockID << " Didn't have an exit IR op as its last instruction" << std::endl;
};
if (!IsBlockExit(Op)) {
HadError |= true;
Errors << "%ssa" << BlockID << " Didn't have a block exit IR op as its last instruction" << std::endl;
}
}
}
}
@@ -89,7 +89,7 @@ namespace {
};
struct RegisterGraph {
std::unique_ptr<IR::RegisterAllocationData, IR::RegisterAllocationDataDeleter> AllocData;
IR::RegisterAllocationData::UniquePtr AllocData;
RegisterSet Set;
std::vector<RegisterNode> Nodes{};
uint32_t NodeCount{};
@@ -153,10 +153,7 @@ namespace {
Graph->Nodes.resize(NodeCount);
Graph->VisitedNodePredecessors.clear();
Graph->AllocData.reset((FEXCore::IR::RegisterAllocationData*)FEXCore::Allocator::malloc(FEXCore::IR::RegisterAllocationData::Size(NodeCount)));
memset(&Graph->AllocData->Map[0], PhysicalRegister::Invalid().Raw, NodeCount);
Graph->AllocData->MapCount = NodeCount;
Graph->AllocData->IsShared = false; // not shared by default
Graph->AllocData = RegisterAllocationData::Create(NodeCount);
Graph->NodeCount = NodeCount;
}
@@ -283,7 +280,7 @@ namespace {
* Top 32bits is the class, lower 32bits is the register
*/
RegisterAllocationData* GetAllocationData() override;
std::unique_ptr<RegisterAllocationData, RegisterAllocationDataDeleter> PullAllocationData() override;
RegisterAllocationData::UniquePtr PullAllocationData() override;
private:
using BlockInterferences = std::vector<IR::NodeID>;
@@ -379,7 +376,7 @@ namespace {
return Graph->AllocData.get();
}
std::unique_ptr<RegisterAllocationData, RegisterAllocationDataDeleter> ConstrainedRAPass::PullAllocationData() {
RegisterAllocationData::UniquePtr ConstrainedRAPass::PullAllocationData() {
return std::move(Graph->AllocData);
}
@@ -560,9 +557,11 @@ namespace {
// Is an OP_LOADREGISTER eligible to read directly from the SRA reg?
auto IsAliasable = [](uint8_t Size, RegisterClassType StaticClass, uint32_t Offset) {
if (StaticClass == GPRFixedClass) {
return (Size == 8 /*|| Size == 4*/) && ((Offset & 7) == 0); // We need more meta info to support not-size-of-reg
// We need more meta info to support not-size-of-reg
return (Size == 8 /*|| Size == 4*/) && ((Offset & 7) == 0);
} else if (StaticClass == FPRFixedClass) {
return (Size == 16 /*|| Size == 8 || Size == 4*/) && ((Offset & 15) == 0); // We need more meta info to support not-size-of-reg
// We need more meta info to support not-size-of-reg
return (Size == 16 /*|| Size == 8 || Size == 4*/) && ((Offset & 15) == 0);
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected static class {}", StaticClass);
}
@@ -578,10 +577,10 @@ namespace {
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[16][0]);
if (Offset >= beginGpr && Offset < endGpr) {
auto reg = (Offset - beginGpr) / 8;
auto reg = (Offset - beginGpr) / Core::CPUState::GPR_REG_SIZE;
return PhysicalRegister(GPRFixedClass, reg);
} else if (Offset >= beginFpr && Offset < endFpr) {
auto reg = (Offset - beginFpr) / 16;
auto reg = (Offset - beginFpr) / Core::CPUState::XMM_REG_SIZE;
return PhysicalRegister(FPRFixedClass, reg);
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected Offset {}", Offset);
@@ -602,10 +601,10 @@ namespace {
auto endFpr = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[16][0]);
if (Offset >= beginGpr && Offset < endGpr) {
auto reg = (Offset - beginGpr) / 8;
auto reg = (Offset - beginGpr) / Core::CPUState::GPR_REG_SIZE;
return &StaticMaps[reg];
} else if (Offset >= beginFpr && Offset < endFpr) {
auto reg = (Offset - beginFpr) / 16;
auto reg = (Offset - beginFpr) / Core::CPUState::XMM_REG_SIZE;
return &StaticMaps[GprSize + reg];
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected offset {}", Offset);
@@ -24,15 +24,15 @@ public:
};
bool IsStaticAllocGpr(uint32_t Offset, RegisterClassType Class) {
const auto begin = offsetof(FEXCore::Core::CPUState, gregs[0]);
const auto end = offsetof(FEXCore::Core::CPUState, gregs[16]);
const auto begin = offsetof(Core::CPUState, gregs[0]);
const auto end = offsetof(Core::CPUState, gregs[16]);
if (Offset >= begin && Offset < end) {
const auto reg = (Offset - begin) / 8;
const auto reg = (Offset - begin) / Core::CPUState::GPR_REG_SIZE;
LOGMAN_THROW_A_FMT(Class == IR::GPRClass, "unexpected Class {}", Class);
// 0..15 -> 16 in total
return reg < 16;
return reg < Core::CPUState::NUM_GPRS;
}
return false;
@@ -43,11 +43,11 @@ bool IsStaticAllocFpr(uint32_t Offset, RegisterClassType Class, bool AllowGpr) {
const auto end = offsetof(FEXCore::Core::CPUState, xmm[16][0]);
if (Offset >= begin && Offset < end) {
const auto reg = (Offset - begin) / 16;
const auto reg = (Offset - begin) / Core::CPUState::XMM_REG_SIZE;
LOGMAN_THROW_A_FMT(Class == IR::FPRClass || (AllowGpr && Class == IR::GPRClass), "unexpected Class {}, AllowGpr {}", Class, AllowGpr);
// 0..15 -> 16 in total
return reg < 16;
return reg < Core::CPUState::NUM_XMMS;
}
return false;
@@ -145,28 +145,23 @@ bool ValueDominanceValidation::Run(IREmitter *IREmit) {
// ...
// We need to walk the predecessors to see if the value comes from there
std::set<IR::OrderedNode *> Predecessors;
std::set<IR::OrderedNode *> Predecessors { BlockNode };
// Recursively gather all predecessors of BlockNode
for (auto NodeIt = Predecessors.begin(); NodeIt != Predecessors.end();) {
auto PredBlock = &OffsetToBlockMap.try_emplace(CurrentIR.GetID(*NodeIt)).first->second;
++NodeIt;
std::function<void(IR::OrderedNode*)> AddPredecessors = [&] (IR::OrderedNode *Node) {
auto PredBlock = &OffsetToBlockMap.try_emplace(CurrentIR.GetID(Node)).first->second;
// Current Block will always be in set
// Walk each predecessor, adding their predecessors
// Leave if all the predecessors are already in the set
for (auto &Pred : PredBlock->Predecessors) {
for (auto *Pred : PredBlock->Predecessors) {
if (Predecessors.insert(Pred).second) {
// If this block didn't exist then walk its predecessors as well
AddPredecessors(Pred);
// New blocks added, so repeat from the beginning to pull in their predecessors
NodeIt = Predecessors.begin();
}
}
};
AddPredecessors(BlockNode);
}
bool FoundPredDefine = false;
for (auto it = Predecessors.begin(); it != Predecessors.end(); ++it) {
IR::OrderedNode *Pred = *it;
for (auto* Pred : Predecessors) {
auto PredIROp = CurrentIR.GetOp<FEXCore::IR::IROp_CodeBlock>(Pred);
if (Arg.ID() >= PredIROp->Begin.ID() &&
+12
View File
@@ -0,0 +1,12 @@
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <string>
namespace FEXCore::Paths {
FEX_DEFAULT_VISIBILITY const char *GetHomeDirectory();
FEX_DEFAULT_VISIBILITY std::string GetCachePath();
FEX_DEFAULT_VISIBILITY std::string GetEntryCachePath();
}
+1
View File
@@ -163,6 +163,7 @@ namespace Type {
FEX_DEFAULT_VISIBILITY void Load();
FEX_DEFAULT_VISIBILITY void ReloadMetaLayer();
FEX_DEFAULT_VISIBILITY std::string FindContainer();
FEX_DEFAULT_VISIBILITY std::string FindContainerPrefix();
FEX_DEFAULT_VISIBILITY void AddLayer(std::unique_ptr<FEXCore::Config::Layer> _Layer);
+40 -22
View File
@@ -11,6 +11,8 @@ $end_info$
#include <cstdint>
#include <string>
#include <memory>
#include <vector>
namespace FEXCore {
@@ -23,6 +25,7 @@ namespace Core {
struct DebugData;
struct ThreadState;
struct CpuStateFrame;
struct InternalThreadState;
}
namespace CodeSerialize {
@@ -30,13 +33,24 @@ namespace CodeSerialize {
}
namespace CPU {
class InterpreterCore;
class JITCore;
class LLVMCore;
struct CPUBackendFeatures {
bool SupportsStaticRegisterAllocation = false;
};
class CPUBackend {
public:
virtual ~CPUBackend() = default;
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
/**
* @param InitialCodeSize - Initial size for the code buffers
* @param MaxCodeSize - Max size for the code buffers
*/
CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize);
virtual ~CPUBackend();
/**
* @return The name of this backend
*/
@@ -61,7 +75,7 @@ class LLVMCore;
[[nodiscard]] virtual void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) = 0;
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
@@ -98,31 +112,35 @@ class LLVMCore;
*/
[[nodiscard]] virtual bool NeedsOpDispatch() = 0;
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
}
virtual void ClearCache() {}
virtual bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const { return false; }
/**
* @brief Does this CPUBackend need its IR to stick around for correct emulation
*
* This should only be used on the interpreter, all other backends can clear their IR
*/
virtual bool NeedsRetainedIRCopy() const { return false; }
/**
* @brief Clear any relocations after JIT compiling
*/
virtual void ClearRelocations() {}
using AsmDispatch = FEX_NAKED void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = FEX_NAKED void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
bool IsAddressInCodeBuffer(uintptr_t Address) const;
JITCallback CallbackPtr{};
protected:
AsmDispatch DispatchPtr{};
FEXCore::Core::InternalThreadState *ThreadState;
size_t InitialCodeSize, MaxCodeSize;
[[nodiscard]] CodeBuffer *GetEmptyCodeBuffer();
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
private:
CodeBuffer AllocateNewCodeBuffer(size_t Size);
void FreeCodeBuffer(CodeBuffer Buffer);
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
// This is the array of code buffers. Unless signals force us to keep more than
// buffer, there will be only one entry here
std::vector<CodeBuffer> CodeBuffers{};
};
}
+20 -1
View File
@@ -11,6 +11,8 @@
#include <ostream>
#include <memory>
#include <set>
#include <mutex>
#include <shared_mutex>
namespace FEXCore {
class CodeLoader;
@@ -33,6 +35,7 @@ namespace FEXCore::HLE {
namespace FEXCore::IR {
struct AOTIRCacheEntry;
class IREmitter;
}
namespace FEXCore::Context {
@@ -51,6 +54,19 @@ namespace FEXCore::Context {
MODE_64BIT,
};
struct CustomIRResult {
void *Creator;
void *Data;
explicit operator bool() const noexcept { return !lock; }
CustomIRResult(std::unique_lock<std::shared_mutex> &&lock, void *Creator, void *Data):
Creator(Creator), Data(Data), lock(std::move(lock)) { }
private:
std::unique_lock<std::shared_mutex> lock;
};
using CustomCPUFactoryType = std::function<std::unique_ptr<FEXCore::CPU::CPUBackend> (FEXCore::Context::Context*, FEXCore::Core::InternalThreadState *Thread)>;
using ExitHandler = std::function<void(uint64_t ThreadId, FEXCore::Context::ExitReason)>;
@@ -95,7 +111,7 @@ namespace FEXCore::Context {
*
* @return true if we loaded code
*/
FEX_DEFAULT_VISIBILITY FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader);
FEX_DEFAULT_VISIBILITY FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, uint64_t InitialRIP, uint64_t StackPointer);
FEX_DEFAULT_VISIBILITY void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler);
FEX_DEFAULT_VISIBILITY ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX);
@@ -250,6 +266,9 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY void FinalizeAOTIRCache(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY void WriteFilesWithCode(FEXCore::Context::Context *CTX, std::function<void(const std::string& fileid, const std::string& filename)> Writer);
FEX_DEFAULT_VISIBILITY void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length);
FEX_DEFAULT_VISIBILITY void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback);
FEX_DEFAULT_VISIBILITY void MarkMemoryShared(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress);
FEX_DEFAULT_VISIBILITY CustomIRResult AddCustomIREntrypoint(FEXCore::Context::Context *CTX, uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator = nullptr, void *Data = nullptr);
}
+60 -38
View File
@@ -2,6 +2,7 @@
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Core/CPUBackend.h>
#include <atomic>
#include <cstddef>
@@ -27,6 +28,20 @@ namespace FEXCore::Core {
} gdt[32];
uint16_t FCW;
uint16_t FTW;
static constexpr size_t FLAG_SIZE = sizeof(flags[0]);
static constexpr size_t GDT_SIZE = sizeof(gdt[0]);
static constexpr size_t GPR_REG_SIZE = sizeof(gregs[0]);
static constexpr size_t XMM_REG_SIZE = sizeof(xmm[0]);
static constexpr size_t MM_REG_SIZE = sizeof(mm[0]);
// Only the first 32 bits are defined.
static constexpr size_t NUM_EFLAG_BITS = 32;
static constexpr size_t NUM_FLAGS = sizeof(flags) / FLAG_SIZE;
static constexpr size_t NUM_GDTS = sizeof(gdt) / GDT_SIZE;
static constexpr size_t NUM_GPRS = sizeof(gregs) / GPR_REG_SIZE;
static constexpr size_t NUM_XMMS = sizeof(xmm) / XMM_REG_SIZE;
static constexpr size_t NUM_MMS = sizeof(mm) / MM_REG_SIZE;
};
static_assert(offsetof(CPUState, xmm) % 16 == 0, "xmm needs to be 128bit aligned!");
@@ -92,13 +107,11 @@ namespace FEXCore::Core {
OPINDEX_MAX,
};
union JITPointers {
struct JITPointers {
struct {
// Process specific
uint64_t LUDIV{};
uint64_t LDIV{};
uint64_t LUREM{};
uint64_t LREM{};
uint64_t PrintValue{};
uint64_t PrintVectorValue{};
uint64_t RemoveThreadCodeEntryFromJIT{};
@@ -106,58 +119,57 @@ namespace FEXCore::Core {
uint64_t CPUIDFunction{};
uint64_t SyscallHandlerObj{};
uint64_t SyscallHandlerFunc{};
uint64_t ExitFunctionLink{};
uint64_t FallbackHandlerPointers[FallbackHandlerIndex::OPINDEX_MAX];
// Thread Specific
uint64_t SignalHandlerRefCountPointer{};
/**
* @name Dispatcher pointers
* @{ */
uint64_t DispatcherLoopTop{};
uint64_t DispatcherLoopTopFillSRA{};
uint64_t ExitFunctionLinker{};
uint64_t ThreadStopHandlerSpillSRA{};
uint64_t ThreadPauseHandlerSpillSRA{};
uint64_t UnimplementedInstructionHandler{};
uint64_t OverflowExceptionHandler{};
uint64_t SignalReturnHandler{};
uint64_t L1Pointer{};
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
uint64_t L2Pointer{};
/** @} */
} AArch64;
} Common;
struct {
// Process specific
uint64_t PrintValue{};
uint64_t PrintVectorValue{};
uint64_t RemoveThreadCodeEntryFromJIT{};
uint64_t CPUIDObj{};
uint64_t CPUIDFunction{};
uint64_t SyscallHandlerObj{};
uint64_t SyscallHandlerFunc{};
union {
struct {
// Process specific
uint64_t LUDIV{};
uint64_t LDIV{};
uint64_t LUREM{};
uint64_t LREM{};
uint64_t FallbackHandlerPointers[FallbackHandlerIndex::OPINDEX_MAX];
// Thread Specific
// Thread Specific
uint64_t SignalHandlerRefCountPointer{};
/**
* @name Dispatcher pointers
* @{ */
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
/** @} */
} AArch64;
/**
* @name Dispatcher pointers
* @{ */
uint64_t DispatcherLoopTop{};
uint64_t DispatcherLoopTopFillSRA{};
uint64_t ThreadStopHandler{};
uint64_t ThreadPauseHandler{};
uint64_t UnimplementedInstructionHandler{};
uint64_t OverflowExceptionHandler{};
uint64_t SignalReturnHandler{};
uint64_t L1Pointer{};
/** @} */
} X86;
struct {
// None so far
} X86;
struct {
uint64_t FragmentExecuter;
using IntCallbackReturn = void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
IntCallbackReturn CallbackReturn;
} Interpreter;
};
};
// Each guest JIT frame has one of these
@@ -177,8 +189,18 @@ namespace FEXCore::Core {
* ARM64:
* - Bit 15: In syscall
* - Bit 14-0: Number of static registers spilled
*/
*/
uint64_t InSyscallInfo{};
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
} SynchronousFaultData;
InternalThreadState* Thread;
// Pointers that the JIT needs to load to remove relocations
@@ -0,0 +1,22 @@
#include <cstddef>
#include <cstdint>
#include <gdb/jit-reader.h>
// everything is stored inline as it is marshaled cross process by gdb
struct blocks_t {
char name[512];
GDB_CORE_ADDR start;
GDB_CORE_ADDR end;
};
struct info_t {
char filename[512];
ptrdiff_t blocks_ofs;
ptrdiff_t lines_ofs;
int nblocks;
int nlines;
};
@@ -8,6 +8,7 @@
#include <FEXCore/Utils/InterruptableConditionVariable.h>
#include <FEXCore/Utils/Threads.h>
#include <map>
#include <unordered_map>
#include <shared_mutex>
@@ -41,10 +42,15 @@ namespace FEXCore::Core {
};
struct DebugDataSubblock {
uintptr_t HostCodeStart;
uint32_t HostCodeOffset;
uint32_t HostCodeSize;
};
struct DebugDataGuestOpcode {
uint64_t GuestEntryOffset;
ptrdiff_t HostEntryOffset;
};
/**
* @brief Contains debug data for a block of code for later debugger analysis
*
@@ -53,6 +59,7 @@ namespace FEXCore::Core {
struct DebugData {
uint64_t HostCodeSize; ///< The size of the code generated in the host JIT
std::vector<DebugDataSubblock> Subblocks;
std::vector<DebugDataGuestOpcode> GuestOpcodes;
std::vector<FEXCore::CPU::Relocation> *Relocations;
};
@@ -67,12 +74,12 @@ namespace FEXCore::Core {
uint64_t StartAddr;
uint64_t Length;
std::unique_ptr<FEXCore::IR::IRListView, FEXCore::IR::IRListViewDeleter> IR;
std::unique_ptr<FEXCore::IR::RegisterAllocationData, FEXCore::IR::RegisterAllocationDataDeleter> RAData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
std::unique_ptr<FEXCore::Core::DebugData> DebugData;
};
struct InternalThreadState {
FEXCore::Core::CpuStateFrame* CurrentFrame = &BaseFrameState;
FEXCore::Core::CpuStateFrame* const CurrentFrame = &BaseFrameState;
struct {
std::atomic_bool Running {false};
@@ -93,7 +100,7 @@ namespace FEXCore::Core {
std::unique_ptr<FEXCore::CPU::CPUBackend> CPUBackend;
std::unique_ptr<FEXCore::LookupCache> LookupCache;
std::unordered_map<uint64_t, LocalIREntry> LocalIRCache;
std::unordered_map<uint64_t, LocalIREntry> DebugStore;
std::unique_ptr<FEXCore::Frontend::Decoder> FrontendDecoder;
std::unique_ptr<FEXCore::IR::PassManager> PassManager;
@@ -107,7 +114,7 @@ namespace FEXCore::Core {
std::shared_mutex ObjectCacheRefCounter{};
bool DestroyedByParent{false}; // Should the parent destroy this thread, or it destory itself
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState{};
};
+1 -1
View File
@@ -186,7 +186,7 @@ struct DecodedInst {
bool DecodedSIB;
DecodedOperand Dest;
DecodedOperand Src[2];
DecodedOperand Src[3];
// Constains the dispatcher handler pointer
X86InstInfo const* TableInfo;
@@ -0,0 +1,96 @@
#pragma once
#include <algorithm>
#include <string>
#include <vector>
#include <memory>
#include <filesystem>
#include <fmt/format.h>
namespace FEXCore::IR {
struct AOTIRCacheEntry;
}
namespace FEXCore::HLE {
struct SourcecodeLineMapping {
uintptr_t FileGuestBegin;
uintptr_t FileGuestEnd;
int LineNumber;
};
struct SourcecodeSymbolMapping {
uintptr_t FileGuestBegin;
uintptr_t FileGuestEnd;
std::string Name;
static std::string SymName(const SourcecodeSymbolMapping *Sym, const std::string &GuestFilename, uintptr_t HostEntry, uintptr_t FileBegin) {
if (Sym) {
auto SymOffset = FileBegin - Sym->FileGuestBegin;
if (SymOffset) {
return fmt::format("{}: {}+{} @{:x}", std::filesystem::path(GuestFilename).stem().string(), Sym->Name,
SymOffset, HostEntry);
} else {
return fmt::format("{}: {} @{:x}", std::filesystem::path(GuestFilename).stem().string(), Sym->Name,
HostEntry);
}
} else {
return fmt::format("{}: +{} @{:x}", std::filesystem::path(GuestFilename).stem().string(), FileBegin,
HostEntry);
}
}
};
struct SourcecodeMap {
std::string SourceFile;
std::vector<SourcecodeLineMapping> SortedLineMappings;
std::vector<SourcecodeSymbolMapping> SortedSymbolMappings;
template<typename F>
void IterateLineMappings(uintptr_t FileBegin, uintptr_t Size, const F &Callback) const {
auto Begin = FileBegin;
auto End = FileBegin + Size;
auto Found = std::lower_bound(SortedLineMappings.cbegin(), SortedLineMappings.cend(), Begin, [](const auto &Range, const auto Position) {
return Range.FileGuestEnd <= Position;
});
while (Found != SortedLineMappings.cend()) {
if (Found->FileGuestBegin < End && Found->FileGuestEnd > Begin) {
Callback(Found);
} else {
break;
}
Found++;
}
}
const SourcecodeLineMapping *FindLineMapping(uintptr_t FileBegin) const {
return Find(FileBegin, SortedLineMappings);
}
const SourcecodeSymbolMapping *FindSymbolMapping(uintptr_t FileBegin) const {
return Find(FileBegin, SortedSymbolMappings);
}
private:
template<typename VecT>
const typename VecT::value_type *Find(uintptr_t FileBegin, const VecT &SortedMappings) const {
auto Found = std::lower_bound(SortedMappings.cbegin(), SortedMappings.cend(), FileBegin, [](const auto &Range, const auto Position) {
return Range.FileGuestEnd <= Position;
});
if (Found != SortedMappings.end() && Found->FileGuestBegin <= FileBegin && Found->FileGuestEnd > FileBegin) {
return &(*Found);
} else {
return {};
}
}
};
class SourcecodeResolver {
public:
virtual std::unique_ptr<SourcecodeMap> GenerateMap(const std::string_view& GuestBinaryFile, const std::string_view& GuestBinaryFileId) = 0;
};
}
+6 -4
View File
@@ -50,9 +50,11 @@ namespace FEXCore::HLE {
};
class SyscallHandler;
class SourcecodeResolver;
struct AOTIRCacheEntryLookupResult {
AOTIRCacheEntryLookupResult(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t Offset, FHU::ScopedSignalMaskWithSharedLock &&lk)
: Entry(Entry), Offset(Offset), lk(std::move(lk))
AOTIRCacheEntryLookupResult(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart, FHU::ScopedSignalMaskWithSharedLock &&lk)
: Entry(Entry), VAFileStart(VAFileStart), lk(std::move(lk))
{
}
@@ -60,7 +62,7 @@ namespace FEXCore::HLE {
AOTIRCacheEntryLookupResult(AOTIRCacheEntryLookupResult&&) = default;
FEXCore::IR::AOTIRCacheEntry *Entry;
uintptr_t Offset;
uintptr_t VAFileStart;
friend class SyscallHandler;
protected:
@@ -79,8 +81,8 @@ namespace FEXCore::HLE {
virtual FEXCore::CodeLoader *GetCodeLoader() const { return nullptr; }
virtual void MarkGuestExecutableRange(uint64_t Start, uint64_t Length) { }
virtual AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(uint64_t GuestAddr) = 0;
virtual std::shared_lock<std::shared_mutex> CompileCodeLock(uint64_t Start) = 0;
virtual SourcecodeResolver *GetSourcecodeResolver() { return nullptr; }
protected:
SyscallOSABI OSABI;
};
+8
View File
@@ -16,6 +16,7 @@
#include <fmt/format.h>
namespace FEXCore::IR {
class OrderedNode;
class RegisterAllocationPass;
class RegisterAllocationData;
@@ -414,6 +415,10 @@ struct SHA256Sum final {
[[nodiscard]] bool operator<(SHA256Sum const &rhs) const {
return memcmp(data, rhs.data, sizeof(data)) < 0;
}
[[nodiscard]] bool operator==(SHA256Sum const &rhs) const {
return memcmp(data, rhs.data, sizeof(data)) == 0;
}
};
class NodeIterator;
@@ -568,6 +573,9 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
return NodeID(NodeOffset / sizeof(IR::OrderedNode));
}
bool IsFragmentExit(FEXCore::IR::IROps Op);
bool IsBlockExit(FEXCore::IR::IROps Op);
} // namespace FEXCore::IR
template <>
+28 -9
View File
@@ -170,23 +170,42 @@ public:
}
}
void Serialize(std::ostream& stream) {
void Serialize(std::ostream& stream) const {
void *nul = nullptr;
//void *IRDataInternal;
stream.write((char*)&nul, sizeof(nul));
stream.write((const char*)&nul, sizeof(nul));
//void *ListDataInternal;
stream.write((char*)&nul, sizeof(nul));
stream.write((const char*)&nul, sizeof(nul));
//size_t DataSize;
stream.write((char*)&DataSize, sizeof(DataSize));
stream.write((const char*)&DataSize, sizeof(DataSize));
//size_t ListSize;
stream.write((char*)&ListSize, sizeof(ListSize));
stream.write((const char*)&ListSize, sizeof(ListSize));
//uint64_t Flags;
uint64_t WrittenFlags = Flags | FLAG_Shared; //on disk format always has the Shared flag
stream.write((char*)&WrittenFlags, sizeof(WrittenFlags));
uint64_t WrittenFlags = FLAG_Shared; //on disk format always has the Shared flag
stream.write((const char*)&WrittenFlags, sizeof(WrittenFlags));
// inline data
stream.write((char*)GetData(), DataSize);
stream.write((char*)GetListData(), ListSize);
stream.write((const char*)GetData(), DataSize);
stream.write((const char*)GetListData(), ListSize);
}
void Serialize(uint8_t *ptr) const {
void *nul = nullptr;
//void *IRDataInternal;
memcpy(ptr, &nul, sizeof(nul)); ptr += sizeof(nul);
//void *ListDataInternal;
memcpy(ptr, &nul, sizeof(nul)); ptr += sizeof(nul);
//size_t DataSize;
memcpy(ptr, &DataSize, sizeof(DataSize)); ptr += sizeof(DataSize);
//size_t ListSize;
memcpy(ptr, &ListSize, sizeof(ListSize)); ptr += sizeof(ListSize);
//uint64_t Flags;
uint64_t WrittenFlags = FLAG_Shared; //on disk format always has the Shared flag
memcpy(ptr, &WrittenFlags, sizeof(WrittenFlags)); ptr += sizeof(WrittenFlags);
// inline data
memcpy(ptr, (const void*)GetData(), DataSize); ptr += DataSize;
memcpy(ptr, (const void*)GetListData(), ListSize); ptr += ListSize;
}
[[nodiscard]] size_t GetInlineSize() const {
@@ -29,6 +29,8 @@ union PhysicalRegister {
static_assert(sizeof(PhysicalRegister) == 1);
struct RegisterAllocationDataDeleter;
// This class is serialized, can't have any holes in the structure
// otherwise ASAN complains about reading uninitialized memory
class FEX_PACKED RegisterAllocationData {
@@ -47,14 +49,12 @@ class FEX_PACKED RegisterAllocationData {
return sizeof(RegisterAllocationData) + NodeCount * sizeof(Map[0]);
}
RegisterAllocationData* CreateCopy() {
auto copy = (RegisterAllocationData*)FEXCore::Allocator::malloc(Size(MapCount));
memcpy((void*)&copy->Map[0], (void*)&Map[0], MapCount * sizeof(Map[0]));
copy->SpillSlotCount = SpillSlotCount;
copy->MapCount = MapCount;
copy->IsShared = IsShared;
return copy;
}
using UniquePtr = std::unique_ptr<FEXCore::IR::RegisterAllocationData, RegisterAllocationDataDeleter>;
static UniquePtr Create(uint32_t NodeCount);
UniquePtr CreateCopy() const;
void Serialize(std::ostream& stream) const {
stream.write((const char*)&SpillSlotCount, sizeof(SpillSlotCount));
stream.write((const char*)&MapCount, sizeof(MapCount));
@@ -67,11 +67,27 @@ class FEX_PACKED RegisterAllocationData {
};
struct RegisterAllocationDataDeleter {
void operator()(RegisterAllocationData* r) {
void operator()(RegisterAllocationData* r) const {
if (!r->IsShared) {
FEXCore::Allocator::free(r);
}
}
};
inline auto RegisterAllocationData::Create(uint32_t NodeCount) -> UniquePtr {
auto Ret = (RegisterAllocationData*)FEXCore::Allocator::malloc(Size(NodeCount));
memset(&Ret->Map[0], PhysicalRegister::Invalid().Raw, NodeCount);
Ret->MapCount = NodeCount;
return UniquePtr { Ret };
}
inline auto RegisterAllocationData::CreateCopy() const -> UniquePtr {
auto copy = (RegisterAllocationData*)FEXCore::Allocator::malloc(Size(MapCount));
memcpy((void*)&copy->Map[0], (void*)&Map[0], MapCount * sizeof(Map[0]));
copy->SpillSlotCount = SpillSlotCount;
copy->MapCount = MapCount;
copy->IsShared = IsShared;
return UniquePtr { copy };
}
}
+1 -1
+1 -1
+175
View File
@@ -0,0 +1,175 @@
#!/usr/bin/python3
import xxhash
import hashlib
import sys
import os
import shutil
import subprocess
def GetDistroInfo():
DistroName = "Unknown"
DistroVersion = "Unknown"
with open("/etc/lsb-release", 'r') as f:
while True:
Line = f.readline()
if not Line:
break
Split = Line.split("=")
if Split[0] == "DISTRIB_ID":
DistroName = Split[1].lower().rstrip()
if Split[0] == "DISTRIB_RELEASE":
DistroVersion = Split[1].rstrip()
return [DistroName, DistroVersion]
def FindBestImageFit(Distro, links_file):
CurrentFitSize = 0
BestFitDistro = None
BestFitDistroVersion = None
BestFitReadableName = None
BestFitImagePath = None
BestFitHash = None
with open(links_file, 'r') as f:
while True:
# Order:
# Distro Name
# Distro Version
# User readable name
# File Path
# Hash
DistroName = f.readline().strip()
if not DistroName:
break
DistroVersion = f.readline().strip()
DistroReadableName = f.readline().strip()
DistroImagePath = f.readline().strip()
DistroHash = f.readline().strip()
FitRate = 0
if (DistroName == Distro[0] or
DistroName == None):
FitRate += 1
if (DistroVersion == Distro[1] or
DistroVersion == None):
FitRate += 1
if FitRate > CurrentFitSize:
CurrentFitSize = FitRate
BestFitDistro = DistroName
BestFitDistroVersion = DistroVersion
BestFitReadableName = DistroReadableName
BestFitImagePath = DistroImagePath
BestFitHash = DistroHash
return [BestFitDistro, BestFitDistroVersion, BestFitReadableName, BestFitImagePath, int(BestFitHash, 16)]
def HashFile(file):
# 32MB buffer size
BUFFER_SIZE = 32 * 1024 * 1024
x = xxhash.xxh3_64(seed=0)
b = bytearray(BUFFER_SIZE)
mv = memoryview(b)
with open(file, 'rb') as f:
while n := f.readinto(mv):
x.update(mv[:n])
return int.from_bytes(x.digest(), "big")
def CheckFilesystemForFS(RootFSMountPath, RootFSPath, DistroFit):
# Check if rootfs mount path exists
if (not os.path.exists(RootFSMountPath) or
not os.path.isdir(RootFSMountPath)):
print("RootFS mount path is wrong")
return False
# Check if rootfs path exists
if (not os.path.exists(RootFSPath) or
not os.path.isdir(RootFSPath)):
# Create this directory
os.makedirs(RootFSPath)
# Check if rootfs path exists
if not os.path.isdir(RootFSPath):
print("RootFS path is not a directory")
return False
# Check rootfs folder for image, copy and extract as necessary
MountRootFSImagePath = RootFSMountPath + DistroFit[3]
RootFSImagePath = RootFSPath + "/" + os.path.basename(DistroFit[3])
NeedsExtraction = False
if not os.path.exists(MountRootFSImagePath):
print("Image {} doesn't exist".format(MountRootFSImagePath))
return False
if not os.path.exists(RootFSImagePath):
# Copy over
print("RootFS image doesn't exist. Copying")
shutil.copyfile(MountRootFSImagePath, RootFSImagePath)
NeedsExtraction = True
# Now hash the image
RootFSHash = HashFile(RootFSImagePath)
if RootFSHash != DistroFit[4]:
print("Hash {} did not match {}, copying new image".format(hex(RootFSHash), hex(DistroFit[4])))
shutil.copyfile(MountRootFSImagePath, RootFSImagePath)
NeedsExtraction = True
# Check if the image needs to be extracted
if not os.path.exists(RootFSPath + "/usr"):
NeedsExtraction = True
if NeedsExtraction:
print("Extracting rootfs")
CmdResult = subprocess.call(["unsquashfs", "-f", "-d", RootFSPath, RootFSImagePath])
if CmdResult != 0:
print("Couldn't extract squashfs")
return False
if not os.path.exists(RootFSPath + "/usr"):
print("Couldn't extract squashfs")
return False
print("RootFS successfully checked and extracted")
return True
def main():
if sys.version_info[0] < 3:
logging.critical ("Python 3 or a more recent version is required.")
FEX_ROOTFS_MOUNT = os.getenv("FEX_ROOTFS_MOUNT")
FEX_ROOTFS_PATH = os.getenv("FEX_ROOTFS_PATH")
if FEX_ROOTFS_MOUNT == None:
print("Need FEX_ROOTFS_MOUNT set")
sys.exit(1)
if FEX_ROOTFS_PATH == None:
print("Need FEX_ROOTFS_PATH set")
sys.exit(1)
if shutil.which("unsquashfs") is None:
print("CI system didn't have unsquashfs installed")
sys.exit(1)
Distro = GetDistroInfo()
DistroFit = FindBestImageFit(Distro, FEX_ROOTFS_MOUNT + "/RootFS_links.txt")
if CheckFilesystemForFS(FEX_ROOTFS_MOUNT, FEX_ROOTFS_PATH, DistroFit) == False:
print("Couldn't load filesystem rootfs")
sys.exit(1)
return 0
if __name__ == "__main__":
# execute only if run as a script
sys.exit(main())
+34
View File
@@ -0,0 +1,34 @@
#!/bin/sh
for binfmt in "$@"
do
result=0
if command -v update-binfmts>/dev/null; then
update-binfmts --find "$binfmt" 1>&- 2>&-
if [ $? -eq 0 ]
then
# If we found the binfmt_misc file passed in then error
result=1
fi
fi
if [ $result -eq 0 ]
then
if [ -f "$binfmt" ]; then
# If the binfmt_misc file exists then error
result=1
fi
fi
if [ $result -eq 1 ]
then
echo "==============================================================="
echo "$binfmt binfmt file is installed!"
echo "This conflicts with FEX-Emu's binfmt_misc!"
echo "This will cause issues when running FEX-Emu through binfmt_misc"
echo "Not installing until you uninstall this binfmt_misc file!"
echo "==============================================================="
exit 1
fi
done
exit 0
+56
View File
@@ -0,0 +1,56 @@
#!/usr/bin/python3
import os
import subprocess
import sys
import tempfile
import platform
def ListContainsRequired(Features, RequiredFeatures):
for Req in RequiredFeatures:
if not Req in Features:
return False
return True
def GetCPUFeaturesVersion():
# Also LOR but kernel doesn't expose this
v8_1Mandatory = ["atomics", "asimdrdm", "crc32"]
v8_2Mandatory = v8_1Mandatory + ["dcpop"]
v8_3Mandatory = v8_2Mandatory + ["fcma", "jscvt", "lrcpc", "paca", "pacg"]
v8_4Mandatory = v8_3Mandatory + ["asimddp", "flagm", "ilrcpc", "uscat"]
# fphp asimdhp asimddp
File = open("/proc/cpuinfo", "r")
Lines = File.readlines()
File.close()
# Minimum spec is ARMv8.0
_ArchVersion = "8.0"
for Line in Lines:
if "Features" in Line:
Features = Line.split(":")[1].strip().split(" ")
# We don't care beyond 8.4 right now
if ListContainsRequired(Features, v8_4Mandatory):
_ArchVersion = "8.4"
elif ListContainsRequired(Features, v8_3Mandatory):
_ArchVersion = "8.3"
elif ListContainsRequired(Features, v8_2Mandatory):
_ArchVersion = "8.2"
elif ListContainsRequired(Features, v8_1Mandatory):
_ArchVersion = "8.1"
break;
return _ArchVersion
def main():
if (platform.machine() == "aarch64"):
print("ARMv{}".format(GetCPUFeaturesVersion()))
elif (platform.machine() == "x86_64"):
print("x64")
sys.exit(0)
if __name__ == "__main__":
sys.exit(main())
+2 -2
View File
@@ -306,12 +306,12 @@ def GetRootFSPath():
return _RootFSPath
def CheckRootFSInstallStatus():
# Matches what is available on https://rootfs.fex-emu.org/file/fex-rootfs/RootFS_links.txt
# Matches what is available on https://rootfs.fex-emu.org/file/fex-rootfs/RootFS_links_XXH3.txt
UbuntuVersionToRootFS = {
"20.04": "Ubuntu_21_04.sqsh",
"21.04": "Ubuntu_21_04.sqsh",
"21.10": "Ubuntu_21_10.sqsh",
"22.04": "Ubuntu_21_10.sqsh",
"22.04": "Ubuntu_22_04.sqsh",
}
return os.path.exists(GetRootFSPath() + UbuntuVersionToRootFS[GetDistro()[1]])
+1 -2
View File
@@ -9,5 +9,4 @@ echo
# These don't have useful documentation at this point
#./Scripts/doc_outline_generator.py "`pwd`" "`pwd`/Scripts" "../"
#./Scripts/doc_outline_generator.py "`pwd`" "`pwd`/Source/Common" "../"
#./Scripts/doc_outline_generator.py "`pwd`" "`pwd`/Source/CommonCore" "../"
#./Scripts/doc_outline_generator.py "`pwd`" "`pwd`/Source/Common" "../"
+3 -2
View File
@@ -1,7 +1,8 @@
#!/bin/env bash
PREVIOUS=FEX-$(date --date='-1 month' +%y%m)
CURRENT=FEX-$(date +%y%m)
# Allow release maintainer to override PREVIOUS and CURRENT by setting it before launching the script
PREVIOUS=${PREVIOUS:-FEX-$(date --date='-1 month' +%y%m)}
CURRENT=${CURRENT:-FEX-$(date +%y%m)}
if ! git rev-list $PREVIOUS > /dev/null 2>&1 ;
then
+9 -7
View File
@@ -14,7 +14,8 @@ known_failures_file = sys.argv[1]
expected_output_file = sys.argv[2]
disabled_tests_file = sys.argv[3]
test_name = sys.argv[4]
fexecutable = sys.argv[5]
mode = sys.argv[5]
fexecutable = sys.argv[6]
known_failures = { }
expected_output = { }
@@ -46,14 +47,15 @@ RunnerArgs = []
RunnerArgs.append(fexecutable)
ROOTFS_ENV = os.getenv("ROOTFS")
if ROOTFS_ENV != None:
RunnerArgs.append("-R")
RunnerArgs.append(ROOTFS_ENV)
if (mode == "guest"):
ROOTFS_ENV = os.getenv("ROOTFS")
if ROOTFS_ENV != None:
RunnerArgs.append("-R")
RunnerArgs.append(ROOTFS_ENV)
# Add the rest of the arguments
for i in range(len(sys.argv) - 6):
RunnerArgs.append(sys.argv[6 + i])
for i in range(len(sys.argv) - 7):
RunnerArgs.append(sys.argv[7 + i])
#print(RunnerArgs)
-1
View File
@@ -1,5 +1,4 @@
add_subdirectory(Common/)
add_subdirectory(CommonCore/)
add_subdirectory(Linux/)
add_subdirectory(Tests/)
add_subdirectory(Tools/)
+2 -3
View File
@@ -3,10 +3,9 @@ set(SRCS
ArgumentLoader.cpp
Config.cpp
EnvironmentLoader.cpp
FEXServerClient.cpp
FileFormatCheck.cpp
RootFSSetup.cpp
StringUtil.cpp
SocketLogging.cpp)
StringUtil.cpp)
add_library(${NAME} STATIC ${SRCS})
target_link_libraries(${NAME} FEXCore_Base cpp-optparse json-maker)
+48 -11
View File
@@ -24,20 +24,20 @@ namespace FEX::Config {
Dest = json_objOpen(Buffer, nullptr);
Dest = json_objOpen(Dest, "Config");
for (auto &it : Layer->GetOptionMap()) {
auto &Name = ConfigToNameLookup.find(it.first)->second;
for (auto &var : it.second) {
Dest = json_str(Dest, Name.c_str(), var.c_str());
}
auto &Name = ConfigToNameLookup.find(it.first)->second;
for (auto &var : it.second) {
Dest = json_str(Dest, Name.c_str(), var.c_str());
}
}
Dest = json_objClose(Dest);
Dest = json_objClose(Dest);
json_end(Dest);
std::ofstream Output (Filename, std::ios::out | std::ios::binary);
if (Output.is_open()) {
Output.write(Buffer, strlen(Buffer));
Output.close();
}
std::ofstream Output (Filename, std::ios::out | std::ios::binary);
if (Output.is_open()) {
Output.write(Buffer, strlen(Buffer));
Output.close();
}
}
std::string LoadConfig(
@@ -69,8 +69,45 @@ namespace FEX::Config {
std::string Program = Args[0];
// These layers load on initialization
auto ProgramName = std::filesystem::path(Program).filename();
bool Wine = false;
std::filesystem::path ProgramName;
for (size_t CurrentProgramNameIndex = 0; CurrentProgramNameIndex < Args.size(); ++CurrentProgramNameIndex) {
auto CurrentProgramName = std::filesystem::path(Args[CurrentProgramNameIndex]).filename();
if (CurrentProgramName == "wine-preloader" ||
CurrentProgramName == "wine64-preloader") {
// Wine preloader is required to be in the format of `wine-preloader <wine executable>`
// The preloader doesn't execve the executable, instead maps it directly itself
// Skip the next argument since we know it is wine (potentially with custom wine executable name)
++CurrentProgramNameIndex;
Wine = true;
}
else if(CurrentProgramName == "wine" ||
CurrentProgramName == "wine64") {
// Next argument, this isn't the program we want
//
// If we are running wine or wine64 then we should check the next argument for the application name instead.
// wine will change the active program name with `setprogname` or `prctl(PR_SET_NAME`.
// Since FEX needs this data far earlier than libraries we need a different check.
Wine = true;
}
else {
if (Wine == true) {
// If this was path separated with '\' then we need to check that.
auto WinSeparator = CurrentProgramName.string().find_last_of('\\');
if (WinSeparator != CurrentProgramName.string().npos) {
// Used windows separators
CurrentProgramName = CurrentProgramName.string().substr(WinSeparator + 1);
}
}
ProgramName = CurrentProgramName;
// Past any wine program names
break;
}
}
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, true));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, false));
return Program;
+25
View File
@@ -0,0 +1,25 @@
#pragma once
#include <FEXCore/Utils/LogManager.h>
#include <fcntl.h>
#include <filesystem>
#include <linux/limits.h>
#include <unistd.h>
namespace FEX {
[[maybe_unused]]
static
std::string get_fdpath(int fd) {
char SymlinkPath[PATH_MAX];
std::filesystem::path Path = std::filesystem::path("/proc/self/fd") / std::to_string(fd);
int Result = readlinkat(AT_FDCWD, Path.c_str(), SymlinkPath, sizeof(SymlinkPath));
if (Result != -1) {
return std::string(SymlinkPath, Result);
}
LOGMAN_MSG_A_FMT("Couldn't get symlink from /proc/self/fd/{}", fd);
return {};
}
}
Loaded 100 of 235 files, more files were not shown because too many files have changed in this diff. Show more