Compare commits

...
233 Commits
Author SHA1 Message Date
Ryan Houdek 33fe6813fc Docs: Update for release FEX-2104 2021-04-02 11:35:29 -07:00
Ryan Houdek c6c2555e1b Merge pull request #938 from Sonicadvance1/FEXConfig_improvements
FEXConfig improvements
2021-04-01 23:55:16 -07:00
Stefanos Kornilios Mitsis Poiitidis 283a5c1293 Merge pull request #937 from Sonicadvance1/fix_inverted_bool
Fixes inverted boolean FEXLoader arguments
2021-04-02 09:38:21 +03:00
Stefanos Kornilios Mitsis Poiitidis b31976d0e2 Merge pull request #936 from Sonicadvance1/fix_cp_config_error
Fixes Thunk Guest libs short arg. C&P error made it not be j
2021-04-02 09:36:53 +03:00
Ryan Houdek f922e6420b Allow opening the first configuration file passed in
Lets us easily open any configuration file
2021-04-01 19:25:17 -07:00
Ryan Houdek 7ff06c828f Loads default config if doesn't exist
Allows us to save a folder with the default if we want it
2021-04-01 19:22:53 -07:00
Ryan Houdek be16cc1456 Fixes window ordering so popup message is visible without config open
This was a child of the config popup before.
Moves it to a child of the workspace instead so that if we had a load error the message is still visible
2021-04-01 19:16:03 -07:00
Ryan Houdek 59df3d13cc Don't set up RootFS folder or inotify thread on config load failure
Can happen in the case that the default file tries to get loaded and doens't exist
2021-04-01 19:15:08 -07:00
Ryan Houdek 1adf518e6f Don't try creating inotify thread twice
If the inotify FD is already opened then just early exit
Can happen in the case that a config tries to get opened and didn't exist
2021-04-01 19:13:43 -07:00
Ryan Houdek 68171ad9b5 Creates RootFS data folder as a user convenience
To make sure users don't mess this up, create the folder if it isn't found
2021-04-01 19:12:56 -07:00
Ryan Houdek 78a923b5b0 Fixes inverted boolean FEXLoader arguments
The default_value doesn't have anything to do with the boolean argument being set.
Setting a boolean argument will always set true, inverted always sets false.
We don't currently have a need to support the inverted case where passing in an argument sets it to false
2021-04-01 19:10:29 -07:00
Ryan Houdek 0c6741438b Update config generator python to check for duplicate options
Checks for duplicate short and long option arguments
2021-04-01 19:08:19 -07:00
Ryan Houdek 5d0c080a9d Fixes Thunk Guest libs short arg. C&P error made it not be j 2021-04-01 19:08:00 -07:00
Ryan Houdek 5f49b57948 Merge pull request #928 from Sonicadvance1/pthreads_threading
Switches FEXCore over to pthreads implementation
2021-04-01 01:15:43 -07:00
Ryan Houdek 38778953a1 Merge pull request #924 from Sonicadvance1/select_named_rootfs
Adds support for Named RootFS folders in FEXConfig
2021-04-01 00:32:15 -07:00
Ryan Houdek 21a364c2ad Switches FEXCore over to pthreads implementation
This defaults to a pthread implementation but it can be switched over to a custom thread
handler if the frontend desires.
2021-04-01 00:31:45 -07:00
Ryan Houdek 3294cc209f Adds support for Named RootFS folders in FEXConfig
Provides a drop down dialog of RootFS folders to select.
Will be empty if the user doesn't have any folders in place.
Just allowing directory paths to be inserted

Fixes #689
2021-04-01 00:15:29 -07:00
Stefanos Kornilios Mitsis Poiitidis ae5d41e46b Merge pull request #926 from Sonicadvance1/disable_seccomp
Disables seccomp
2021-04-01 10:03:36 +03:00
Ryan Houdek 43f83fd6d2 Merge pull request #923 from Sonicadvance1/named_rootfs
Adds support for named rootfs configurations
2021-04-01 00:02:11 -07:00
Ryan Houdek 8cf47f8221 Adds support for named rootfs configurations
If the relative or absolute folder doesn't exist then FEX will search in the data folder for a rootFS named the same thing
If the path exists then it is used.

For example if my data folder contains `$HOME/.fex-emu/RootFS/Ubuntu_main/` and I set the RootFS option to `Ubuntu_main`
Then this rootfs will be used. Allows you to easily select a rootfs in that folder without having the full file path.
2021-03-31 23:46:07 -07:00
Ryan Houdek 43b6cc1cb9 Adds GetDataDirectory config helper
This ends up in either $HOME/.fex-emu/ or $XDG_DATA_HOME/.fex-emu/
2021-03-31 23:46:07 -07:00
Ryan Houdek 1ab6726498 Moves Config directory finding to FEXCore
This will be necessary in the next commit
2021-03-31 23:46:07 -07:00
Ryan Houdek c21acd0d45 Merge pull request #916 from Sonicadvance1/fix_execve_again
Fix execve again
2021-03-31 23:45:06 -07:00
Ryan Houdek e637751112 ptrace behaviour is now changed 2021-03-31 23:38:14 -07:00
Ryan Houdek fac6377ed3 Fixes a crash in openat 2021-03-31 23:37:57 -07:00
Ryan Houdek 4f3cb933cb Fixes execve 2021-03-31 23:37:56 -07:00
Ryan Houdek b2825bd848 Adds handler in ELFLoader to determine ELF type. 2021-03-31 23:37:56 -07:00
Ryan Houdek 83eb7f4ad3 Return EPERM in ptrace. A Feral launcher is attempting to use this to ensure it can't ptrace. 2021-03-31 23:37:56 -07:00
Ryan Houdek c3854c211a Fix finding home directory when we have no environment variables. 2021-03-31 23:37:56 -07:00
Ryan Houdek 61cd3eb3ce Fix crash in ELFDB if it tried loading something that wasn't an ELF 2021-03-31 23:37:56 -07:00
Ryan Houdek 490352f568 Merge pull request #921 from Sonicadvance1/update_imgui
Updates external imgui which fixes keypad enter in FEXConfig
2021-03-31 23:24:42 -07:00
Ryan Houdek 4004d5a3b7 Disables seccomp
FEX doesn't support seccomp in userspace and allowing these through causes chromium secure sandbox to break.
Disabling these with EINVAL allows FEX to behave as if seccomp isn't enabled in the kernel config

Necessary to get the Civ6 launcher further
2021-03-31 23:23:40 -07:00
Ryan Houdek 761447467e Updates external imgui which fixes keypad enter in FEXConfig 2021-03-30 15:39:11 -07:00
Ryan Houdek 3f38ed94ca Merge pull request #922 from Sonicadvance1/disable_oomscore
Disables gvisor test proc_pid_oomscore_test
2021-03-30 15:38:33 -07:00
Ryan Houdek ad877d4088 Disables gvisor test proc_pid_oomscore_test
Depending on runner this passes or fails.
The test expects to be able to open `/proc/self/oom_score_adj` as writable to adjust the oom score.
This is expected to fail on all three of our runners, but sometimes the solidrun board manages to open it.
Disable as it is a flake. Could be a kernel bug on the solid run board
2021-03-30 15:29:52 -07:00
Ryan Houdek 92fc6b1909 Merge pull request #920 from lioncash/insert
Config: Make lookup map assignment match behavior of comment
2021-03-30 14:17:37 -07:00
Lioncache 4b11dcb72a Config: Make lookup map assignment match behavior of comment
emplace() only performs an insertion or assignment if the key doesn't
already exist within the map, which is at odds with what the comment
above the line indicates should happen.

insert_or_assign() better models what the comment indicates.
2021-03-30 11:51:52 -04:00
Lioncache dd7c78dcd7 Config: Perform lookups by character overloads where applicable
These are marginally better to perform than the string equivalents
2021-03-30 11:41:57 -04:00
Lioncache a20ef403c1 Config: std::move strings within MergeEnvironmentVariables()
Same behavior, minus potential heap duplicate allocations
2021-03-30 11:40:42 -04:00
Ryan Houdek be06511d3d Merge pull request #919 from FEX-Emu/skmp/fix-build
Build: Fix for ubuntu 20.04
2021-03-30 03:20:43 -07:00
Stefanos Kornilios Mitsis Poiitidis b7a60337e8 Build: Fix for ubuntu 20.04 2021-03-30 13:08:19 +03:00
Stefanos Kornilios Mitsis Poiitidis 315b8e9c1e Merge pull request #917 from FEX-Emu/skmp/lightweight-docs-3
Docs: Add tags, unittest Readme.md, generated SourceOutline.md
2021-03-30 12:52:18 +03:00
Ryan Houdek 155a2b0194 Merge pull request #918 from lioncash/fallthrough
CMakeLists: Flag unannotated implicit fallthrough as an error
2021-03-30 02:37:07 -07:00
Stefanos Kornilios Mitsis Poiitidis 8c874f4540 Docs: Commit generated SourceOutline.md, update Readme.md 2021-03-30 12:21:18 +03:00
Stefanos Kornilios Mitsis Poiitidis 942b8549c6 Docs: Add tags to the source code 2021-03-30 12:21:18 +03:00
Ryan Houdek fa577b4527 Merge pull request #910 from FEX-Emu/skmp/lightweight-docs
Docs: lightweight doc autogen
2021-03-30 02:17:41 -07:00
Stefanos Kornilios Mitsis Poiitidis 15537f8c8c Scripts: Also add unittests folder for doc outline 2021-03-30 11:58:18 +03:00
Lioncache 434a74a88a CMakeLists: Flag unannotated implicit fallthrough as an error
Prevents a class of sneaky logic bugs from slipping through into the
codebase.

This also resolves a case of such a bug within the Decorder's ReadData()
where all 3 byte reads would be performed as if they were a 4 byte read.
2021-03-30 04:56:12 -04:00
Stefanos Kornilios Mitsis Poiitidis 1fdd6fbb14 Scripts: Add a documentation comment in changelog_generator.py 2021-03-30 11:55:39 +03:00
Stefanos Kornilios Mitsis Poiitidis 458bebf598 Scripts: Add a documentation comment in doc_outline_generator.py 2021-03-30 11:53:11 +03:00
Stefanos Kornilios Mitsis Poiitidis 2b10b9792b Scripts: Fix generate_release, don't use markdown for changelogs 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 0dd02da57c Docs: Update outline scripts 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 309139e203 Scripts: Add generate_release script 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 350c33ada4 Docs: Update outline script 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis bdd35e5743 docs: Update outline generator 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis ceb7082e37 Docs/Versioning: Add Changelog Generator 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 2d6dc80039 Docs: Adds tag-based outline generator w/ glossary support 2021-03-30 11:22:56 +03:00
Ryan Houdek c92df21627 Merge pull request #915 from lioncash/fallthrough
OpcodeDispatcher: Add missing break for UD2 in INTOp()
2021-03-29 22:56:19 -07:00
Lioncache 539de0492e OpcodeDispatcher: Add missing break for UD2 in INTOp()
Fixes a case where the reason gets overwritten.
2021-03-29 23:42:26 -04:00
Ryan Houdek efd41de8ea Merge pull request #914 from lioncash/optional
Context: Make use of std::optional with GetFilenameHash()
2021-03-29 16:44:08 -07:00
Lioncache 99ae862875 Context: Eliminate undefined behavior in LoadEntryList()
This breaks the strict-aliasing rule. We can use std::memcpy here to
make this code well-defined.
2021-03-29 12:17:36 -04:00
Lioncache b81ea43601 Context: Make use of std::optional with GetFilenameHash()
Same behavior, minus the need for an out parameter. While we're at it,
we can tidy up some of the file handling code.
2021-03-29 12:17:32 -04:00
Ryan Houdek fcd4974cff Merge pull request #912 from Sonicadvance1/config_docs_improvements
Configuration option improvements
2021-03-29 04:47:45 -07:00
Stefanos Kornilios Mitsis Poiitidis 8550c9f2dd Merge pull request #913 from Sonicadvance1/fix_32bit_mask
Adds a 32-bit mask for multiblock RIP calculation
2021-03-29 14:35:32 +03:00
Ryan Houdek 15dd70487f Removes legacy difference between FEX config env variable and enum names 2021-03-29 00:26:11 -07:00
Ryan Houdek 8fad0c8fdb Adds a 32-bit mask for multiblock RIP calculation
In the case of overflow then it'll just mask this result.
The rest of the logic already matches what 32-bit expects.

Fixes #703
2021-03-28 23:39:53 -07:00
Ryan Houdek b6c0a9f2a0 Configuration option improvements
This commit does three things that are bounded to each other
1) Moves configs from ConfigValues.inl to Config.json
2) Uses the json to generate a man file with the options inside of it
3) Generates the code required for FEXLoader to automatically parse defined options

Moving the configuration options to a parseable format was required to generate the man pages.
Can't really include an inl file in the man page
Man page gets generated and installed through the regular cmake install process
`man FEX` to get the man page

With this change, FEXLoader's argument parsing now will automatically be generated from the json.
This way whatever is in the json file matches what is in the man page, in the json, AND what is returned in FEXLoader --help.
It will stay in sync now.

FEXConfig is the only application that stays out of sync for now as it requires some more thought to plan out.

A minor improvement that this brings as well is that every boolean option that has a long argument also gains the inversion of that property.
aka, `--gdb` also gains `--no-gdb` The use case for this is minor but boolean arguments should always allow negated variants.
2021-03-28 23:28:41 -07:00
Ryan Houdek b5a6dd031b Merge pull request #911 from Sonicadvance1/std_size_life
Replaces FEX's usage of classic C array element count with std::size
2021-03-28 05:58:55 -07:00
Ryan Houdek 650877adca Replaces FEX's usage of classic C array element count with std::size
Just learned that this is a c++17 feature and it's a nice little
cleanup.

Only one remains in our codebase which can't easily be replaced
2021-03-28 05:39:35 -07:00
Ryan Houdek 432b6fccf2 Merge pull request #907 from Sonicadvance1/replace_glfw
Replaces FEXConfig usage of GLFW with SDL2
2021-03-27 05:53:00 -07:00
Ryan Houdek 47f3c983ee Merge pull request #905 from Sonicadvance1/removes_warnings_aarch64
Removes remaining warnings from AArch64 JIT.
2021-03-27 05:52:51 -07:00
Ryan Houdek 9565247ada Merge pull request #904 from Sonicadvance1/stop_compiling_twice
Stop compiling FEXCore twice
2021-03-27 05:52:39 -07:00
Ryan Houdek c59b9fbede Merge pull request #903 from Sonicadvance1/fix_silent_logs
Disables silent logging on unit tests
2021-03-27 05:52:28 -07:00
Ryan Houdek 7c96893579 Updates README with SDL2 requirement 2021-03-27 05:26:00 -07:00
Ryan Houdek 8ebef457fc Replaces FEXConfig usage of GLFW with SDL2
Sometimes GLFW3 can just consume a full CPU core, spinning on nothing.

Fixes #906
2021-03-27 05:23:58 -07:00
Ryan Houdek 9eee879be0 Removes remaining warnings from AArch64 JIT.
There are a couple of warnings remaining but need restructuring to solve.
I believe @phire is going to fix the one warning about a virtual destructor in a upcoming PR.
Fixes #679
2021-03-26 18:49:40 -07:00
Ryan Houdek 64affa8c8e Stop compiling FEXCore twice
This now creates an object binary then static and shared libraries from that.
This stops us outputting warnings twice in FEXCore.
Also improves compilation time even with ccache enabled.
Went from 12s to 8s compilation WITH ccache.
I'm sure without ccache it is even better.
2021-03-26 18:27:23 -07:00
Ryan Houdek dc041bdf0e Disables silent logging on unit tests
We need these for our CI artifacts
2021-03-26 18:04:17 -07:00
Ryan Houdek 51c43a0761 Adds option to disable silent logging 2021-03-26 18:00:23 -07:00
Ryan Houdek aa67142e6f Merge pull request #902 from Azkali/patch-1
Dockerfile add arm64 compatibility
2021-03-25 22:36:41 -07:00
The Great Wizard Azkali ae569da895 Update Dockerfile 2021-03-26 06:30:17 +01:00
The Great Wizard Azkali d77acdf474 Dockerfile add arm64 compatibility
Python3, Python3-dev and linux-headers prevent from building on arm64
2021-03-26 02:18:13 +01:00
Ryan Houdek 264ec44276 Merge pull request #892 from Sonicadvance1/install_thunks
Allows installing of FEXThunks in our data directory
2021-03-25 14:50:54 -07:00
Ryan Houdek da49eb3394 Allows installing of FEXThunks in our data directory
This is necessary for building FEX packages that contain some initial thunk libs.
Gives an initial foothold for a default location for the host and guest thunk folders
2021-03-24 03:27:38 -07:00
Ryan Houdek 5447ec3ec8 Merge pull request #884 from Sonicadvance1/fix_empty_string_crash
Fixes crash on empty path config string
2021-03-24 02:34:20 -07:00
Ryan Houdek ff3974c4a1 Merge pull request #890 from Sonicadvance1/merged_environment_variables
Fixes environment variable configuration from multiple layers
2021-03-24 02:34:14 -07:00
Ryan Houdek 358f12e074 Merge pull request #891 from Sonicadvance1/installed_global_app_configs
Implements support for installing global application profiles
2021-03-24 02:34:00 -07:00
Ryan Houdek c5648ac84c Implements support for installing global application profiles
These need to be used sparingly, we don't want to proactively crush user options.
2021-03-23 22:34:01 -07:00
Ryan Houdek 30e1f871ad Fixes environment variable configuration from multiple layers
This was a missing feature that I had skipped previously.
Before this commit, each layer would overwrite all previous environment variables if they had anything defined.

After this commit, the layers will now merge in priority order.
Meaning if a higher priority layer has the same environment variable defined, it will overwrite that specific variable
Ending up with a superset of all the layer's environment variables now.
2021-03-23 22:01:09 -07:00
Ryan Houdek 5c624684b7 Merge pull request #887 from Sonicadvance1/FEXBashLoader
Adds a FEXBash helper program
2021-03-23 21:37:45 -07:00
Ryan Houdek ef11b534ef Adds a FEXBash helper program
This is a helper program to execute a bash command through FEX.
Works around an edge case of
eg:
FEXInterpreter /bin/sh -c "echo A"
versus
FEXBash "echo A"
Argument expansion ends up being a pain point
This isn't currently used but will be very shortly
2021-03-23 21:27:30 -07:00
Ryan Houdek cd32eeb428 Adds FEXInterpreter as a build target
Makes it easier to have FEXInterpreter around
2021-03-23 21:27:29 -07:00
Ryan Houdek 43984c810d Merge pull request #886 from Sonicadvance1/namespace_perm
Returns EPERM on clone with namespace
2021-03-23 20:56:20 -07:00
Ryan Houdek a50689aa8b Merge pull request #885 from Sonicadvance1/cpuid_leaf
Adds support for CPUID leafs
2021-03-23 20:55:33 -07:00
Scott Mansell 311d6d4385 Merge pull request #883 from phire/thread_cleanup
Cleanup threads when they exit
2021-03-24 16:22:15 +13:00
Ryan Houdek a19c59f5b2 Merge pull request #888 from Sonicadvance1/default_silentlogs
Default to silent logging
2021-03-23 20:13:07 -07:00
Scott Mansell 6a8f68022f Join threads after fork.
Also, add some comments to document things and
rename function to be clear about it's purpose.
2021-03-24 15:41:09 +13:00
Ryan Houdek fdcb6206da Default to silent logging
Useful for debugging but we need to be silent by default now.
Without this anything that checks stdout output of applications from bash would fail

Steam with lspci, lsusb, uname, etc
2021-03-23 19:30:42 -07:00
Ryan Houdek 6ea977ba34 Remove proc_pid_uid_gid_map from known failures list
Now instead of crashing with namespacing, it skips the tests if it gets EPERM.
2021-03-23 19:28:12 -07:00
Ryan Houdek 9b35cb4408 Returns EPERM on clone with namespace
Notably this allows applications to work that don't require the namespace but check up front if they are able to clone with it.
Civ 6's launcher checks this as an example
2021-03-23 19:28:06 -07:00
Ryan Houdek afa869e1f2 Adds Config.h generated file
Gives us the install path and FEXInterpreter locations
2021-03-23 19:12:18 -07:00
Ryan Houdek 43ddba4e82 Adds support for CPUID leafs
Leafs come from ECX but only some CPUID functions support this.
This adds the initial infrastructure but doesn't yet add support for the CPUID functions to consume the leaf.
2021-03-23 18:13:14 -07:00
Ryan Houdek e3380957ba Fixes crash on empty path config string
If the config exists as an empty string then this would crash
2021-03-23 18:02:56 -07:00
Scott Mansell 21687f4dc5 Also destroy the parent thread 2021-03-24 03:14:23 +13:00
Scott Mansell 41a1a9d400 Cleanup threads when they exit 2021-03-24 03:14:23 +13:00
Stefanos Kornilios Mitsis Poiitidis 82c2b49a2e Merge pull request #878 from Sonicadvance1/fexconfig_actual_default
Changes FEXConfig to actually load default configuration
2021-03-22 18:24:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fbc49c1648 Merge pull request #876 from Sonicadvance1/QOL_config
Implements a couple quality of life configuration option handling
2021-03-22 18:23:52 +02:00
Stefanos Kornilios Mitsis Poiitidis d2f636893d Merge pull request #880 from Sonicadvance1/github_label_testing
Support disabling unit tests based on runner label
2021-03-22 18:18:08 +02:00
Ryan Houdek e5a9bd4d3e Fixes incorrect result on cmpxchg reg, reg failure when regsize = 32bit 2021-03-22 09:03:48 -07:00
Ryan Houdek ca24ea1d36 Fixes missing HandledLock on XCHG 2021-03-22 09:03:48 -07:00
Ryan Houdek 998e53aa70 Only disabled unaligned atomics test on ARMv8.0 2021-03-22 09:03:48 -07:00
Ryan Houdek de890e7387 Have unit tests check for runner label 2021-03-22 09:03:48 -07:00
Ryan Houdek 042a71be96 Get the runner's label for testing disabled tests later 2021-03-22 09:03:48 -07:00
Ryan Houdek 0a61741596 special case logging option stdout/stderr 2021-03-21 02:26:56 -07:00
Ryan Houdek 91a5a625ac Implements a couple quality of life configuration option handling
This takes the changes from #829 and moves it to the correct location to be picked up from any loader.
This also fixes #873.
Expands the config paths. Anything that has ~ or is relative will be converted to an absolute path.
2021-03-21 02:25:30 -07:00
Stefanos Kornilios Mitsis Poiitidis f68593ea97 Merge pull request #877 from Sonicadvance1/fix_absolute_linker
Adds some additional logic for finding absolute soft linked linker
2021-03-21 11:08:46 +02:00
Stefanos Kornilios Mitsis Poiitidis e15730782b Merge pull request #879 from Sonicadvance1/update_readme
Updates readme with some more direct information
2021-03-21 11:07:35 +02:00
Ryan Houdek fc27893f20 Updates readme with some more direct information
Some of this taken from Stef's recent changes. Some taken from the wiki.
combination of the two best bits
2021-03-21 01:53:51 -07:00
Ryan Houdek 5b8bf8dd24 Adds some additional logic for finding absolute soft linked linker
Ubuntu's x86_64 rootfs does a soft link from `/lib64/ld-linux-x86-64.so.2` to an absolute path of `/lib/x86_64-linux-gnu/ld-2.32.so`
Our frontend ELF loader logic would resolve this symlink to be just `/lib/x86_64-linux-gnu/ld-2.32.so` which would then fail.
Adds some additional logic to our frontend that if the symlink ends up being a softlink to an absolute address then it'll first resolve that symlink.
Then combines rootfs + absolute path. If that still fails then it'll fall down the regular linker path, which could still find the linker on the host side.

This is all a bit of a kludge to just work around Ubuntu doing an absolute address rather than a relative one.

Fixes #869
2021-03-21 01:11:47 -07:00
Ryan Houdek 0c48f1409e Changes FEXConfig to actually load default configuration
This was previously just some default values that I put in as placeholder
Now that we actually have defaults configured somewhere, use those.

Fixed #874
2021-03-21 00:48:51 -07:00
Ryan Houdek 8c1f2eb0b3 Merge pull request #870 from FEX-Emu/skmp/lock-validation
OpDisp: Validate LOCK handling, add missing segment offsets
2021-03-20 15:53:06 -07:00
Ryan Houdek 07613456d2 Merge pull request #872 from FEX-Emu/skmp/x87-fixes
X87: Init on X87FNSAVE, fix FNINIT
2021-03-20 15:52:54 -07:00
Ryan Houdek 493bb3bef3 Merge pull request #875 from phire/test-no-multibock
Explictly as for --no-multiblock in tests
2021-03-20 15:52:44 -07:00
Scott Mansell cf8572e090 Explictly as for --no-multiblock in tests
We switched the default over a while back, so we haven't
been getting test coverage with multiblock off
2021-03-21 04:57:47 +13:00
Stefanos Kornilios Mitsis Poiitidis d120f12c93 X87: Init on X87FNSAVE, fix FNINIT 2021-03-19 17:04:51 +02:00
Stefanos Kornilios Mitsis Poiitidis 695322d446 OpDisp: Validate LOCK handling, add missing segment offsets 2021-03-19 16:32:29 +02:00
Stefanos Kornilios Mitsis Poiitidis d34cde12ae Merge pull request #850 from Sonicadvance1/struct_verifier
libclang based Struct verifier written in python
2021-03-19 10:06:34 +02:00
Stefanos Kornilios Mitsis Poiitidis a7dc9d00d2 Merge pull request #864 from Sonicadvance1/cpuid_15h
Implements CPUID 15h
2021-03-19 10:02:02 +02:00
Ryan Houdek dbde2e400e Merge pull request #863 from Sonicadvance1/enable_invariant_tsc
Enables Invariant TSC CPUID bit
2021-03-18 12:39:08 -07:00
Ryan Houdek bdeefa7bf0 Merge pull request #861 from Sonicadvance1/add_version_config
Implements a --version argument
2021-03-18 12:39:00 -07:00
Ryan Houdek 51e513df14 Merge pull request #865 from FEX-Emu/skmp/fix-emufiles-lock
EmulatedFiles: Lock around FDToNameMap accesses
2021-03-18 04:11:58 -07:00
Stefanos Kornilios Mitsis Poiitidis 0106b362d6 EmulatedFiles: Lock around FDToNameMap accesses 2021-03-18 12:45:27 +02:00
Ryan Houdek c7f265158f Implements CPUID 15h
ARMv8 allows you query the cycle counter register very from userspace which allows us to emulate this easily.

My IceLake device returns 1.5Ghz through this interface
My AMD Zen+ device doesn't support this but with measurements has 3Ghz TSC
My Snapdragon 865 device has a TSC frequency of 19.20Mhz
2021-03-18 01:42:44 -07:00
Ryan Houdek fbf5325bdd Enables Invariant TSC CPUID bit
I keep forgetting to enable this bit. Fixes #855
2021-03-18 00:15:48 -07:00
Ryan Houdek 9b2bd8df28 Implements a --version argument
Nicer for the user to determine what their version is
2021-03-17 00:00:57 -07:00
Ryan Houdek 2a074204bc Merge pull request #859 from Sonicadvance1/improve_fexinterpreter_msg
Adds an installer message to the FEXInterpreter install
2021-03-16 23:13:12 -07:00
Ryan Houdek d79fb37135 Merge pull request #858 from Sonicadvance1/remove_testharness
Removes TestHarness
2021-03-16 23:13:03 -07:00
Ryan Houdek d01b40c5aa Adds an installer message to the FEXInterpreter install
Otherwise it doesn't feel like FEXInterpreter is installing
2021-03-16 22:36:58 -07:00
Ryan Houdek e871a43387 Merge pull request #857 from phire/deletething
Remove old and unused HostCore
2021-03-16 22:30:30 -07:00
Ryan Houdek ff45b37904 Merge pull request #856 from Sonicadvance1/remove_libcap
Removes libcap-dev dependency
2021-03-16 22:30:21 -07:00
Ryan Houdek 62aab57ef7 Removes TestHarness
This has been completely deprecated for favour of the TestHarnessRunner instead.
Is unused.
2021-03-16 22:14:49 -07:00
Scott Mansell 208fa0d1fc Remove old and unused HostCore
The thing we are actually using is Source/CommonCore/HostFactory.cpp
2021-03-17 18:09:49 +13:00
Ryan Houdek 634d4fead0 Removes libcap-dev dependency
We are only using this for syscalls getcap and setcap.
These are only passthrough pointers and don't need to be handled from a library.

Noticed this while writing user documentation
2021-03-16 22:08:15 -07:00
Ryan Houdek 0bdddedfe6 Merge pull request #853 from FEX-Emu/skmp/dont-lse-though-syscalls
RCLSE: Invalidate around OP_SYSCALLs, Syscalls might read the context
2021-03-16 17:54:54 -07:00
Stefanos Kornilios Mitsis Poiitidis 971740991b RCLSE: Invalidate around OP_SYSCALLs, Syscalls might read the context and need the latest version of it
Fixes steam w/ all optimization passes enabled
2021-03-17 02:42:32 +02:00
Ryan Houdek cb672e035c Merge pull request #852 from Sonicadvance1/fix_fexconfig_idle
Fixes FEXConfig trying to update full refresh
2021-03-16 00:53:40 -07:00
Ryan Houdek 4f64ba582c Patch up the remaining 32bit structs 2021-03-15 15:51:59 -07:00
Ryan Houdek fbe2583a04 Fix truncating to not set the log files to 21MB 2021-03-15 15:24:49 -07:00
Ryan Houdek befe9dcbae Adds struct verifier to github yaml workflow file
This way CI tests this on each commit
2021-03-15 15:24:49 -07:00
Ryan Houdek aacb0f6891 Adds new struct_verifier ctest to cmake
Currently only testing 32bit syscall struct definitions
2021-03-15 15:24:49 -07:00
Ryan Houdek 040cc746e7 Fixes FEXConfig trying to update full refresh
This burns a decent amount of CPU time just idling and we don't have active screen elements to matter here
It's still responsive because it'll update as the window gets any events
2021-03-15 09:46:49 -07:00
Ryan Houdek e0c5840f2c Adds libclang based struct verifier
This requires multiarch on the targets to work.
Will run a header through multiple architectures and ensure that the struct packing all works
2021-03-15 06:58:06 -07:00
Ryan Houdek 48027f1d2e Moves BUILD_TESTS check up the list
This way Source/ can have its own tests
2021-03-15 06:58:06 -07:00
Ryan Houdek 9dc0717cdb Attributes a few 32bit syscall types
This will be necessary for struct verification in the next commit
rusage needed to be updated to match the real rusage. Has unions with two named types
2021-03-15 06:58:06 -07:00
Ryan Houdek 8de7dd8fd0 Merge pull request #848 from Sonicadvance1/remove_warnings
Remove most warnings from FEX
2021-03-15 06:57:42 -07:00
Ryan Houdek 79e9477b14 Removes warnings from RAPass 2021-03-14 13:27:40 -07:00
Ryan Houdek 0b8f29000e Removes warnings from ConstProp 2021-03-14 13:27:39 -07:00
Ryan Houdek e530e3676f Removes warnings from IRDumper 2021-03-14 13:27:39 -07:00
Ryan Houdek 0c5e5dbdf2 Removes warnings from OpcodeDispatcher 2021-03-14 13:27:39 -07:00
Ryan Houdek d8591a8f14 Removes warnings from x86 JITCore 2021-03-14 13:27:39 -07:00
Ryan Houdek 941107120c Removes warnings from x86 BranchOps 2021-03-14 13:27:39 -07:00
Ryan Houdek ffc17c90c5 Removes warnings from InterpreterOps 2021-03-14 13:27:39 -07:00
Ryan Houdek 5285f4baa6 Adds CMake option to enable -Werror
We aren't error free so can't be default enabled
2021-03-14 13:27:39 -07:00
Ryan Houdek 037781c4d0 Merge pull request #849 from Sonicadvance1/fix_kernel_version
Fixes uname version and wraps /proc/version
2021-03-14 11:09:05 -07:00
Scott Mansell 4d2221a456 Merge pull request #831 from phire/one_dispatcher_to_rule_them_all
Unify all four dispatchers
2021-03-15 05:42:21 +13:00
Scott Mansell 466dc03c19 Remove stray cmake changes 2021-03-15 04:49:09 +13:00
Scott Mansell 5e1e09a43d Fix review issues 2021-03-15 04:41:14 +13:00
Ryan Houdek 3a0b475cfa Wrap /proc/version where we were leaking host kernel information
drops the correct FEX version in to the string as well
2021-03-12 20:29:35 -08:00
Ryan Houdek 4ece56ac4d Fixes uname version string
This wasn't following the correct format and it never actually had a working FEX_VERSION define.
Now it pulls in the GIT_DESCRIBE_STRING and brings in the correct date + time format
2021-03-12 20:29:35 -08:00
Ryan Houdek ea7af248e8 Ensure programs can pull in git_version.h 2021-03-12 20:29:35 -08:00
Scott Mansell 42b186a76c Don't generate a default vixl::codebuffer 2021-03-12 12:47:29 +13:00
Scott Mansell bdca7109a4 Force XBYAK64
Because XBYAK's autoconfig doesn't quite do the right thing
2021-03-11 23:46:07 +13:00
Scott Mansell 074bedf25c Automatically use correct ContextBackup type 2021-03-11 23:46:07 +13:00
Scott Mansell bdb199b6d6 Unify all four dispatchers 2021-03-11 23:46:07 +13:00
Stefanos Kornilios Mitsis Poiitidis e9afabc4cb Merge pull request #827 from FEX-Emu/skmp/thunks-json
Thunks: Add Thunk json, thunk guest folder
2021-03-11 11:05:23 +02:00
Ryan Houdek da1532f3b8 Merge pull request #836 from Sonicadvance1/fexcore_config_nodefault_value
Adds config value option constructor without default initializer
2021-03-11 01:04:58 -08:00
Ryan Houdek c73d2b2bca Merge pull request #843 from Sonicadvance1/locked_not
Adds support for locked NOT
2021-03-11 00:47:01 -08:00
Ryan Houdek da43db8787 Merge pull request #842 from Sonicadvance1/locked_adc_sbb
Adds support for locked ADC and SBB
2021-03-11 00:46:53 -08:00
Ryan Houdek ab8aff6e0d Adds unittests for locked not 2021-03-11 00:38:44 -08:00
Ryan Houdek 8a7caa82b2 Adds support for locked NOT
This was trivial
2021-03-11 00:38:12 -08:00
Ryan Houdek 2706c9aaae Adds unittests for locked ADC and SBB 2021-03-11 00:24:03 -08:00
Ryan Houdek 96f2a461f8 Adds support for locked ADC and SBB
These were fairly straightforward
2021-03-11 00:22:35 -08:00
Ryan Houdek cf966eb2e0 Adds config value option constructor without default initializer
This has the expected behaviour that the configuration option will be available and not default initialized.
To enforce this fact, it will assert if it tries to get a value that doesn't exist yet.
2021-03-10 23:52:19 -08:00
Ryan Houdek 6b860cd68e Merge pull request #839 from Sonicadvance1/more_32bit_syscalls
Adds a couple new 32bit syscalls found while running Steam
2021-03-10 21:34:31 -08:00
Ryan Houdek 6b4ccc46a8 Merge pull request #838 from Sonicadvance1/fix_cmpxchg_32bit_reg
Fixes an edge case of 32bit cmpxchg <reg>, <reg>
2021-03-10 21:34:24 -08:00
Ryan Houdek 8acb18b2b9 Merge pull request #837 from Sonicadvance1/fix_logmanager
Fixes crash with large strings through LogManager
2021-03-10 21:34:17 -08:00
Ryan Houdek 832a634f64 Merge pull request #841 from phire/vixl_cmake_improvements
Improve vixl cmakefiles
2021-03-10 21:20:57 -08:00
Scott Mansell ec196b47e2 Improve Vixl cmakefiles 2021-03-11 16:43:55 +13:00
Ryan Houdek 4dc2cf2b32 Adds more cmpxchg <reg>, <reg> unit tests
This will catch the previous fix
2021-03-10 17:47:54 -08:00
Ryan Houdek e264194453 Fixes an edge case of 32bit cmpxchg <reg>, <reg>
The comparison was using the wrong values which means this cmpxchg would fail
2021-03-10 17:47:54 -08:00
Ryan Houdek ddce28df11 Merge pull request #834 from Sonicadvance1/cleanup_clang_tidy_iwyu
Moves clang-tidy arguments to root cmakelists
2021-03-10 17:47:19 -08:00
Ryan Houdek fc5bf5d261 Merge pull request #833 from Sonicadvance1/add_perf_warning
Adds x86-64 host performance warning
2021-03-10 17:47:07 -08:00
Ryan Houdek 45e5e58c7e Merge pull request #832 from Sonicadvance1/cleanup_error_interpreter
Extends error message about not being able to find guest interpreter
2021-03-10 17:46:52 -08:00
Ryan Houdek 0dc7fd75e1 Adds a couple new 32bit syscalls found while running Steam
These are straightforward to implement. Only a few dozen more syscalls missing on 32bit
2021-03-10 17:43:21 -08:00
Ryan Houdek 10fbf60259 Converts the printf usages in FEXCore to LogManager
Now that we can safely handle long strings these will now work
2021-03-10 16:38:46 -08:00
Ryan Houdek 4e55f29589 Fixes crash with large strings through LogManager
va_list types have an edge where after operating on it, then it is undefined behaviour to do anything other than va_end on the object.
So in order to call vsnprintf on it multiple times we must create copies with va_copy.
Easy enough to work around by doing a few more copies of the va_list
2021-03-10 16:36:33 -08:00
Ryan Houdek 1ddfa3e7ed Moves clang-tidy arguments to root cmakelists
This way we don't need to redeclare the arguments twice
Also moves IWYU lower so it doesn't hit any external projects other than FEXCore

Still not running these since everything needs to be cleaned up anyway
2021-03-10 12:38:16 -08:00
Ryan Houdek ec3883181b Adds x86-64 host performance warning
Lets users know that x86-64 isn't our optimal target but can still be used if passed in a new cmake argument.
Easy enough just pass in -DENABLE_X86_HOST_DEBUG=True to cmake.

Closes #776
2021-03-10 11:36:19 -08:00
Ryan Houdek e2ec645855 Extends error message about not being able to find guest interpreter
Passes the result back up to the frontend as well which allows us to early exit correctly.
Also ensures that we return ENOEXEC on these error cases so if someone is waiting on a return value, they don't just get zero
Fixes #757
2021-03-10 11:11:36 -08:00
Stefanos Kornilios Mitsis Poiitidis f0d3004363 Thunks: Add Thunk json, thunk guest folder 2021-03-10 15:51:34 +02:00
Stefanos Kornilios Mitsis Poiitidis e708b83dbb Merge pull request #826 from Sonicadvance1/disable_rcpc
Disables RCPC on ARM64 JIT
2021-03-10 12:31:17 +02:00
Ryan Houdek a9d3851684 Disables RCPC on ARM64 JIT
This is currently bugged on Snapdragon 865.
Seems to only affect its prime core. Smells like errata.
2021-03-10 02:13:16 -08:00
Ryan Houdek 7e4ba77d72 Merge pull request #825 from FEX-Emu/skmp/fix-sigchld
Signals: SA_NOCLDSTOP only blocks CLD_CONTINUED/STOPPED/TRAPPED
2021-03-10 01:50:04 -08:00
Stefanos Kornilios Mitsis Poiitidis 449ea3e80b Signals: SA_NOCLDSTOP only blocks CLD_CONTINUED/STOPPED/TRAPPED 2021-03-10 11:25:40 +02:00
Stefanos Kornilios Mitsis Poiitidis 8fd91a5b7a Merge pull request #821 from Sonicadvance1/update_vixl
Updates vixl submodule
2021-03-10 09:12:40 +02:00
Ryan Houdek e3d6db2eb5 Merge pull request #823 from FEX-Emu/skmp/fix-secondaryalu-atomics
OpDisp: Add atomic logic for SecondaryALUOp
2021-03-09 18:15:18 -08:00
Stefanos Kornilios Mitsis Poiitidis d930de30ef OpDisp: Add atomic logic for SecondaryALUOp 2021-03-09 18:32:13 +02:00
Scott Mansell b36378369a Merge pull request #822 from FEX-Emu/skmp/fix-scm-bug
ConstProp: Restrict imm code motion around selects to matching sizes, fixes dav1d
2021-03-09 22:12:50 +13:00
Stefanos Kornilios Mitsis Poiitidis cbf4c9a687 ConstProp: Restrict imm code motion around selects to matching sizes, fixes dav1d 2021-03-09 11:06:16 +02:00
Ryan Houdek bcebe7d47a Updates vixl submodule
Works around the no enum enum conversion warnings
2021-03-08 23:57:58 -08:00
Ryan Houdek e9b057bc8b Merge pull request #817 from phire/SeparateThreadAndState
Separate thread and state
2021-03-08 23:41:56 -08:00
Scott Mansell afda855caf disptacher fixes/cleanups 2021-03-09 20:28:53 +13:00
Scott Mansell 6ed36e3304 Refactor syscalls to take Frame
The vast majoirty of syscalls don't need anything in thread or frame.
So lets save an indirection for all those syscalls.

Most of the syscalls which do need Thread (or CTX via
Thread are in Thread.cpp or Memory.cpp
These have all been modifiy to fetch Thread from Frame
2021-03-09 20:28:50 +13:00
Scott Mansell 35c01038ba Seperate per-thread and per-frame state 2021-03-09 20:27:44 +13:00
Scott Mansell 43a087d554 Allow both interpeter dispatchers to be built
This commit doesn't include cmakefile changes to actually do this.
2021-03-09 20:25:14 +13:00
Ryan Houdek a1fed323df Merge pull request #820 from Sonicadvance1/remove_config_duplication
Deduplicates some configuration data
2021-03-08 23:20:04 -08:00
Ryan Houdek 0bcd1c7df5 Merge pull request #816 from Sonicadvance1/cpuid_80000005
Implements CPUID 0x8000'0005 for L1 cacheline information
2021-03-08 23:15:03 -08:00
Ryan Houdek a3ba1ae199 Deduplicates some configuration data
Configuration mapping was duplicated between three different tables.
Additionally default configuration values were strewn about. Making it confusing as to what the default value would end up being

Adds a new ConfigValues.inl header that defines a few things right next to each other.
Defines the enum name as usual.
Defines the JSON config option name.
Defines the Environment config option name
Defines the default value that the configuration should be
2021-03-08 22:53:41 -08:00
Ryan Houdek a2278cfb1e Adds StrConv type for enum conversion 2021-03-08 22:50:42 -08:00
Scott Mansell 98cb54e3e4 Merge pull request #812 from phire/both_sides
Allow both ARM64 and X86_64 jits to be compiled at the same time
2021-03-09 14:59:11 +13:00
Scott Mansell cb17127ccb Remove last traces of vixl simulator mode 2021-03-08 16:30:25 +13:00
Scott Mansell 93b6dbdc2b Fix Arm64JitCore instantiation on x86 2021-03-08 16:30:25 +13:00
Scott Mansell bea0adc3d1 Allow both jitcores to be enabled simultaneously 2021-03-08 16:30:25 +13:00
Scott Mansell 9c339c85c6 Cmake: allow independant control of both jits 2021-03-08 16:30:25 +13:00
Scott Mansell 9cf7545f8d Abstract mcontext out of JITs
This will help with unification of common code later
2021-03-08 16:30:25 +13:00
Ryan Houdek c657165e56 Implements CPUID 0x8000'0005 for L1 cacheline information
@skmp commented about this in the L2 cacheline information PR.
Quickly implement it since it also has cacheline information.

Fairly trivial since every piece of information is just some 8bit variables.
2021-03-06 08:32:47 -08:00
Ryan Houdek 05ae7d3c5e Merge pull request #814 from Sonicadvance1/cacheline_cpuid
Implements CPUID 0x8000'0006 for cacheline information
2021-03-06 08:23:03 -08:00
Ryan Houdek f94529ebb2 Implements CPUID 0x8000'0006 for cacheline information
This exposes some cache information including cacheline information.

Fills out this data structure as well just incase we hit another cacheline size bug
2021-03-06 07:30:25 -08:00
Ryan Houdek 852ab4e245 Merge pull request #815 from Sonicadvance1/sse4_1_movnt
Implements MOVNTDQA
2021-03-06 07:25:56 -08:00
Ryan Houdek bb6e7360a2 Merge pull request #813 from Sonicadvance1/disable_python_dev_check
Disabled cmake check for python development
2021-03-06 07:25:16 -08:00
Ryan Houdek 50ae18cbaf Adds MOVNTDQA unit test 2021-03-06 06:23:27 -08:00
Ryan Houdek 1adbf66fee Implements MOVNTDQA
Found this when running Steamlink under FEX on my laptop.
Turns out the Iris video driver now uses this unconditionally in a couple of locations.
Makes sense there since all the Iris targets are expected to support SSE4.1 atm.
2021-03-06 06:23:18 -08:00
Ryan Houdek 8222acdf20 Disabled cmake check for python development
We only need the interpreter
2021-03-06 06:06:59 -08:00
263 changed files with 8880 additions and 5520 deletions

No files matched your search

+16 -2
View File
@@ -25,6 +25,9 @@ jobs:
steps:
- uses: actions/checkout@v2
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
# Need to update submodules
run: git submodule update --init --depth 1
@@ -45,7 +48,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -113,13 +116,24 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
+111 -18
View File
@@ -12,6 +12,8 @@ option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -27,14 +29,6 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_IWYU)
find_program(IWYU_EXE "iwyu")
if (IWYU_EXE)
message(STATUS "IWYU enabled")
set(CMAKE_CXX_INCLUDE_WHAT_YOU_USE "${IWYU_EXE}")
endif()
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -83,6 +77,14 @@ set (CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" Be warned: FEX isn't optimized for x86_64 hosts!\n"
" Support for x86_64 hosts is only for debugging and convenience!\n"
" Don't expect amazing performance or optimal code generation!\n"
" Pass -DENABLE_X86_HOST_DEBUG=True to bypass this message!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
@@ -91,19 +93,17 @@ endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
set(_M_ARM_64 1)
add_definitions(-D_M_ARM_64=1)
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
if(CMAKE_BUILD_TYPE MATCHES DEBUG)
add_definitions(-DVIXL_DEBUG=1)
endif()
endif()
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
message(FATAL_ERROR "FEX doesn't support getting compiled with GCC!")
endif()
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter Development)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
@@ -149,6 +149,14 @@ if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
# Disable some Werror that can add frustration when developing
add_compile_options(-Wno-error=unused-variable)
endif()
endif()
if(_M_ARM_64)
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
@@ -165,16 +173,82 @@ if(_M_ARM_64)
endif()
endif()
if (ENABLE_IWYU)
find_program(IWYU_EXE "iwyu")
if (IWYU_EXE)
message(STATUS "IWYU enabled")
set(CMAKE_CXX_INCLUDE_WHAT_YOU_USE "${IWYU_EXE}")
endif()
endif()
if (ENABLE_CLANG_FORMAT)
find_program(CLANG_TIDY_EXE "clang-tidy")
if (NOT CLANG_TIDY_EXE)
message(FATAL_ERROR "Couldn't find clang-tidy")
endif()
set(CLANG_TIDY_FLAGS
"-checks=*"
"-fuchsia*"
"-bugprone-macro-parentheses"
"-clang-analyzer-core.*"
"-cppcoreguidelines-pro-type-*"
"-cppcoreguidelines-pro-bounds-array-to-pointer-decay"
"-cppcoreguidelines-pro-bounds-pointer-arithmetic"
"-cppcoreguidelines-avoid-c-arrays"
"-cppcoreguidelines-avoid-magic-numbers"
"-cppcoreguidelines-pro-bounds-constant-array-index"
"-cppcoreguidelines-no-malloc"
"-cppcoreguidelines-special-member-functions"
"-cppcoreguidelines-owning-memory"
"-cppcoreguidelines-macro-usage"
"-cppcoreguidelines-avoid-goto"
"-google-readability-function-size"
"-google-readability-namespace-comments"
"-google-readability-braces-around-statements"
"-google-build-using-namespace"
"-hicpp-*"
"-llvm-namespace-comment"
"-llvm-include-order" # Messes up with case sensitivity
"-llvmlibc-*"
"-misc-unused-parameters"
"-modernize-loop-convert"
"-modernize-use-auto"
"-modernize-avoid-c-arrays"
"-modernize-use-nodiscard"
"readability-*"
"-readability-function-size"
"-readability-implicit-bool-conversion"
"-readability-braces-around-statements"
"-readability-else-after-return"
"-readability-magic-numbers"
"-readability-named-parameter"
"-readability-uppercase-literal-suffix"
"-cert-err34-c"
"-cert-err58-cpp"
"-bugprone-exception-escape"
)
string(REPLACE ";" "," CLANG_TIDY_FLAGS "${CLANG_TIDY_FLAGS}")
set(CMAKE_CXX_CLANG_TIDY ${CLANG_TIDY_EXE} "${CLANG_TIDY_FLAGS}")
endif()
add_compile_options(-Wall)
add_subdirectory(External/FEXCore)
add_subdirectory(Source/)
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/Config.h)
if (BUILD_TESTS)
include(CTest)
enable_testing()
message(STATUS "Unit tests are enabled")
endif()
add_subdirectory(External/FEXCore)
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
@@ -185,16 +259,35 @@ if (BUILD_THUNKS)
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
CMAKE_ARGS "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
install(
CODE "MESSAGE(\"-- Installing: host-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkHostsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Host
)"
DEPENDS host-libs
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}" "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
install(
CODE "MESSAGE(\"-- Installing: guest-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkGuestsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
DEPENDS guest-libs
)
endif()
+25
View File
@@ -0,0 +1,25 @@
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS *.json)
file(GLOB GEN_CONFIG_SOURCES CONFIGURE_DEPENDS *.json.in)
# Any application configuration json file gets installed
foreach(CONFIG_SRC ${CONFIG_SOURCES})
install(FILES ${CONFIG_SRC}
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
endforeach()
# Any configuration file json file that needs to be generated
# First generate then install it
foreach(GEN_CONFIG_SRC ${GEN_CONFIG_SOURCES})
# Get the filename only component
get_filename_component(CONFIG_NAME ${GEN_CONFIG_SRC} NAME_WE)
# Configure it
configure_file(
${GEN_CONFIG_SRC}
${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json)
# Then install the configured json
install(
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
endforeach()
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"Env": "STEAM_GAME_LAUNCH_SHELL=@CMAKE_INSTALL_PREFIX@/bin/FEXBash"
}
}
+2 -1
View File
@@ -4,7 +4,8 @@ FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-10 llvm-10 nasm ninja-build libnuma-dev \
libcap-dev libglfw3-dev libepoxy-dev
libcap-dev libglfw3-dev libepoxy-dev python3-dev \
python3 linux-headers-generic
COPY . /opt/FEX
+12 -12
View File
@@ -4,6 +4,17 @@ project(${PROJECT_NAME}
VERSION 0.01
LANGUAGES CXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
set(_M_ARM_64 1)
endif()
set(ENABLE_JIT_X86_64 ${_M_X86_64} CACHE BOOL "Enable the x86_64 JIT")
set(ENABLE_JIT_ARM64 ${_M_ARM_64} CACHE BOOL "Enable the ARM64 JIT")
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_JITSYMBOLS "Enable visibility of JITSymbols in profiling tools" FALSE)
@@ -17,22 +28,11 @@ set(CMAKE_INCLUDE_CURRENT_DIR ON)
include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
message(STATUS "Enabling x86-64 JIT")
set(ENABLE_JIT 1)
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
message(STATUS "Enabling AArch64 JIT")
set(_M_ARM_64 1)
set(ENABLE_JIT 1)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
endif()
endif()
set(CMAKE_CXX_STANDARD 20)
+432
View File
@@ -0,0 +1,432 @@
import datetime
import json
import sys
def print_header():
header = '''#ifndef OPT_BASE
#define OPT_BASE(type, group, enum, json, default)
#endif
#ifndef OPT_BOOL
#define OPT_BOOL(group, enum, json, default) OPT_BASE(bool, group, enum, json, default)
#endif
#ifndef OPT_UINT8
#define OPT_UINT8(group, enum, json, default) OPT_BASE(uint8_t, group, enum, json, default)
#endif
#ifndef OPT_INT32
#define OPT_INT32(group, enum, json, default) OPT_BASE(int32_t, group, enum, json, default)
#endif
#ifndef OPT_UINT32
#define OPT_UINT32(group, enum, json, default) OPT_BASE(uint32_t, group, enum, json, default)
#endif
#ifndef OPT_UINT64
#define OPT_UINT64(group, enum, json, default) OPT_BASE(uint64_t, group, enum, json, default)
#endif
#ifndef OPT_STR
#define OPT_STR(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#endif
#ifndef OPT_STRARRAY
#define OPT_STRARRAY(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#endif
'''
output_file.write(header)
def print_tail():
tail = '''#undef OPT_BASE
#undef OPT_BOOL
#undef OPT_UINT8
#undef OPT_INT32
#undef OPT_UINT32
#undef OPT_UINT64
#undef OPT_STR
#undef OPT_STRARRAY
'''
output_file.write(tail)
def print_config(type, group_name, json_name, default_value):
output_file.write("OPT_{0} ({1}, {2}, {3}, {4})\n".format(type.upper(), group_name.upper(), json_name.upper(), json_name, default_value))
def print_options(options):
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
print_config(
op_vals["Type"],
op_group,
op_key,
default)
output_file.write("\n")
def print_unnamed_options(options):
output_file.write("// Unnamed configuration options\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
print_config(
op_vals["Type"],
op_group,
op_key.upper(), # KEY is the enum here, there is no json configuration for these
default)
output_file.write("\n")
def print_man_option(short, long, desc, default):
if (short != None):
output_man.write(".It Fl {0} , ".format(short))
else:
output_man.write(".It ")
output_man.write("Fl Fl {0}=".format(long))
output_man.write("\n");
# Print description
for line in desc:
output_man.write(".Pp\n")
output_man.write("{0}\n".format(line))
output_man.write(".Pp\n")
output_man.write("\\fBdefault:\\fR {0}\n".format(default))
output_man.write(".Pp\n\n")
def print_man_env_option(name, desc, default):
output_man.write("\\fBFEX_{0}\\fR\n".format(name))
# Print description
for line in desc:
output_man.write(".Pp\n")
output_man.write("{0}\n".format(line))
output_man.write(".Pp\n")
output_man.write("\\fBdefault:\\fR {0}\n".format(default))
output_man.write(".Pp\n\n")
def print_man_options(options):
output_man.write(".Sh OPTIONS\n")
output_man.write(".Bl -tag -width -indent\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
default = op_vals["Default"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_option(
short,
long,
op_vals["Desc"],
default
)
output_man.write(".El\n")
def print_man_environment(options):
output_man.write(".Sh ENVIRONMENT\n")
output_man.write(".Bl -tag -width -indent\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_env_option(
op_key.upper(),
op_vals["Desc"],
default
)
output_man.write(".El\n")
def print_man_header():
header ='''.Dd {0}
.Dt FEX
.Os Linux
.Sh NAME
.Nm FEXLoader
.Nm FEXInterpreter
.Nm FEXBash
.Nd Fast x86-64 and x86 emulation.
.Sh SYNOPSIS
.Nm
.Op options
.Op Ar --
.Ar Application
<args> ...
.Pp
.Nm FEXInterpreter
.Ar Application
<args> ...
.Pp
.Nm FEXBash
.Ar <args> ...
.Sh DESCRIPTION
FEX allows you to run x86 and x86-64 binaries on an AArch64 host, similar to qemu-user and box86.
It has native support for a rootfs overlay, so you don't need to chroot, as well as some thunklibs so it can forward things like GL to the host.
FEX presents a Linux 5.0 interface to the guest, and supports both AArch64 and x86-64 as hosts.
FEX is very much work in progress, so expect things to change.
'''
output_man.write(header.format(datetime.datetime.now().strftime("%d-%m-%Y")))
def print_man_tail():
tail ='''.Sh FILES
.Bl -tag -width "$prefix/share/fex-emu/GuestThunks" -compact
.It Pa $XDG_HOME_DIR/.fex-emu
Default FEX user configuration directory
.It Pa $prefix/share/fex-emu/AppConfig
System level application configuration files
.It Pa $prefix/share/fex-emu/GuestThunks
guest-side thunk data libraries
.It Pa $prefix/lib/fex-emu/HostThunks
host-side thunks for guest communication
.El
'''
output_man.write(tail)
def print_config_option(type, group_name, json_name, default_value, short, choices, desc):
if (type == "bool"):
# Bool gets some special handling to add an inverted case
output_argloader.write("{0}Group".format(group_name))
options = ""
AddedArg = False
if (short != None):
AddedArg = True
options += "\"-{0}\"".format(short)
if (AddedArg):
options += ", "
options += "\"--{0}\"".format(json_name.lower())
output_argloader.write(".add_option({0})".format(options))
output_argloader.write("\n")
output_argloader.write("\t.action(\"store_true\")\n")
output_argloader.write("\t.dest(\"{0}\")\n".format(json_name));
# help
output_argloader.write("\t.help(\n")
desc_line_ender = ""
if (len(desc) > 1):
desc_line_ender = "\\n"
for line in desc:
output_argloader.write("\t\t\"{0}{1}\"\n".format(line, desc_line_ender))
output_argloader.write("\t)\n")
output_argloader.write("\t.set_default({0});\n\n".format(default_value));
output_argloader.write("{0}Group".format(group_name))
output_argloader.write(".add_option(\"--no-{0}\")\n".format(json_name.lower()))
# Inverted case sets the bool to false
output_argloader.write("\t.action(\"store_false\")\n")
output_argloader.write("\t.dest(\"{0}\");\n".format(json_name));
else:
output_argloader.write("{0}Group".format(group_name))
options = ""
AddedArg = False
if (short != None):
AddedArg = True
options += "\"-{0}\"".format(short)
if (AddedArg):
options += ", "
options += "\"--{0}\"".format(json_name.lower())
output_argloader.write(".add_option({0})".format(options))
output_argloader.write("\n")
output_argloader.write("\t.dest(\"{0}\")\n".format(json_name));
if (choices != None):
output_argloader.write("\t.choices({\n")
for choice in choices:
output_argloader.write("\t\t\"{0}\",\n".format(choice))
output_argloader.write("\t})\n")
# help
output_argloader.write("\t.help(\n")
desc_line_ender = ""
if (len(desc) > 1):
desc_line_ender = "\\n"
for line in desc:
output_argloader.write("\t\t\"{0}{1}\"\n".format(line, desc_line_ender))
output_argloader.write("\t)\n")
output_argloader.write("\t.set_default({0});\n".format(default_value));
output_argloader.write("\n");
def print_argloader_options(options):
output_argloader.write("#ifdef BEFORE_PARSE\n")
output_argloader.write("#undef BEFORE_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = "\"" + op_vals["TextDefault"] + "\""
short = None
choices = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if ("Choices" in op_vals):
choices = op_vals["Choices"]
print_config_option(
op_vals["Type"],
op_group,
op_key,
default,
short,
choices,
op_vals["Desc"])
output_argloader.write("\n")
output_argloader.write("#endif\n")
def print_parse_argloader_options(options):
output_argloader.write("#ifdef AFTER_PARSE\n")
output_argloader.write("#undef AFTER_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
output_argloader.write("if (Options.is_set_by_user(\"{0}\")) {{\n".format(op_key))
value_type = op_vals["Type"]
NeedsString = False
conversion_func = "std::to_string"
if ("ArgumentHandler" in op_vals):
NeedsString = True
conversion_func = "FEX::Handler::{0}".format(op_vals["ArgumentHandler"])
if (value_type == "str"):
NeedsString = True
conversion_func = ""
if (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
output_argloader.write("\tfor (auto iter = Array.begin(); iter != Array.end(); ++iter) {\n")
output_argloader.write("\t\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, *iter);\n".format(op_key.upper()))
output_argloader.write("\t}\n")
else:
if (NeedsString):
output_argloader.write("\tstd::string UserValue = Options[\"{0}\"];\n".format(op_key))
else:
output_argloader.write("\t{0} UserValue = Options.get(\"{1}\");\n".format(value_type, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, {1}(UserValue));\n".format(op_key.upper(), conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def check_for_duplicate_options(options):
short_map = []
long_map = []
# Spin through all the items and see if we have a duplicate option
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
long_invert = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if (op_vals["Type"] == "bool"):
long_invert = "no-" + long
# Check for short key duplication
if (short != None):
if (short in short_map):
raise Exception("Short config '{0}' for option '{1}' has duplicate entry!".format(short, op_key))
else:
short_map.append(short)
# Check for long key duplication
if (long in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long))
else:
long_map.append(long)
# Check for long key duplication
if (long_invert != None):
if (long_invert in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long_invert))
else:
long_map.append(long_invert)
if (len(sys.argv) < 5):
sys.exit()
output_filename = sys.argv[2]
output_man_page = sys.argv[3]
output_argumentloader_filename = sys.argv[4]
json_file = open(sys.argv[1], "r")
json_text = json_file.read()
json_file.close()
json_object = json.loads(json_text)
options = json_object["Options"]
unnamed_options = json_object["UnnamedOptions"]
check_for_duplicate_options(options)
# Generate config include file
output_file = open(output_filename, "w")
print_header()
print_options(options)
print_unnamed_options(unnamed_options)
print_tail()
output_file.close()
# Generate man file
output_man = open(output_man_page, "w")
print_man_header()
print_man_options(options)
print_man_environment(options)
print_man_tail()
output_man.close()
# Generate argument loader code
output_argloader = open(output_argumentloader_filename, "w")
print_argloader_options(options);
print_parse_argloader_options(options);
output_argloader.close()
+110 -93
View File
@@ -1,49 +1,4 @@
if (ENABLE_CLANG_FORMAT)
find_program(CLANG_TIDY_EXE "clang-tidy")
set(CLANG_TIDY_FLAGS
"-checks=*"
"-fuchsia*"
"-bugprone-macro-parentheses"
"-clang-analyzer-core.*"
"-cppcoreguidelines-pro-type-*"
"-cppcoreguidelines-pro-bounds-array-to-pointer-decay"
"-cppcoreguidelines-pro-bounds-pointer-arithmetic"
"-cppcoreguidelines-avoid-c-arrays"
"-cppcoreguidelines-avoid-magic-numbers"
"-cppcoreguidelines-pro-bounds-constant-array-index"
"-cppcoreguidelines-no-malloc"
"-cppcoreguidelines-special-member-functions"
"-cppcoreguidelines-owning-memory"
"-cppcoreguidelines-macro-usage"
"-cppcoreguidelines-avoid-goto"
"-google-readability-function-size"
"-google-readability-namespace-comments"
"-google-readability-braces-around-statements"
"-google-build-using-namespace"
"-hicpp-*"
"-llvm-namespace-comment"
"-llvm-include-order" # Messes up with case sensitivity
"-misc-unused-parameters"
"-modernize-loop-convert"
"-modernize-use-auto"
"-modernize-avoid-c-arrays"
"-modernize-use-nodiscard"
"readability-*"
"-readability-function-size"
"-readability-implicit-bool-conversion"
"-readability-braces-around-statements"
"-readability-else-after-return"
"-readability-magic-numbers"
"-readability-named-parameter"
"-readability-uppercase-literal-suffix"
"-cert-err34-c"
"-cert-err58-cpp"
"-bugprone-exception-escape"
)
string(REPLACE ";" "," CLANG_TIDY_FLAGS "${CLANG_TIDY_FLAGS}")
set(CMAKE_CXX_CLANG_TIDY ${CLANG_TIDY_EXE} "${CLANG_TIDY_FLAGS}")
endif()
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (SRCS
Common/Paths.cpp
@@ -130,6 +85,11 @@ set (SRCS
Interface/Core/X86Tables.cpp
Interface/Core/X86DebugInfo.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64_stubs.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/X86Tables/BaseTables.cpp
@@ -164,58 +124,58 @@ set (SRCS
Utils/ELFLoader.cpp
Utils/ELFSymbolDatabase.cpp
Utils/LogManager.cpp
Utils/Threads.cpp
)
if (_M_X86_64)
list(APPEND SRCS Interface/Core/Interpreter/x86_64Dispatcher.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp
Interface/Core/Interpreter/Arm64Dispatcher.cpp)
Interface/Core/ArchHelpers/Arm64.cpp)
endif()
set (JIT_LIBS )
if (ENABLE_JIT)
if (_M_X86_64)
add_definitions(-D_M_X86_64=1)
if (NOT FORCE_AARCH64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
endif()
endif()
if(_M_ARM_64)
add_definitions(-D_M_ARM_64=1)
add_definitions(-DVIXL_INCLUDE_TARGET_AARCH64=1)
add_definitions(-DVIXL_CODE_BUFFER_MMAP=1)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
list(APPEND JIT_LIBS vixl)
endif()
set(DEFINES )
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
endif()
if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
list(APPEND DEFINES -DJIT_X86_64)
endif()
if (ENABLE_JIT_ARM64)
list(APPEND DEFINES -DJIT_ARM64)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
endif()
if (ENABLE_JITSYMBOLS)
add_definitions(-DENABLE_JITSYMBOLS=1)
list(APPEND DEFINES -DENABLE_JITSYMBOLS=1)
endif()
# Generate IR include file
@@ -251,20 +211,60 @@ add_custom_command(
set_source_files_properties(${OUTPUT_IR_NAME} PROPERTIES
GENERATED TRUE)
# Create teh target
# Create the target
add_custom_target(IR_INC
DEPENDS "${OUTPUT_NAME}"
DEPENDS "${OUTPUT_IR_DOC}")
# Generate the configuration include file
set(OUTPUT_CONFIG_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/Config")
set(OUTPUT_CONFIG_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigValues.inl")
set(OUTPUT_CONFIG_OPTION_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigOptions.inl")
set(INPUT_CONFIG_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_CONFIG_NAME}"
OUTPUT "${OUTPUT_CONFIG_OPTION_NAME}"
OUTPUT "${OUTPUT_MAN_NAME}"
DEPENDS "${INPUT_CONFIG_NAME}"
DEPENDS CREATE_CONFIG_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py" "${INPUT_CONFIG_NAME}" "${OUTPUT_CONFIG_NAME}" "${OUTPUT_MAN_NAME}"
"${OUTPUT_CONFIG_OPTION_NAME}"
)
set_source_files_properties(${OUTPUT_CONFIG_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_MAN_NAME} PROPERTIES
GENERATED TRUE)
# Create the target
add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_CONFIG_NAME}"
DEPENDS "${OUTPUT_CONFIG_OPTION_NAME}"
DEPENDS "${OUTPUT_MAN_NAME}")
# Install the man page
install(FILES ${OUTPUT_MAN_NAME} DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
check_cxx_compiler_flag(-fdiagnostics-color=always GCC_COLOR)
check_cxx_compiler_flag(-fcolor-diagnostics CLANG_COLOR)
function(AddLibrary Name Type)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} pthread rt ${JIT_LIBS} ${LINUX_LIBS} dl)
add_dependencies(${Name} CONFIG_INC)
target_link_libraries(${Name} pthread rt vixl ${LINUX_LIBS} dl)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
@@ -275,9 +275,15 @@ function(AddLibrary Name Type)
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
target_compile_definitions(${Name} PRIVATE ${DEFINES})
target_compile_options(${Name}
PRIVATE
-Wno-trigraphs -Wall)
-Wall
-Werror=implicit-fallthrough
-Wno-trigraphs
)
if (GCC_COLOR)
target_compile_options(${Name}
@@ -291,6 +297,17 @@ function(AddLibrary Name Type)
endif()
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} pthread rt vixl ${LINUX_LIBS} dl)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
endfunction()
AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
@@ -96,6 +96,7 @@ extFloat80_t
switch ( roundingMode ) {
case softfloat_round_near_even:
if ( !(sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF )) ) break;
__attribute__((fallthrough));
case softfloat_round_near_maxMag:
if ( exp == 0x3FFE ) goto mag1;
break;
+7
View File
@@ -38,4 +38,11 @@ namespace FEXCore::StrConv {
*Result = Value;
return true;
}
template <typename T,
typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]] static bool Conv(std::string_view Value, T *Result) {
*Result = static_cast<T>(std::stoull(std::string(Value), nullptr, 0));
return true;
}
}
+224 -105
View File
@@ -4,118 +4,101 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Config/Config.h>
#include <filesystem>
#include <pwd.h>
#include <map>
#include <sys/sysinfo.h>
#include <unistd.h>
namespace FEXCore::Config {
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
switch (Option) {
case FEXCore::Config::CONFIG_MULTIBLOCK:
CTX->Config.Multiblock = Config != 0;
break;
case FEXCore::Config::CONFIG_MAXBLOCKINST:
CTX->Config.MaxInstPerBlock = Config;
break;
case FEXCore::Config::CONFIG_DEFAULTCORE:
CTX->Config.Core = static_cast<FEXCore::Config::ConfigCore>(Config);
break;
case FEXCore::Config::CONFIG_VIRTUALMEMSIZE:
CTX->Config.VirtualMemSize = Config;
break;
case FEXCore::Config::CONFIG_SINGLESTEP:
CTX->Config.RunningMode = Config != 0 ? FEXCore::Context::CoreRunningMode::MODE_SINGLESTEP : FEXCore::Context::CoreRunningMode::MODE_RUN;
break;
case FEXCore::Config::CONFIG_GDBSERVER:
Config != 0 ? CTX->StartGdbServer() : CTX->StopGdbServer();
break;
case FEXCore::Config::CONFIG_IS64BIT_MODE:
CTX->Config.Is64BitMode = Config != 0;
break;
case FEXCore::Config::CONFIG_TSO_ENABLED:
CTX->Config.TSOEnabled = Config != 0;
break;
case FEXCore::Config::CONFIG_SMC_CHECKS:
CTX->Config.SMCChecks = static_cast<FEXCore::Config::ConfigSMCChecks>(Config);
break;
case FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS:
CTX->Config.ABILocalFlags = Config != 0;
break;
case FEXCore::Config::CONFIG_ABI_NO_PF:
CTX->Config.ABINoPF = Config != 0;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
CTX->Config.ValidateIRarser = Config != 0;
break;
case FEXCore::Config::CONFIG_AOTIR_GENERATE:
CTX->Config.AOTIRCapture = Config != 0;
break;
case FEXCore::Config::CONFIG_AOTIR_LOAD:
CTX->Config.AOTIRLoad = Config != 0;
break;
default: LogMan::Msg::A("Unknown configuration option");
char const* FindUserHomeThroughUID() {
auto passwd = getpwuid(geteuid());
if (passwd) {
return passwd->pw_dir;
}
return nullptr;
}
const char *GetHomeDirectory() {
char const *HomeDir = getenv("HOME");
// Try to get home directory from uid
if (!HomeDir) {
HomeDir = FindUserHomeThroughUID();
}
// try the PWD
if (!HomeDir) {
HomeDir = getenv("PWD");
}
// Still doesn't exit? You get local
if (!HomeDir) {
HomeDir = ".";
}
return HomeDir;
}
std::string GetConfigDirectory(bool Global) {
std::string ConfigDir;
if (Global) {
ConfigDir = GLOBAL_DATA_DIRECTORY;
}
else {
char const *HomeDir = GetHomeDirectory();
char const *ConfigXDG = getenv("XDG_CONFIG_HOME");
ConfigDir = ConfigXDG ? ConfigXDG : HomeDir;
ConfigDir += "/.fex-emu/";
// Ensure the folder structure is created for our configuration
if (!std::filesystem::exists(ConfigDir) &&
!std::filesystem::create_directories(ConfigDir)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigDir.c_str());
// Let's go local in this case
return "./";
}
}
return ConfigDir;
}
std::string GetConfigFileLocation() {
std::string ConfigFile = GetConfigDirectory(false) + "Config.json";
return ConfigFile;
}
std::string GetApplicationConfig(std::string &Filename, bool Global) {
std::string ConfigFile = GetConfigDirectory(Global);
if (!Global &&
!std::filesystem::exists(ConfigFile) &&
!std::filesystem::create_directories(ConfigFile)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigFile.c_str());
// Let's go local in this case
return "./";
}
ConfigFile += "AppConfig/" + Filename + ".json";
return ConfigFile;
}
std::string GetDataDirectory() {
std::string DataDir{};
char const *HomeDir = GetHomeDirectory();
char const *DataXDG = getenv("XDG_DATA_HOME");
DataDir = DataXDG ?: HomeDir;
DataDir += "/.fex-emu/";
return DataDir;
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, std::string const &Config) {
switch (Option) {
case CONFIG_ROOTFSPATH:
CTX->Config.RootFSPath = Config;
break;
case CONFIG_THUNKLIBSPATH:
CTX->Config.ThunkLibsPath = Config;
break;
case FEXCore::Config::CONFIG_DUMPIR:
CTX->Config.DumpIR = Config;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
}
uint64_t GetConfig(FEXCore::Context::Context *CTX, ConfigOption Option) {
switch (Option) {
case FEXCore::Config::CONFIG_MULTIBLOCK:
return CTX->Config.Multiblock;
break;
case FEXCore::Config::CONFIG_MAXBLOCKINST:
return CTX->Config.MaxInstPerBlock;
break;
case FEXCore::Config::CONFIG_DEFAULTCORE:
return CTX->Config.Core;
break;
case FEXCore::Config::CONFIG_VIRTUALMEMSIZE:
return CTX->Config.VirtualMemSize;
break;
case FEXCore::Config::CONFIG_SINGLESTEP:
return CTX->Config.RunningMode == FEXCore::Context::CoreRunningMode::MODE_SINGLESTEP ? 1 : 0;
case FEXCore::Config::CONFIG_GDBSERVER:
return CTX->GetGdbServerStatus();
break;
case FEXCore::Config::CONFIG_IS64BIT_MODE:
return CTX->Config.Is64BitMode;
break;
case FEXCore::Config::CONFIG_TSO_ENABLED:
return CTX->Config.TSOEnabled;
break;
case FEXCore::Config::CONFIG_SMC_CHECKS:
return CTX->Config.SMCChecks;
break;
case FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS:
return CTX->Config.ABILocalFlags;
break;
case FEXCore::Config::CONFIG_ABI_NO_PF:
return CTX->Config.ABINoPF;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
return CTX->Config.ValidateIRarser;
break;
case FEXCore::Config::CONFIG_AOTIR_GENERATE:
return CTX->Config.AOTIRCapture;
break;
case FEXCore::Config::CONFIG_AOTIR_LOAD:
return CTX->Config.AOTIRLoad;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
return 0;
}
@@ -149,6 +132,7 @@ namespace FEXCore::Config {
private:
void MergeConfigMap(const LayerOptions &Options);
void MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value);
};
void MetaLayer::Load() {
@@ -163,10 +147,56 @@ namespace FEXCore::Config {
}
}
void MetaLayer::MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
if (MetaEnvironment == OptionMap.end()) {
// Doesn't exist, just insert
OptionMap.insert_or_assign(Option, Value);
return;
}
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
std::unordered_map<std::string, std::string> LookupMap;
const auto AddToMap = [&LookupMap](FEXCore::Config::LayerValue const &Value) {
for (const auto &EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == std::string::npos) {
// Broken environment variable
// Skip
continue;
}
auto Key = std::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = std::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
// Add the key to the map, overwriting whatever previous value was there
LookupMap.insert_or_assign(std::move(Key), std::move(Value));
}
};
AddToMap(MetaEnvironment->second);
AddToMap(Value);
// Now with the two layers merged in the map
// Add all the values to the option
Erase(Option);
for (auto &Val : LookupMap) {
// Set will emplace multiple options in to its list
Set(Option, Val.first + "=" + Val.second);
}
}
void MetaLayer::MergeConfigMap(const LayerOptions &Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto &it : Options) {
OptionMap.insert_or_assign(it.first, it.second);
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV) {
MergeEnvironmentVariables(it.first, it.second);
}
else {
OptionMap.insert_or_assign(it.first, it.second);
}
}
}
@@ -189,8 +219,91 @@ namespace FEXCore::Config {
}
}
std::string ExpandPath(std::string PathName) {
if (PathName.empty()) {
return {};
}
std::filesystem::path Path{PathName};
// Expand home if it exists
if (Path.is_relative()) {
std::string Home = getenv("HOME") ?: "";
// Home expansion only works if it is the first character
// This matches bash behaviour
if (PathName.at(0) == '~') {
PathName.replace(0, 1, Home);
return PathName;
}
// Expand relative path to absolute
Path = std::filesystem::absolute(Path);
// Only return if it exists
if (std::filesystem::exists(Path)) {
return Path;
}
}
return {};
}
void ReloadMetaLayer() {
Meta->Load();
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THREADS)) {
FEX_CONFIG_OPT(Cores, THREADS);
if (Cores == 0) {
// When the number of emulated CPU cores is zero then auto detect
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THREADS, std::to_string(get_nprocs_conf()));
}
}
auto ExpandPathIfExists = [](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
}
};
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
std::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
if (std::filesystem::exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKHOSTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKHOSTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKHOSTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKGUESTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKGUESTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKGUESTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKCONFIG)) {
FEX_CONFIG_OPT(PathName, THUNKCONFIG);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKCONFIG, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_OUTPUTLOG)) {
FEX_CONFIG_OPT(PathName, OUTPUTLOG);
if (PathName() != "stdout" && PathName() != "stderr") {
ExpandPathIfExists(FEXCore::Config::CONFIG_OUTPUTLOG, PathName());
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_SINGLESTEP)) {
// Single stepping also enforces single instruction size blocks
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, std::to_string(1u));
}
}
void AddLayer(std::unique_ptr<FEXCore::Config::Layer> _Layer) {
@@ -252,8 +365,14 @@ namespace FEXCore::Config {
}
}
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template int8_t Value<int8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int8_t Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template int16_t Value<int16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int16_t Default);
template uint16_t Value<uint16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint16_t Default);
template int32_t Value<int32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int32_t Default);
template uint32_t Value<uint32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint32_t Default);
template int64_t Value<int64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int64_t Default);
template uint64_t Value<uint64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint64_t Default);
// Constructor
+232
View File
@@ -0,0 +1,232 @@
{
"Options": {
"CPU": {
"Core": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigCore::CONFIG_IRJIT",
"TextDefault": "irjit",
"ShortArg": "c",
"Choices": [ "irint", "irjit", "host" ],
"ArgumentHandler": "CoreHandler",
"Desc": [
"Which CPU core to use",
"host only exists on x86_64",
"[irint, irjit, host]"
]
},
"Multiblock": {
"Type": "bool",
"Default": "true",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation"
]
},
"MaxInst": {
"Type": "int32",
"Default": "5000",
"ShortArg": "n",
"Desc": [
"Maximum number of instruction to store in a block"
]
},
"Threads": {
"Type": "uint32",
"Default": "1",
"ShortArg": "T",
"Desc": [
"Number of physical hardware threads to tell the process we have.",
"0 will auto detect."
]
}
},
"Emulation": {
"RootFS": {
"Type": "str",
"Default": "",
"ShortArg": "R",
"Desc": [
"Which Root filesystem prefix to use",
"This can be a filesystem path",
"\teg: ~/RootFS/Debian_x86_64",
"Or this can be a name of a rootfs",
"If the named rootfs exists in the FEX data folder then it will use that one",
"\teg: $HOME/.fex-emu/RootFS/<RootFS name>/",
"Or if you have XDG_DATA_HOME the config will search in that directory",
"\teg: $XDG_DATA_HOME/.fex-emu/RootFS/<RootFS name>/"
]
},
"ThunkHostLibs": {
"Type": "str",
"Default": "",
"ShortArg": "t",
"Desc": [
"Folder to find the host-side thunking libraries."
]
},
"ThunkGuestLibs": {
"Type": "str",
"Default": "",
"ShortArg": "j",
"Desc": [
"Folder to find the guest-side thunking libraries."
]
},
"ThunkConfig": {
"Type": "str",
"Default": "",
"ShortArg": "k",
"Desc": [
"A json file specifying where to overlay the thunks."
]
},
"Env": {
"Type": "strarray",
"Default": "",
"ShortArg": "E",
"Desc": [
"Adds an environment variable to the emulated environment."
]
}
},
"Debug": {
"SingleStep": {
"Type": "bool",
"Default": "false",
"ShortArg": "S",
"Desc": [
"Single stepping configuration."
]
},
"GdbServer": {
"Type": "bool",
"Default": "false",
"ShortArg": "G",
"Desc": [
"Enables the GDB server."
]
},
"DumpIR": {
"Type": "str",
"Default": "no",
"Desc": [
"Folder to dump the IR in to.",
"[no, stdout, stderr, <Folder>]"
]
},
"DumpGPRs": {
"Type": "bool",
"Default": "false",
"ShortArg": "g",
"Desc": [
"When the test harness ends, print the GPR state."
]
},
"O0": {
"Type": "bool",
"Default": "false",
"ShortArg": "O0",
"Desc": [
"Disables optimizations passes for debugging."
]
}
},
"Logging": {
"SilentLog": {
"Type": "bool",
"Default": "true",
"ShortArg": "s",
"Desc": [
"Disables logging"
]
},
"OutputLog": {
"Type": "str",
"Default": "stdout",
"ShortArg": "o",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, <Filename>]"
]
}
},
"Hacks": {
"SMCChecks": {
"Type": "uint8",
"Default": "FEXCore::Config::CONFIG_SMC_MMAN",
"TextDefault": "mman",
"ArgumentHandler": "SMCCheckHandler",
"Desc": [
"Checks code for modification before execution.",
"\tnone: No checks",
"\tmman: Invalidate on mmap, mprotect, munmap",
"\tfull: Validate code before every run (slow)"
]
},
"TSOEnabled": {
"Type": "bool",
"Default": "true",
"Desc": [
"Controls TSO IR ops.",
"Highly likely to break any multithreaded application if disabled."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
"Desc": [
"When enabled enables an optimization around flags.",
"Assumes flags are not used across cals.",
"Hand-written assembly can violate this assumption."
]
},
"ABINoPF": {
"Type": "bool",
"Default": "false",
"Desc": [
"When enabled enables an optimization around parity flag calculation.",
"Removes the calculation of the parity flag from GPR instructions.",
"Assuming no uses rely on it"
]
}
},
"Misc": {
"AOTIRCapture": {
"Type": "bool",
"Default": "false",
"Desc": [
"Captures IR and generates an AOT IR cache.",
"Captures both the loaded executable and libraries it loads."
]
},
"AOTIRLoad": {
"Type": "bool",
"Default": "false",
"Desc": [
"Loads an AOT IR cache for the loaded executable."
]
}
}
},
"UnnamedOptions": {
"Misc": {
"IS_INTERPRETER": {
"Type": "bool",
"Default": "false"
},
"INTERPRETER_INSTALLED": {
"Type": "bool",
"Default": "false"
},
"APP_FILENAME": {
"Type": "str",
"Default": ""
},
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
}
}
}
}
+12 -29
View File
@@ -23,6 +23,9 @@ namespace FEXCore::Context {
}
void DestroyContext(FEXCore::Context::Context *CTX) {
if (CTX->ParentThread) {
CTX->DestroyThread(CTX->ParentThread);
}
delete CTX;
}
@@ -65,11 +68,11 @@ namespace FEXCore::Context {
}
void GetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, &CTX->ParentThread->State.State, sizeof(FEXCore::Core::CPUState));
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void SetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(&CTX->ParentThread->State.State, State, sizeof(FEXCore::Core::CPUState));
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
void Pause(FEXCore::Context::Context *CTX) {
@@ -84,10 +87,6 @@ namespace FEXCore::Context {
CTX->CustomCPUFactory = std::move(Factory);
}
void SetFallbackCPUBackendFactory(FEXCore::Context::Context *CTX, CustomCPUFactoryType Factory) {
CTX->FallbackCPUFactory = std::move(Factory);
}
bool AddVirtualMemoryMapping([[maybe_unused]] FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
return false;
}
@@ -123,20 +122,12 @@ namespace FEXCore::Context {
CTX->StopThread(Thread);
}
void DeleteForkedThreads(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : CTX->Threads) {
if (DeadThread == Thread) {
continue;
}
void DestroyThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->DestroyThread(Thread);
}
// Setting running to false ensures that when they are shutdown we won't send signals to kill them
DeadThread->State.RunningEvents.Running = false;
}
// We now only have one thread
CTX->IdleWaitRefCount = 1;
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
@@ -147,8 +138,8 @@ namespace FEXCore::Context {
CTX->SyscallHandler = Handler;
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, [[maybe_unused]] uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function);
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function, Leaf);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::istream>(const std::string&)> CacheReader) {
@@ -178,10 +169,6 @@ namespace Debug {
return CTX->GetRuntimeStatsForThread(Thread);
}
FEXCore::Core::CPUState GetCPUState(FEXCore::Context::Context *CTX) {
return CTX->GetCPUState();
}
bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
return CTX->GetDebugDataForRIP(RIP, Data);
}
@@ -198,10 +185,6 @@ namespace Debug {
// void SetIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir) {
// CTX->SetIRForRIP(RIP, ir);
// }
FEXCore::Core::ThreadState *GetThreadState(FEXCore::Context::Context *CTX) {
return CTX->GetThreadState();
}
}
}
+48 -36
View File
@@ -12,14 +12,15 @@
#include <FEXCore/Utils/Event.h>
#include <stdint.h>
#include <memory>
#include <map>
#include <unordered_map>
#include <set>
#include <mutex>
#include <istream>
#include <ostream>
#include <functional>
#include <istream>
#include <map>
#include <memory>
#include <mutex>
#include <optional>
#include <ostream>
#include <set>
#include <unordered_map>
namespace FEXCore {
class ThunkHandler;
@@ -28,7 +29,8 @@ class GdbServer;
class SiganlDelegator;
namespace CPU {
class JITCore;
class Arm64JITCore;
class X86JITCore;
}
namespace HLE {
class SyscallHandler;
@@ -52,34 +54,37 @@ namespace FEXCore::Context {
struct Context {
friend class FEXCore::HLE::SyscallHandler;
friend class FEXCore::CPU::JITCore;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
#ifdef JIT_X86_64
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::IR::Validation::IRValidation;
struct {
bool Multiblock {false};
bool BreakOnFrontendFailure {true};
int64_t MaxInstPerBlock {-1LL};
uint64_t VirtualMemSize {1ULL << 36};
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
FEXCore::Config::ConfigCore Core {FEXCore::Config::CONFIG_INTERPRETER};
bool GdbServer {false};
std::string RootFSPath;
std::string ThunkLibsPath;
bool Is64BitMode {true};
bool TSOEnabled {true};
FEXCore::Config::ConfigSMCChecks SMCChecks {FEXCore::Config::CONFIG_SMC_MMAN};
bool ABILocalFlags {false};
bool ABINoPF {false};
bool AOTIRCapture {false};
bool AOTIRLoad {false};
std::string DumpIR;
uint64_t VirtualMemSize{1ULL << 36};
// this is for internal use
bool ValidateIRarser { false };
FEX_CONFIG_OPT(Multiblock, MULTIBLOCK);
FEX_CONFIG_OPT(SingleStepConfig, SINGLESTEP);
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(ABINoPF, ABINOPF);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(Core, CORE);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(DumpIR, DUMPIR);
} Config;
using IntCallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
@@ -105,7 +110,6 @@ namespace FEXCore::Context {
std::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
CustomCPUFactoryType CustomCPUFactory;
CustomCPUFactoryType FallbackCPUFactory;
std::function<void(uint64_t ThreadId, FEXCore::Context::ExitReason)> CustomExitHandler;
struct AOTIRCacheEntry {
@@ -162,24 +166,27 @@ namespace FEXCore::Context {
static void RemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState
static void RemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveCodeEntry(Frame->Thread, GuestRIP);
}
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(uint64_t Thread);
FEXCore::Core::CPUState GetCPUState();
bool GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data);
bool FindHostCodeForRIP(uint64_t RIP, uint8_t **Code);
// XXX:
// bool FindIRForRIP(uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir);
// void SetIRForRIP(uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir);
FEXCore::Core::ThreadState *GetThreadState();
void LoadEntryList();
std::tuple<FEXCore::IR::IRListView *, FEXCore::IR::RegisterAllocationData *, uint64_t, uint64_t, uint64_t, uint64_t> GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<void *, FEXCore::IR::IRListView *, FEXCore::Core::DebugData *, FEXCore::IR::RegisterAllocationData *, bool, uint64_t, uint64_t> CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
bool LoadAOTIRCache(std::istream &stream);
bool WriteAOTIRCache(std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter);
@@ -191,6 +198,9 @@ namespace FEXCore::Context {
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
void RunThread(FEXCore::Core::InternalThreadState *Thread);
void DestroyThread(FEXCore::Core::InternalThreadState *Thread);
void CleanupAfterFork(FEXCore::Core::InternalThreadState *ExceptForThread);
std::vector<FEXCore::Core::InternalThreadState*> *const GetThreads() { return &Threads; }
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
@@ -200,13 +210,15 @@ namespace FEXCore::Context {
FEXCore::JITSymbols Symbols;
#endif
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
private:
void WaitForIdleWithTimeout();
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void NotifyPause();
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length);
@@ -214,7 +226,7 @@ namespace FEXCore::Context {
FEXCore::CodeLoader *LocalLoader{};
// Entry Cache
bool GetFilenameHash(std::string const &Filename, std::string &Hash);
std::optional<std::string> GetFilenameHash(std::string const &Filename) const;
void AddThreadRIPsToEntryList(FEXCore::Core::InternalThreadState *Thread);
void SaveEntryList();
std::set<uint64_t> EntryList;
@@ -224,8 +236,8 @@ namespace FEXCore::Context {
std::unique_ptr<GdbServer> DebugServer;
bool StartPaused = false;
FEXCore::Config::Value<std::string> AppFilename{FEXCore::Config::CONFIG_APP_FILENAME, ""};
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
};
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::InternalThreadState *Thread, FEXCore::HLE::SyscallArguments *Args);
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args);
}
@@ -38,8 +38,8 @@ static bool StoreCAS8(uint8_t &Expected, uint8_t Val, uint64_t Addr) {
return Atom->compare_exchange_strong(Expected, Val);
}
bool HandleCASPAL(void *_mcontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = reinterpret_cast<mcontext_t*>(_mcontext);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -827,8 +827,8 @@ std::tuple<uint64_t, bool> DoCAS64(
}
bool HandleCASAL(void *_mcontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = reinterpret_cast<mcontext_t*>(_mcontext);
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -921,8 +921,8 @@ bool HandleCASAL(void *_mcontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleAtomicMemOp(void *_mcontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = reinterpret_cast<mcontext_t*>(_mcontext);
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
+3 -3
View File
@@ -24,7 +24,7 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t ATOMIC_UMIN_OP = 0b0111;
constexpr uint32_t ATOMIC_SWAP_OP = 0b1000;
bool HandleCASPAL(void *_mcontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_mcontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_mcontext, void *_info, uint32_t Instr);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr);
}
@@ -0,0 +1,220 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include "aarch64/cpu-aarch64.h"
namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
if (SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
SetCPUFeatures(Features);
if (!SupportsAtomics) {
WARN_ONCE("Host CPU doesn't support atomics. Expect bad performance");
}
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
if (Is64Bit && ((~Constant)>> 16) == 0) {
movn(Reg, (~Constant) & 0xFFFF);
return;
}
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
}
}
}
void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
{x25, x26},
{x27, x28},
{x29, x30},
}};
for (auto &RegPair : CalleeSaved) {
stp(RegPair.first, RegPair.second, PairOffset);
}
// Additionally we need to store the lower 64bits of v8-v15
// Here's a fun thing, we can use two ST4 instructions to store everything
// We just need a single sub to sp before that
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v8, v9, v10, v11},
{v12, v13, v14, v15},
}};
uint32_t VectorSaveSize = sizeof(uint64_t) * 8;
sub(sp, sp, VectorSaveSize);
// SP supporting move
// We just saved x19 so it is safe
add(x19, sp, 0);
MemOperand QuadOffset(x19, 32, PostIndex);
for (auto &RegQuad : FPRs) {
st4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
}
void Arm64Emitter::PopCalleeSavedRegisters() {
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v12, v13, v14, v15},
{v8, v9, v10, v11},
}};
MemOperand QuadOffset(sp, 32, PostIndex);
for (auto &RegQuad : FPRs) {
ld4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
{x23, x24},
{x21, x22},
{x19, x20},
}};
for (auto &RegPair : CalleeSaved) {
ldp(RegPair.first, RegPair.second, PairOffset);
}
}
void Arm64Emitter::SpillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
void Arm64Emitter::FillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
void Arm64Emitter::PushDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
sub(sp, sp, SPOffset);
int i = 0;
for (auto RA : RAFPR)
{
str(RA.Q(), MemOperand(sp, i * 8));
i+=2;
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
str(RA, MemOperand(sp, i * 8));
i++;
}
#endif
str(lr, MemOperand(sp, i * 8));
}
void Arm64Emitter::PopDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
int i = 0;
for (auto RA : RAFPR)
{
ldr(RA.Q(), MemOperand(sp, i * 8));
i+=2;
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
ldr(RA, MemOperand(sp, i * 8));
i++;
}
#endif
ldr(lr, MemOperand(sp, i * 8));
add(sp, sp, SPOffset);
}
void Arm64Emitter::ResetStack() {
if (SpillSlots == 0)
return;
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetBuffer()->GetOffsetAddress<uint64_t>(GetCursorOffset());
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
nop();
}
}
}
@@ -0,0 +1,74 @@
#pragma once
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
// All but x29 are caller saved
const std::array<aarch64::Register, 16> SRA64 = {
x4, x5, x6, x7, x8, x9, x10, x11,
x12, x18, x17, x16, x15, x14, x13, x29
};
// All are callee saved
const std::array<aarch64::Register, 9> RA64 = {
x20, x21, x22, x23, x24, x25, x26, x27,
x19
};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA64Pair = {{
{x20, x21},
{x22, x23},
{x24, x25},
{x26, x27},
}};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA32Pair = {{
{w20, w21},
{w22, w23},
{w24, w25},
{w26, w27},
}};
// All are caller saved
const std::array<aarch64::VRegister, 16> SRAFPR = {
v16, v17, v18, v19, v20, v21, v22, v23,
v24, v25, v26, v27, v28, v29, v30, v31
};
// v8..v15 = (lower 64bits) Callee saved
const std::array<aarch64::VRegister, 12> RAFPR = {
/*v0, v1, v2, v3,*/v4, v5, v6, v7, // v0 ~ v3 are used as temps
v8, v9, v10, v11, v12, v13, v14, v15
};
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(size_t size);
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs();
void FillStaticRegs();
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void ResetStack();
void Align16B();
uint32_t SpillSlots{};
};
}
@@ -0,0 +1,26 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore::ArchHelpers::Arm64 {
#ifndef _M_ARM_64
// These are stub implementations that exist only to allow instantiating the arm64 jit
// on non arm platforms.
// Obvously such a configuration can't do the actual arm64-specific stuff
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASPAL Not Implemented");
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASAL Not Implemented");
}
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleAtomicMemOp Not Implemented");
}
#endif
}
@@ -0,0 +1,212 @@
#pragma once
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/UContext.h>
#include <signal.h>
#include <string.h>
#include <ucontext.h>
#include <stdint.h>
#include <type_traits>
namespace FEXCore::ArchHelpers::Context {
struct X86ContextBackup {
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
static constexpr int RedZoneSize = 128;
};
struct ArmContextBackup {
// Host State
uint64_t GPRs[31];
uint64_t PrevSP;
uint64_t PrevPC;
uint64_t PState;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
// Arm64 doesn't have a red zone
static constexpr int RedZoneSize = 0;
};
static inline mcontext_t* GetMContext(void* ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
return &_context->uc_mcontext;
}
#ifdef _M_ARM_64
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->sp;
}
static inline uint64_t GetPc(void* ucontext) {
return GetMContext(ucontext)->pc;
}
static inline void SetSp(void* ucontext, uint64_t val) {
GetMContext(ucontext)->sp = val;
}
static inline void SetPc(void* ucontext, uint64_t val) {
GetMContext(ucontext)->pc = val;
}
static inline uint64_t GetState(void* ucontext) {
return GetMContext(ucontext)->regs[28];
}
static inline void SetState(void* ucontext, uint64_t val) {
GetMContext(ucontext)->regs[28] = val;
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
return GetMContext(ucontext)->regs[id];
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
GetMContext(ucontext)->regs[id] = val;
}
constexpr uint32_t FPR_MAGIC = 0x46508001U;
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
using ContextBackup = ArmContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
memcpy(&Backup->GPRs[0], &_mcontext->regs[0], 31 * sizeof(uint64_t));
Backup->PrevSP = ArchHelpers::Context::GetSp(ucontext);
Backup->PrevPC = ArchHelpers::Context::GetPc(ucontext);
Backup->PState = _mcontext->pstate;
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Backup->FPRs[0], 32 * sizeof(__uint128_t));
HostState->FPCR = Backup->FPCR;
HostState->FPSR = Backup->FPSR;
// Restore GPRs and other state
_mcontext->pstate = Backup->PState;
ArchHelpers::Context::SetPc(ucontext, Backup->PrevPC);
ArchHelpers::Context::SetSp(ucontext, Backup->PrevSP);
memcpy(&_mcontext->regs[0], &Backup->GPRs[0], 31 * sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
#endif
#ifdef _M_X86_64
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_RSP];
}
static inline uint64_t GetPc(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_RIP];
}
static inline void SetSp(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_RSP] = val;
}
static inline void SetPc(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_RIP] = val;
}
static inline uint64_t GetState(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_R14];
}
static inline void SetState(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_R14] = val;
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not impelented for x86 host");
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
ERROR_AND_DIE("Not impelented for x86 host");
}
using ContextBackup = X86ContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
memcpy(&Backup->GPRs[0], &_mcontext->gregs[0], sizeof(X86ContextBackup::GPRs));
// Copy the FPRState
memcpy(&Backup->FPRState, _mcontext->fpregs, sizeof(X86ContextBackup::FPRState));
// XXX: Save 256bit and 512bit AVX register state
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Backup->GPRs[0], sizeof(X86ContextBackup::GPRs));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Backup->FPRState, sizeof(X86ContextBackup::FPRState));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
#endif
} // namespace FEXCore::ArchHelpers::Context
+124 -1
View File
@@ -1,11 +1,42 @@
/*
$info$
tags: opcodes|cpuid
desc: Handles presented capability bits for guest cpu
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "git_version.h"
#include <cstring>
#ifdef _M_X86_64
#include <cpuid.h>
#endif
namespace FEXCore {
//#define CPUID_AMD
#ifdef _M_ARM_64
static uint32_t GetCycleCounterFrequency() {
uint64_t Result{};
__asm("mrs %[Res], CNTFRQ_EL0"
: [Res] "=r" (Result));
return Result;
}
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x15) {
__cpuid(0x15, eax, ebx, ecx, edx);
if (eax && ebx && ecx) {
return ecx * ebx / eax;
}
}
return 0;
}
#endif
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h() {
FEXCore::CPUID::FunctionResults Res{};
@@ -251,6 +282,18 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h() {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h() {
FEXCore::CPUID::FunctionResults Res{};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1;
Res.ecx = FrequencyHz;
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h() {
FEXCore::CPUID::FunctionResults Res{};
@@ -375,10 +418,81 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h() {
return Res;
}
// L1 Cache and TLB identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0005h() {
FEXCore::CPUID::FunctionResults Res{};
// L1 TLB Information for 2MB and 4MB pages
Res.eax =
(64 << 0) | // Number of TLB instruction entries
(255 << 8) | // instruction TLB associativity type (full)
(64 << 16) | // Number of TLB data entries
(255 << 24); // data TLB associativity type (full)
// L1 TLB Information for 4KB pages
Res.ebx =
(64 << 0) | // Number of TLB instruction entries
(255 << 8) | // instruction TLB associativity type (full)
(64 << 16) | // Number of TLB data entries
(255 << 24); // data TLB associativity type (full)
// L1 data cache identifiers
Res.ecx =
(64 << 0) | // L1 data cache size line in bytes
(1 << 8) | // L1 data cachelines per tag
(8 << 16) | // L1 data cache associativity
(32 << 24); // L1 data cache size in KB
// L1 instruction cache identifiers
Res.edx =
(64 << 0) | // L1 instruction cache line size in bytes
(1 << 8) | // L1 instruction cachelines per tag
(4 << 16) | // L1 instruction cache associativity
(64 << 24); // L1 instruction cache size in KB
return Res;
}
// L2 Cache identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0006h() {
FEXCore::CPUID::FunctionResults Res{};
// L2 TLB Information for 2MB and 4MB pages
Res.eax =
(1024 << 0) | // Number of TLB instruction entries
(6 << 12) | // instruction TLB associativity type
(1536 << 16) | // Number of TLB data entries
(3 << 28); // data TLB associativity type
// L2 TLB Information for 4KB pages
Res.ebx =
(1024 << 0) | // Number of TLB instruction entries
(6 << 12) | // instruction TLB associativity type
(1536 << 16) | // Number of TLB data entries
(5 << 28); // data TLB associativity type
// L2 cache identifiers
Res.ecx =
(64 << 0) | // cacheline size
(1 << 8) | // cachelines per tag
(6 << 12) | // cache associativity
(512 << 16); // L2 cache size in KB
// L3 cache identifiers
Res.edx =
(64 << 0) | // cacheline size
(1 << 8) | // cachelines per tag
(6 << 12) | // cache associativity
(16 << 18); // L2 cache size in KB
return Res;
}
// Advanced power management
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0007h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // APIC timer not affected by p-state
Res.edx =
(1 << 8); // Invariant TSC
return Res;
}
@@ -408,7 +522,11 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x12: Intel SGX capability enumeration
// 0x13: Reserved
// 0x14: Intel Processor trace
// 0x15: Timestamp counter information
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
RegisterFunction(0x15, std::bind(&CPUIDEmu::Function_15h, this));
#endif
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
@@ -423,7 +541,12 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// Processor brand string continued
RegisterFunction(0x8000'0004, std::bind(&CPUIDEmu::Function_8000_0004h, this));
// 0x8000'0005: L1 Cache and TLB identifiers
#ifdef CPUID_AMD
// This is full reserved on Intel platforms
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_8000_0005h, this));
#endif
// 0x8000'0006: L2 Cache identifiers
RegisterFunction(0x8000'0006, std::bind(&CPUIDEmu::Function_8000_0006h, this));
// Advanced power management information
RegisterFunction(0x8000'0007, std::bind(&CPUIDEmu::Function_8000_0007h, this));
// 0x8000'0008: Virtual and physical address sizes
+4 -1
View File
@@ -23,7 +23,7 @@ private:
public:
void Init(FEXCore::Context::Context *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function) {
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, [[maybe_unused]] uint32_t Leaf) {
auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end()) {
@@ -51,11 +51,14 @@ private:
FEXCore::CPUID::FunctionResults Function_02h();
FEXCore::CPUID::FunctionResults Function_06h();
FEXCore::CPUID::FunctionResults Function_07h();
FEXCore::CPUID::FunctionResults Function_15h();
FEXCore::CPUID::FunctionResults Function_8000_0000h();
FEXCore::CPUID::FunctionResults Function_8000_0001h();
FEXCore::CPUID::FunctionResults Function_8000_0002h();
FEXCore::CPUID::FunctionResults Function_8000_0003h();
FEXCore::CPUID::FunctionResults Function_8000_0004h();
FEXCore::CPUID::FunctionResults Function_8000_0005h();
FEXCore::CPUID::FunctionResults Function_8000_0006h();
FEXCore::CPUID::FunctionResults Function_8000_0007h();
FEXCore::CPUID::FunctionResults Function_Reserved();
+11 -7
View File
@@ -5,6 +5,12 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore {
static void* ThreadHandler(void *Arg) {
FEXCore::CompileService *This = reinterpret_cast<FEXCore::CompileService*>(Arg);
This->ExecutionThread();
return nullptr;
}
CompileService::CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ParentThread {Thread} {
@@ -16,9 +22,7 @@ namespace FEXCore {
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
WorkerThread = std::thread([this]() {
ExecutionThread();
});
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
}
void CompileService::Initialize() {
@@ -30,7 +34,7 @@ namespace FEXCore {
ShuttingDown = true;
// Kick the working thread
StartWork.NotifyAll();
WorkerThread.join();
WorkerThread->join(nullptr);
}
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread) {
@@ -92,7 +96,7 @@ namespace FEXCore {
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->State.ThreadManager.TID.load());
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
pthread_setname_np(pthread_self(), ThreadName);
while (true) {
@@ -121,10 +125,10 @@ namespace FEXCore {
if (Item) {
// Make sure it's not in lookup cache by accident
LogMan::Throw::A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->State.State.rip = Item->RIP;
CompileThreadData->CurrentFrame->State.rip = Item->RIP;
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
+5 -2
View File
@@ -2,6 +2,7 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <memory>
#include <thread>
@@ -45,12 +46,14 @@ class CompileService final {
WorkItem *CompileCode(uint64_t RIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread);
// Public for threading
void ExecutionThread();
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
void ExecutionThread();
std::thread WorkerThread;
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::unique_ptr<FEXCore::Core::InternalThreadState> CompileThreadData;
std::mutex QueueMutex{};
+212 -129
View File
@@ -1,3 +1,12 @@
/*
$info$
category: glue ~ Logic that binds various parts together
meta: glue|driver ~ Emulation mainloop related glue logic
tags: glue|driver
desc: Glues Frontend, OpDispatcher and IR Opts & Compilation, LookupCache, Dispatcher and provides the Execution loop entrypoint
$end_info$
*/
#include "Common/MathUtils.h"
#include "Common/Paths.h"
@@ -25,6 +34,7 @@
#include <fstream>
#include <unistd.h>
#include <filesystem>
#include <algorithm>
#include "Interface/Core/GdbServer.h"
@@ -55,12 +65,12 @@ namespace {
v = 0;
switch (len & 7) {
case 7: v ^= (uint64_t)pos2[6] << 48;
case 6: v ^= (uint64_t)pos2[5] << 40;
case 5: v ^= (uint64_t)pos2[4] << 32;
case 4: v ^= (uint64_t)pos2[3] << 24;
case 3: v ^= (uint64_t)pos2[2] << 16;
case 2: v ^= (uint64_t)pos2[1] << 8;
case 7: v ^= (uint64_t)pos2[6] << 48; [[fallthrough]];
case 6: v ^= (uint64_t)pos2[5] << 40; [[fallthrough]];
case 5: v ^= (uint64_t)pos2[4] << 32; [[fallthrough]];
case 4: v ^= (uint64_t)pos2[3] << 24; [[fallthrough]];
case 3: v ^= (uint64_t)pos2[2] << 16; [[fallthrough]];
case 2: v ^= (uint64_t)pos2[1] << 8; [[fallthrough]];
case 1: v ^= (uint64_t)pos2[0];
h ^= mix(v);
h *= m;
@@ -158,7 +168,7 @@ namespace DefaultFallbackCore {
bool NeedsOpDispatch() override { return false; }
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override {
LogMan::Msg::E("Fell back to default code handler at RIP: 0x%lx", ThreadState->State.State.rip);
LogMan::Msg::E("Fell back to default code handler at RIP: 0x%lx", ThreadState->CurrentFrame->State.rip);
return nullptr;
}
@@ -175,29 +185,32 @@ namespace DefaultFallbackCore {
namespace FEXCore::Context {
Context::Context() {
FallbackCPUFactory = FEXCore::Core::DefaultFallbackCore::CPUCreationFactory;
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
if (Config.GdbServer) {
StartGdbServer();
}
else {
StopGdbServer();
}
}
bool Context::GetFilenameHash(std::string const &Filename, std::string &Hash) {
std::optional<std::string> Context::GetFilenameHash(std::string const &Filename) const {
// Calculate a hash for the input file
std::ifstream Input (Filename.c_str(), std::ios::in | std::ios::binary | std::ios::ate);
if (Input.is_open()) {
std::streampos Size;
Size = Input.tellg();
Input.seekg(0, std::ios::beg);
std::string Data;
Data.resize(Size);
Input.read(&Data.at(0), Size);
Input.close();
std::hash<std::string> string_hash;
Hash = std::to_string(string_hash(Data));
return true;
std::ifstream Input(Filename, std::ios::in | std::ios::binary | std::ios::ate);
if (!Input) {
return std::nullopt;
}
return false;
const auto Size = static_cast<size_t>(Input.tellg());
Input.seekg(0, std::ios::beg);
std::string Data(Size, '\0');
Input.read(Data.data(), Size);
Input.close();
std::hash<std::string> string_hash;
return std::to_string(string_hash(Data));
}
void Context::AddThreadRIPsToEntryList(FEXCore::Core::InternalThreadState *Thread) {
@@ -208,45 +221,48 @@ namespace FEXCore::Context {
void Context::SaveEntryList() {
std::string const &Filename = AppFilename();
std::string hash_string;
if (GetFilenameHash(Filename, hash_string)) {
if (auto const hash = GetFilenameHash(Filename)) {
auto DataPath = FEXCore::Paths::GetEntryCachePath();
DataPath += "Entries_" + hash_string;
DataPath += "Entries_" + *hash;
std::ofstream Output (DataPath.c_str(), std::ios::out | std::ios::binary);
if (Output.is_open()) {
for (auto Entry : EntryList) {
Output.write(reinterpret_cast<char const*>(&Entry), sizeof(Entry));
}
Output.close();
std::ofstream Output(DataPath, std::ios::out | std::ios::binary);
if (!Output) {
return;
}
for (auto Entry : EntryList) {
Output.write(reinterpret_cast<char const*>(&Entry), sizeof(Entry));
}
}
}
void Context::LoadEntryList() {
std::string const &Filename = AppFilename();
std::string hash_string;
if (GetFilenameHash(Filename, hash_string)) {
if (auto const hash = GetFilenameHash(Filename)) {
auto DataPath = FEXCore::Paths::GetEntryCachePath();
DataPath += "Entries_" + hash_string;
DataPath += "Entries_" + *hash;
std::ifstream Input (DataPath.c_str(), std::ios::in | std::ios::binary | std::ios::ate);
if (Input.is_open()) {
std::streampos Size;
Size = Input.tellg();
Input.seekg(0, std::ios::beg);
std::string Data;
Data.resize(Size);
Input.read(&Data.at(0), Size);
Input.close();
size_t EntryCount = Size / sizeof(uint64_t);
uint64_t *Entries = reinterpret_cast<uint64_t*>(&Data.at(0));
std::ifstream Input(DataPath, std::ios::in | std::ios::binary | std::ios::ate);
if (!Input) {
return;
}
for (size_t i = 0; i < EntryCount; ++i) {
EntryList.insert(Entries[i]);
}
auto const Size = static_cast<size_t>(Input.tellg());
Input.seekg(0, std::ios::beg);
std::string Data(Size, '\0');
if (!Input.read(Data.data(), Size)) {
return;
}
Input.close();
size_t const EntryCount = Size / sizeof(uint64_t);
for (size_t i = 0; i < EntryCount; ++i) {
uint64_t Entry = 0;
std::memcpy(&Entry, &Data[i * sizeof(Entry)], sizeof(Entry));
EntryList.insert(Entry);
}
}
}
@@ -254,8 +270,8 @@ namespace FEXCore::Context {
Context::~Context() {
{
for (auto &Thread : Threads) {
if (Thread->ExecutionThread.joinable()) {
Thread->ExecutionThread.join();
if (Thread->ExecutionThread->joinable()) {
Thread->ExecutionThread->join(nullptr);
}
}
@@ -313,12 +329,12 @@ namespace FEXCore::Context {
Loader->MapMemoryRegion();
Thread->State.State.gregs[X86State::REG_RSP] = Loader->SetupStack();
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = Loader->SetupStack();
Loader->LoadMemory();
Loader->GetInitLocations(&InitLocations);
Thread->State.State.rip = StartingRIP = Loader->DefaultRIP();
Thread->CurrentFrame->State.rip = StartingRIP = Loader->DefaultRIP();
InitializeThreadData(Thread);
@@ -338,7 +354,7 @@ namespace FEXCore::Context {
void Context::HandleCallback(uint64_t RIP) {
auto Thread = Core::ThreadData.Thread;
Thread->CPUBackend->CallbackPtr(Thread, RIP);
Thread->CPUBackend->CallbackPtr(Thread->CurrentFrame, RIP);
}
void Context::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func) {
@@ -382,9 +398,9 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE);
if (Thread->State.RunningEvents.Running.load()) {
if (Thread->RunningEvents.Running.load()) {
// Only attempt to stop this thread if it is running
tgkill(Thread->State.ThreadManager.PID, Thread->State.ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
}
@@ -403,7 +419,7 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN);
Thread->State.RunningEvents.WaitingToStart.store(true);
Thread->RunningEvents.WaitingToStart.store(true);
}
for (auto &Thread : Threads) {
@@ -456,17 +472,17 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
if (IgnoreCurrentThread &&
Thread->State.ThreadManager.TID == tid) {
Thread->ThreadManager.TID == tid) {
// If we are callign stop from the current thread then we can ignore sending signals to this thread
// This means that this thread is already gone
continue;
}
else if (Thread->State.ThreadManager.TID == tid) {
else if (Thread->ThreadManager.TID == tid) {
// We need to save the current thread for last to ensure all threads receive their stop signals
CurrentThread = Thread;
continue;
}
if (Thread->State.RunningEvents.Running.load()) {
if (Thread->RunningEvents.Running.load()) {
StopThread(Thread);
} else {
LogMan::Msg::D("Skipping thread %p: Already stopped", Thread);
@@ -481,16 +497,16 @@ namespace FEXCore::Context {
}
void Context::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->State.RunningEvents.Running.exchange(false)) {
if (Thread->RunningEvents.Running.exchange(false)) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::SIGNALEVENT_STOP);
tgkill(Thread->State.ThreadManager.PID, Thread->State.ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
void Context::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->State.RunningEvents.Running.load()) {
if (Thread->RunningEvents.Running.load()) {
Thread->SignalReason.store(Event);
tgkill(Thread->State.ThreadManager.PID, Thread->State.ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
@@ -534,11 +550,26 @@ namespace FEXCore::Context {
LogMan::Msg::D("Done", EntryList.size());
}
struct ExecutionThreadHandler {
FEXCore::Context::Context *This;
FEXCore::Core::InternalThreadState *Thread;
};
static void *ThreadHandler(void* Data) {
ExecutionThreadHandler *Handler = reinterpret_cast<ExecutionThreadHandler*>(Data);
Handler->This->ExecutionThread(Handler->Thread);
free(Handler);
return nullptr;
}
void Context::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
InitializeThreadData(Thread);
// This will create the execution thread but it won't actually start executing
Thread->ExecutionThread = std::thread(&Context::ExecutionThread, this, Thread);
ExecutionThreadHandler *Arg = reinterpret_cast<ExecutionThreadHandler*>(malloc(sizeof(ExecutionThreadHandler)));
Arg->This = this;
Arg->Thread = Thread;
Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
// Wait for the thread to have started
Thread->ThreadWaiting.Wait();
@@ -558,7 +589,7 @@ namespace FEXCore::Context {
State->PassManager->RegisterExitHandler([this]() {
Stop(false /* Ignore current thread */);
});
#if _M_ARM_64
bool DoSRA = true;
#else
@@ -579,10 +610,18 @@ namespace FEXCore::Context {
break;
case FEXCore::Config::CONFIG_IRJIT:
State->PassManager->InsertRegisterAllocationPass(DoSRA);
State->CPUBackend.reset(FEXCore::CPU::CreateJITCore(this, State, CompileThread));
#if (_M_X86_64 && JIT_X86_64)
State->CPUBackend.reset(FEXCore::CPU::CreateX86JITCore(this, State, CompileThread));
#elif (_M_ARM_64 && JIT_ARM64)
State->CPUBackend.reset(FEXCore::CPU::CreateArm64JITCore(this, State, CompileThread));
#else
ERROR_AND_DIE("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM: State->CPUBackend.reset(CustomCPUFactory(this, &State->State)); break;
default: LogMan::Msg::A("Unknown core configuration");
case FEXCore::Config::CONFIG_CUSTOM: State->CPUBackend.reset(CustomCPUFactory(this, State)); break;
default: ERROR_AND_DIE("Unknown core configuration");
}
}
@@ -593,20 +632,76 @@ namespace FEXCore::Context {
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
Thread = Threads.emplace_back(new FEXCore::Core::InternalThreadState{});
Thread->State.ThreadManager.TID = ++ThreadID;
Thread->ThreadManager.TID = ++ThreadID;
}
// Copy over the new thread state to the new object
memcpy(&Thread->State.State, NewThreadState, sizeof(FEXCore::Core::CPUState));
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
Thread->CurrentFrame->Thread = Thread;
// Set up the thread manager state
Thread->State.ThreadManager.parent_tid = ParentTID;
Thread->ThreadManager.parent_tid = ParentTID;
InitializeCompiler(Thread, false);
return Thread;
}
void Context::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
// remove new thread object
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
auto It = std::find(Threads.begin(), Threads.end(), Thread);
LogMan::Throw::A(It != Threads.end(), "Thread wasn't in Threads");
Threads.erase(It);
}
if (Thread->ExecutionThread &&
Thread->ExecutionThread->IsSelf()) {
// To be able to delete a thread from itself, we need to detached the std::thread object
Thread->ExecutionThread->detach();
}
delete Thread;
}
void Context::CleanupAfterFork(FEXCore::Core::InternalThreadState *LiveThread) {
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : Threads) {
if (DeadThread == LiveThread) {
continue;
}
// Setting running to false ensures that when they are shutdown we won't send signals to kill them
DeadThread->RunningEvents.Running = false;
// Despite what google searches may susgest, glibc actually has special code to handle forks
// with multiple active threads.
// It cleans up the stacks of dead threads and marks them as terminated.
// It also cleans up a bunch of internal mutexes.
// FIXME: TLS is probally still alive. Investigate
// Deconstructing the Interneal thread state should clean up most of the state.
// But if anything on the now deleted stack is holding a refrence to the heap, it will be leaked
delete DeadThread;
// FIXME: Make sure sure nothing gets leaked via the heap. Ideas:
// * Make sure nothing is allocated on the heap without ref in InternalThreadState
// * Surround any code that heap allocates with a per-thread mutex.
// Before forking, the the forking thread can lock all thread mutexes.
}
// Remove all threads but the live thread from Threads
Threads.clear();
Threads.push_back(LiveThread);
// We now only have one thread
IdleWaitRefCount = 1;
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length) {
Thread->LookupCache->AddBlockMapping(Address, Ptr, Start, Length);
}
@@ -633,10 +728,6 @@ namespace FEXCore::Context {
uint64_t TotalInstructionsLength {0};
if (!Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP)) {
if (Config.BreakOnFrontendFailure) {
LogMan::Msg::E("Had Frontend decoder error");
Stop(false /* Ignore Current Thread */);
}
return { nullptr, nullptr, 0, 0, 0, 0 };
}
@@ -669,6 +760,7 @@ namespace FEXCore::Context {
TableInfo = Block.DecodedInstructions[i].TableInfo;
DecodedInfo = &Block.DecodedInstructions[i];
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
@@ -684,7 +776,7 @@ namespace FEXCore::Context {
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_RemoveCodeEntry();
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_Constant(Block.Entry + BlockInstructionsLength));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
@@ -693,15 +785,13 @@ namespace FEXCore::Context {
if (TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
Thread->OpDispatcher->HandledLock = false;
std::invoke(Fn, Thread->OpDispatcher, DecodedInfo);
if (Thread->OpDispatcher->HadDecodeFailure()) {
if (Config.BreakOnFrontendFailure) {
LogMan::Msg::E("Had OpDispatcher error at 0x%lx", GuestRIP);
Stop(false /* Ignore Current Thread */);
}
HadDispatchError = true;
}
else {
LogMan::Throw::A(Thread->OpDispatcher->HandledLock == IsLocked, "Missing LOCK HANDLER at 0x%lx{'%s'}\n", Block.Entry + BlockInstructionsLength, TableInfo->Name);
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
++TotalInstructions;
@@ -740,15 +830,15 @@ namespace FEXCore::Context {
FILE* f = nullptr;
bool CloseAfter = false;
if (Thread->CTX->Config.DumpIR=="stderr") {
if (Thread->CTX->Config.DumpIR() =="stderr") {
f = stderr;
}
else if (Thread->CTX->Config.DumpIR=="stdout") {
else if (Thread->CTX->Config.DumpIR() =="stdout") {
f = stdout;
}
else {
std::stringstream fileName;
fileName << Thread->CTX->Config.DumpIR << "/" << std::hex << GuestRIP << (RA ? "-post.ir" : "-pre.ir");
fileName << Thread->CTX->Config.DumpIR() << "/" << std::hex << GuestRIP << (RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.str().c_str(), "w");
CloseAfter = true;
@@ -766,7 +856,7 @@ namespace FEXCore::Context {
}
};
if (Thread->CTX->Config.DumpIR != "no") {
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(nullptr);
}
@@ -786,8 +876,8 @@ namespace FEXCore::Context {
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
printf("one:\n %s\n", out.str().c_str());
printf("two:\n %s\n", out2.str().c_str());
LogMan::Msg::I("one:\n %s", out.str().c_str());
LogMan::Msg::I("two:\n %s", out2.str().c_str());
LogMan::Msg::A("Parsed ir doesn't match\n");
}
delete reparsed;
@@ -796,7 +886,7 @@ namespace FEXCore::Context {
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(Thread->OpDispatcher.get());
if (Thread->CTX->Config.DumpIR != "no") {
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->GetAllocationData() : nullptr);
}
@@ -804,7 +894,7 @@ namespace FEXCore::Context {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->GetAllocationData() : nullptr);
printf("IR 0x%lx:\n%s\n@@@@@\n", GuestRIP, out.str().c_str());
LogMan::Msg::I("IR 0x%lx:\n%s\n@@@@@\n", GuestRIP, out.str().c_str());
}
auto RAData = Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->PullAllocationData() : nullptr;
@@ -850,7 +940,7 @@ namespace FEXCore::Context {
}
auto AOTEntry = Mod->find(GuestRIP - file->second.Start + file->second.Offset);
if (AOTEntry != Mod->end()) {
// verify hash
auto MappedStart = AOTEntry->second.start + file->second.Start - file->second.Offset;
@@ -935,7 +1025,7 @@ namespace FEXCore::Context {
stream.read((char*)&addr, sizeof(addr));
if (!stream)
return false;
stream.read((char*)&start, sizeof(start));
if (!stream)
return false;
@@ -958,9 +1048,9 @@ namespace FEXCore::Context {
}
IR::RegisterAllocationData *RAData = (IR::RegisterAllocationData *)malloc(IR::RegisterAllocationData::Size(RASize));
RAData->MapCount = RASize;
stream.read((char*)&RAData->Map[0], sizeof(RAData->Map[0]) * RASize);
if (!stream) {
delete IR;
return false;
@@ -1024,8 +1114,9 @@ namespace FEXCore::Context {
}
uintptr_t Context::CompileBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
// Is the code in the cache?
// The backends only check L1 and L2, not L3
if (auto HostCode = Thread->LookupCache->FindBlock(GuestRIP)) {
@@ -1098,12 +1189,12 @@ namespace FEXCore::Context {
// Add to AOT cache if aot generation is enabled
if (Config.AOTIRCapture && RAData) {
std::lock_guard<std::mutex> lk(AOTIRCacheLock);
RAData->IsShared = true;
IRList->IsShared = true;
auto hash = fasthash64((void*)StartAddr, Length, 0);
auto file = AddrToFile.lower_bound(StartAddr);
if (file != AddrToFile.begin()) {
--file;
@@ -1121,11 +1212,6 @@ namespace FEXCore::Context {
AddBlockMapping(Thread, GuestRIP, CodePtr, StartAddr, Length);
return (uintptr_t)CodePtr;
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
return 0;
}
void Context::ExecutionThread(FEXCore::Core::InternalThreadState *Thread) {
@@ -1133,14 +1219,14 @@ namespace FEXCore::Context {
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
// Let's do some initial bookkeeping here
Thread->State.ThreadManager.TID = ::gettid();
Thread->State.ThreadManager.PID = ::getpid();
Thread->ThreadManager.TID = ::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
++IdleWaitRefCount;
LogMan::Msg::D("[%d] Waiting to run", Thread->State.ThreadManager.TID.load());
LogMan::Msg::D("[%d] Waiting to run", Thread->ThreadManager.TID.load());
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
@@ -1150,25 +1236,25 @@ namespace FEXCore::Context {
Thread->StartRunning.Wait();
}
LogMan::Msg::D("[%d] Running", Thread->State.ThreadManager.TID.load());
LogMan::Msg::D("[%d] Running", Thread->ThreadManager.TID.load());
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
Thread->State.RunningEvents.Running = true;
Thread->RunningEvents.Running = true;
Thread->CPUBackend->ExecuteDispatch(Thread);
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->State.RunningEvents.WaitingToStart = false;
Thread->State.RunningEvents.Running = false;
Thread->RunningEvents.WaitingToStart = false;
Thread->RunningEvents.Running = false;
// If it is the parent thread that died then just leave
// XXX: This doesn't make sense when the parent thread doesn't outlive its children
if (Thread->State.ThreadManager.parent_tid == 0) {
if (Thread->ThreadManager.parent_tid == 0) {
CoreShuttingDown.store(true);
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
if (CustomExitHandler) {
CustomExitHandler(Thread->State.ThreadManager.TID, Thread->ExitReason);
CustomExitHandler(Thread->ThreadManager.TID, Thread->ExitReason);
}
}
@@ -1176,6 +1262,11 @@ namespace FEXCore::Context {
IdleWaitCV.notify_all();
SignalDelegation->UninstallTLSState(Thread);
// If the parent thread is waiting to join, then we can't destroy our thread object
if (!Thread->DestroyedByParent && Thread != Thread->CTX->ParentThread) {
Thread->CTX->DestroyThread(Thread);
}
}
void FlushCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
@@ -1183,9 +1274,9 @@ namespace FEXCore::Context {
if (Thread->CTX->Config.SMCChecks == FEXCore::Config::CONFIG_SMC_MMAN) {
auto lower = Thread->LookupCache->CodePages.lower_bound(Start >> 12);
auto upper = Thread->LookupCache->CodePages.upper_bound((Start + Length) >> 12);
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second)
for (auto Address: it->second)
Context::RemoveCodeEntry(Thread, Address);
it->second.clear();
}
@@ -1199,16 +1290,16 @@ namespace FEXCore::Context {
// Debug interface
void Context::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
uint64_t RIPBackup = Thread->State.State.rip;
Thread->State.State.rip = RIP;
uint64_t RIPBackup = Thread->CurrentFrame->State.rip;
Thread->CurrentFrame->State.rip = RIP;
// Erase the RIP from all the storage backings if it exists
RemoveCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CompileBlock(Thread, RIP);
CompileBlock(Thread->CurrentFrame, RIP);
Thread->State.State.rip = RIPBackup;
Thread->CurrentFrame->State.rip = RIPBackup;
}
uint64_t Context::GetThreadCount() const {
@@ -1219,10 +1310,6 @@ namespace FEXCore::Context {
return &Threads[Thread]->Stats;
}
FEXCore::Core::CPUState Context::GetCPUState() {
return ParentThread->State.State;
}
bool Context::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
auto it = ParentThread->LocalIRCache.find(RIP);
if (it == ParentThread->LocalIRCache.end()) {
@@ -1243,20 +1330,16 @@ namespace FEXCore::Context {
return true;
}
FEXCore::Core::ThreadState *Context::GetThreadState() {
return &ParentThread->State;
}
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::InternalThreadState *Thread, FEXCore::HLE::SyscallArguments *Args) {
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args) {
uint64_t Result{};
Result = Handler->HandleSyscall(Thread, Args);
Result = Handler->HandleSyscall(Frame, Args);
return Result;
}
void Context::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
auto base_filename = std::filesystem::path(filename).filename().string();
if (base_filename.size()) {
auto filename_hash = fasthash64(filename.c_str(), filename.size(), 0xBAADF00D);
@@ -0,0 +1,363 @@
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
auto Buffer = GetBuffer();
DispatchPtr = Buffer->GetOffsetAddress<CPUBackend::AsmDispatch>(GetCursorOffset());
// while (true) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// Ptr();
// }
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
Literal l_VirtualMemory {VirtualMemorySize};
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_L1Ptr {Thread->LookupCache->GetL1Pointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
Literal l_ExitFunctionLink {config.ExitFunctionLink};
Literal l_ExitFunctionLinkThis {config.ExitFunctionLinkThis};
// Push all the register we need to save
PushCalleeSavedRegisters();
// Push our memory base to the correct register
// Move our thread pointer to the correct register
// This is passed in to parameter 0 (x0)
mov(STATE, x0);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
AbsoluteLoopTopAddressFillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled) {
FillStaticRegs();
}
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
aarch64::Label FullLookup{};
aarch64::Label LoopTop{};
aarch64::Label ExitSpillSRA{};
aarch64::Label ThreadPauseHandler{};
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
auto RipReg = x2;
if (!config.ExecuteBlocksWithCall) {
// L1 Cache
ldr(x0, &l_L1Ptr);
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
ldp(x1, x0, MemOperand(x0));
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
br(x1);
}
// L1C check failed, do a full lookup
bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, &l_PagePtr);
// Mask the address by the virtual address size so we can check for aliases
if (__builtin_popcountl(VirtualMemorySize) == 1) {
and_(x3, RipReg, Thread->LookupCache->GetVirtualMemorySize() - 1);
}
else {
ldr(x3, &l_VirtualMemory);
and_(x3, RipReg, x3);
}
aarch64::Label NoBlock;
{
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
// Load the pointer from the offset
ldr(x0, MemOperand(x0, x1, Shift::LSL, 3));
// If page pointer is zero then we have no block
cbz(x0, &NoBlock);
// Steal the page offset
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x3, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x3, &NoBlock);
// If we've made it here then we have a real compiled block
{
// Jump to the block
if (!config.ExecuteBlocksWithCall) {
// update L1 cache
ldr(x0, &l_L1Ptr);
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
br(x3);
} else {
mov(x0, STATE);
blr(x3);
}
}
if (config.ExecuteBlocksWithCall) {
// Interpreter continues execution here
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, &l_CTX);
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
}
else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
}
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
PopCalleeSavedRegisters();
// Return from the function
// LR is set to the correct return location now
ret();
}
{
ExitFunctionLinkerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
ldr(x3, &l_ExitFunctionLink);
blr(x3);
if (SRAEnabled)
FillStaticRegs();
br(x0);
}
// Need to create the block
{
bind(&NoBlock);
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileBlock);
if (SRAEnabled)
SpillStaticRegs();
// X2 contains our guest RIP
blr(x3); // { CTX, Frame, RIP}
if (SRAEnabled)
FillStaticRegs();
b(&LoopTop);
}
{
SignalHandlerReturnAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// Now to get back to our old location we need to do a fault dance
// We can't use SIGTRAP here since gdb catches it and never gives it to the application!
hlt(0);
}
{
ThreadPauseHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
// Call our sleep handler
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x2, &l_Sleep);
blr(x2);
PauseReturnInstruction = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// Fault to start running again
hlt(0);
}
{
// The expectation here is that a thunked function needs to call back in to the JIT in a reentrant safe way
// To do this safely we need to do some state tracking and register saving
//
// eg:
// JIT Call->
// Thunk->
// Thunk callback->
//
// The thunk callback needs to execute JIT code and when it returns, it needs to safely return to the thunk rather than JIT space
// This is handled by pushing a return address trampoline to the stack so when the guest address returns it hits our custom thunk return
// - This will safely return us to the thunk
//
// On return to the thunk, the thunk can get whatever its return value is from the thread context depending on ABI handling on its end
// When the thunk itself returns, it'll do its regular return logic there
// void ReentrantCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
CallbackPtr = Buffer->GetOffsetAddress<CPUBackend::JITCallback>(GetCursorOffset());
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
// First thing we need to move the thread state pointer back in to our register
mov(STATE, x0);
// Make sure to adjust the refcounter so we don't clear the cache now
LoadConstant(x0, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
ldr(w2, MemOperand(x0));
add(w2, w2, 1);
str(w2, MemOperand(x0));
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// load static regs
if (SRAEnabled)
FillStaticRegs();
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
place(&l_VirtualMemory);
place(&l_PagePtr);
place(&l_L1Ptr);
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
place(&l_ExitFunctionLink);
place(&l_ExitFunctionLinkThis);
FinalizeCode();
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
#endif
}
void Arm64Dispatcher::SpillSRA(void *ucontext) {
for(int i = 0; i < SRA64.size(); i++) {
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
// TODO: Also recover FPRs, not sure where the neon context is
// This is usually not needed
/*
for(int i = 0; i < SRAFPR.size(); i++) {
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
}
*/
}
#ifdef _M_ARM_64
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Dispatcher = new Arm64Dispatcher(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
}
#endif
}
@@ -0,0 +1,18 @@
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
protected:
void SpillSRA(void *ucontext) override;
};
}
@@ -0,0 +1,343 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Common/MathUtils.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <FEXCore/Core/X86Enums.h>
namespace FEXCore::CPU {
void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ArchHelpers::Context::ContextBackup);
// We need to back up behind the host's red zone
// We do this on the guest side as well
// (does nothing on arm hosts)
NewSP -= ArchHelpers::Context::ContextBackup::RedZoneSize;
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
ArchHelpers::Context::BackupContext(ucontext, Context);
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, ThreadState->CurrentFrame, sizeof(FEXCore::Core::CPUState));
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
SignalFrames.push(NewSP);
}
void Dispatcher::RestoreThreadState(void *ucontext) {
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uintptr_t NewSP = OldSP;
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(ThreadState->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
StoreThreadState(Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
uint64_t OldGuestSP = Frame->State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
if (GuestAction->sa_flags & SA_SIGINFO) {
if (SRAEnabled) {
if (!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
LogMan::Throw::A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
} else {
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
}
}
// Setup ucontext a bit
if (CTX->Config.Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::ucontext_t);
uint64_t UContextLocation = NewGuestSP;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
// We have extended float information
guest_uctx->uc_flags |= FEXCore::x86_64::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = &guest_uctx->__fpregs_mem;
#define COPY_REG(x) \
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x] = Frame->State.gregs[X86State::REG_##x];
COPY_REG(R8);
COPY_REG(R9);
COPY_REG(R10);
COPY_REG(R11);
COPY_REG(R12);
COPY_REG(R13);
COPY_REG(R14);
COPY_REG(R15);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
// Copy float registers
memcpy(guest_uctx->__fpregs_mem._st, Frame->State.mm, sizeof(Frame->State.mm));
memcpy(guest_uctx->__fpregs_mem._xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = Frame->State.FCW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
(Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] << 11) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] << 8) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] << 9) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] << 10) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] << 14);
// Copy over signal stack information
guest_uctx->uc_stack.ss_flags = GuestStack->ss_flags;
guest_uctx->uc_stack.ss_sp = GuestStack->ss_sp;
guest_uctx->uc_stack.ss_size = GuestStack->ss_size;
// XXX: siginfo_t(RSI)
Frame->State.gregs[X86State::REG_RSI] = 0x4142434445460000;
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
// XXX: 32bit Support
NewGuestSP -= sizeof(FEXCore::x86::ucontext_t);
uint64_t UContextLocation = 0; // NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::siginfo_t);
uint64_t SigInfoLocation = 0; // NewGuestSP;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = UContextLocation;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SigInfoLocation;
}
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
if (CTX->Config.Is64BitMode) {
Frame->State.gregs[X86State::REG_RDI] = Signal;
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
}
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
LogMan::Throw::A(CTX->X86CodeGen.SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
}
return true;
}
bool Dispatcher::HandleSIGILL(int Signal, void *info, void *ucontext) {
if (ArchHelpers::Context::GetPc(ucontext) == SignalHandlerReturnAddress) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
if (ArchHelpers::Context::GetPc(ucontext) == PauseReturnInstruction) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
return false;
}
bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = ThreadState->SignalReason.load();
auto Frame = ThreadState->CurrentFrame;
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LogMan::Throw::A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
}
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the core and get out safely
ArchHelpers::Context::SetSp(ucontext, Frame->ReturningStackLocation);
// Our ref counting doesn't matter anymore
SignalHandlerRefCounter = 0;
// Set the new PC
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LogMan::Throw::A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
}
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
uint64_t Dispatcher::GetCompileBlockPtr() {
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::CpuStateFrame *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast CompileBlockPtr;
CompileBlockPtr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
return CompileBlockPtr.Data;
}
void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
for (auto iter = CodeBuffers.begin(); iter != CodeBuffers.end(); ++iter) {
auto [start, end] = *iter;
if (start == reinterpret_cast<uint64_t>(start_to_remove)) {
CodeBuffers.erase(iter);
return;
}
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher) {
return IsAddressInDispatcher(Address);
}
return false;
}
}
@@ -0,0 +1,84 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "Interface/Context/Context.h"
#include <stack>
namespace FEXCore::CPU {
struct DispatcherConfig {
bool ExecuteBlocksWithCall = false;
uintptr_t ExitFunctionLink = 0;
uintptr_t ExitFunctionLinkThis = 0;
bool StaticRegisterAssignment = false;
};
class Dispatcher {
public:
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
/**
* @name Dispatch Helper functions
* @{ */
uint64_t ThreadStopHandlerAddress{};
uint64_t ThreadStopHandlerAddressSpillSRA{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t AbsoluteLoopTopAddressFillSRA{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
uint64_t Start{};
uint64_t End{};
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
void RegisterCodeBuffer(uint8_t* start, size_t size) {
CodeBuffers.emplace_back(reinterpret_cast<uint64_t>(start),
reinterpret_cast<uint64_t>(start + size));
}
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true);
bool IsAddressInDispatcher(uint64_t Address) {
return Address >= Start && Address < End;
}
protected:
Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ThreadState {Thread} {}
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
private:
std::vector<std::tuple<uint64_t, uint64_t>> CodeBuffers; // Start, End
};
}
@@ -0,0 +1,325 @@
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE) {
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
// r10, r11
//
// Callee Saved
// rbx, rbp, r12, r13, r14, r15
//
// 1St Argument: rdi <ThreadState>
// XMM:
// All temp
// while (true) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Bunch of exit state stuff
// x86-64 ABI has the stack aligned when /call/ happens
// Which means the destination has a misaligned stack at that point
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
mov(STATE, rdi);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)], rsp);
Label LoopTop;
Label FullLookup;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
L(LoopTop);
AbsoluteLoopTopAddressFillSRA = AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
if (!config.ExecuteBlocksWithCall)
{
// L1 Cache
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
jne(FullLookup);
jmp(qword[r13 + rax + 0]);
}
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
// Load page pointer
mov(rdi, qword [r13 + rax * 8]);
cmp(rdi, 0);
je(NoBlock);
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
cmp(rcx, rdx);
jne(NoBlock);
// Load the block pointer
mov(rax, qword [rdi + rax]);
cmp(rax, 0);
je(NoBlock);
// Update L1
if (config.ExecuteBlocksWithCall) {
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
mov(qword[r13 + rcx*8 + 8], rdx);
mov(qword[r13 + rcx*8 + 0], rax);
}
// Real block if we made it here
if (!config.ExecuteBlocksWithCall) {
jmp(rax);
} else {
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
}
{
L(ExitBlock);
ThreadStopHandlerAddress = getCurr<uint64_t>();
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
// Block creation
{
L(NoBlock);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::CpuStateFrame *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
mov(rax, Ptr.Data);
call(rax);
// rdx already contains RIP here
jmp(LoopTop);
}
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
// {rdi, rsi, rdx}
mov(rdi, config.ExitFunctionLinkThis);
mov(rsi, STATE);
mov(rdx, rax); // rax is set at the block end
mov(rax, config.ExitFunctionLink);
call(rax);
jmp(rax);
}
{
// Pause handler
ThreadPauseHandlerAddress = getCurr<uint64_t>();
L(ThreadPauseHandler);
mov(rdi, reinterpret_cast<uintptr_t>(CTX));
mov(rsi, STATE);
mov(rax, reinterpret_cast<uint64_t>(SleepThread));
call(rax);
// XXX: Unsupported atm
PauseReturnInstruction = getCurr<uint64_t>();
ud2();
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
// First thing we need to move the thread state pointer back in to our register
mov(STATE, rdi);
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
add(dword [rax], 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
mov(rax, CTX->X86CodeGen.CallbackReturn);
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])]);
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], rsi);
// Back to the loop top now
jmp(LoopTop);
}
{
// Signal return handler
SignalHandlerReturnAddress = getCurr<uint64_t>();
ud2();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
// rsi = rsp
mov(rsp, rsi);
// Now jump back to the thunk
// XXX: XMM?
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
ready();
Start = reinterpret_cast<uint64_t>(getCode());
End = Start + getSize();
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(Start, End-Start, Name);
#endif
}
X86Dispatcher::~X86Dispatcher() {
}
#ifdef _M_X86_64
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Dispatcher = new X86Dispatcher(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
}
#endif
}
@@ -0,0 +1,17 @@
#pragma once
#include "Interface/Core/Dispatcher/Dispatcher.h"
#define XBYAK64
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
virtual ~X86Dispatcher() override;
};
}
+24 -19
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: frontend|x86-meta-blocks
desc: Extracts instruction & block meta info, frontend multiblock logic
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/InternalThreadState.h"
@@ -129,26 +136,17 @@ uint8_t Decoder::PeekByte(uint8_t Offset) {
}
uint64_t Decoder::ReadData(uint8_t Size) {
uint64_t Res{};
#define READ_DATA(x, y) \
case x: { \
y const *Data = reinterpret_cast<y const*>(&InstStream[InstructionSize]); \
Res = *Data; \
} \
break
switch (Size) {
case 0: return 0;
READ_DATA(1, uint8_t);
READ_DATA(2, uint16_t);
case 3: memcpy(&Res, &InstStream[InstructionSize], Size);
READ_DATA(4, uint32_t);
READ_DATA(8, uint64_t);
default:
LogMan::Msg::A("Unknown data size to read");
return 0;
if (Size == 0) {
return 0;
}
#undef READ_DATA
if (Size > sizeof(uint64_t)) {
LogMan::Msg::A("Unknown data size to read");
return 0;
}
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
#ifndef NDEBUG
for(size_t i = 0; i < Size; ++i) {
@@ -157,6 +155,7 @@ uint64_t Decoder::ReadData(uint8_t Size) {
#else
SkipBytes(Size);
#endif
return Res;
}
@@ -919,6 +918,7 @@ void Decoder::BranchTargetInMultiblockRange() {
// If the RIP setting is conditional AND within our symbol range then it can be considered for multiblock
uint64_t TargetRIP = 0;
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
bool Conditional = true;
switch (DecodeInst->OP) {
@@ -946,6 +946,11 @@ void Decoder::BranchTargetInMultiblockRange() {
break;
}
if (GPRSize == 4) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
TargetRIP &= 0xFFFFFFFFU;
}
// If the target RIP is within the symbol ranges then we are golden
if (TargetRIP >= SymbolMinAddress && TargetRIP < SymbolMaxAddress) {
// Update our conditional branch ranges before we return
+22 -10
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: glue|gdbserver
desc: Provides a gdb interface to the guest state
$end_info$
*/
#include <cstdlib>
#include <cstdio>
#include <iomanip>
@@ -232,17 +239,17 @@ std::string GdbServer::readRegs() {
bool Found = false;
for (auto &Thread : *Threads) {
if (Thread->State.ThreadManager.GetTID() != CurrentDebuggingThread) {
if (Thread->ThreadManager.GetTID() != CurrentDebuggingThread) {
continue;
}
state = Thread->State.State;
memcpy(&state, Thread->CurrentFrame, sizeof(state));
Found = true;
break;
}
if (!Found) {
// If set to an invalid thread then just get the parent thread ID
state = CTX->GetCPUState();
memcpy(&state, CTX->ParentThread->CurrentFrame, sizeof(state));
}
// Encode the GDB context definition
@@ -284,17 +291,17 @@ GdbServer::HandledPacketType GdbServer::readReg(std::string& packet) {
bool Found = false;
for (auto &Thread : *Threads) {
if (Thread->State.ThreadManager.GetTID() != CurrentDebuggingThread) {
if (Thread->ThreadManager.GetTID() != CurrentDebuggingThread) {
continue;
}
state = Thread->State.State;
memcpy(&state, Thread->CurrentFrame, sizeof(state));
Found = true;
break;
}
if (!Found) {
// If set to an invalid thread then just get the parent thread ID
state = CTX->GetCPUState();
memcpy(&state, CTX->ParentThread->CurrentFrame, sizeof(state));
}
@@ -525,7 +532,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(std::string &packet) {
ss << "<threads>\n";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << "\t<thread id=\"" << std::hex << Thread->State.ThreadManager.GetTID() << "\" core=\"" << i << "\" name=\"" << getThreadName(Thread->State.ThreadManager.GetTID()) << "\">\n";
ss << "\t<thread id=\"" << std::hex << Thread->ThreadManager.GetTID() << "\" core=\"" << i << "\" name=\"" << getThreadName(Thread->ThreadManager.GetTID()) << "\">\n";
ss << "\t</thread>\n";
}
@@ -653,7 +660,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
ss << "m";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << std::hex << Thread->State.ThreadManager.TID << ",";
ss << std::hex << Thread->ThreadManager.TID << ",";
}
return {ss.str(), HandledPacketType::TYPE_ACK};
}
@@ -672,7 +679,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
if (match("qC")) {
// Returns the current Thread ID
std::ostringstream ss;
ss << "m" << std::hex << CTX->ParentThread->State.ThreadManager.TID;
ss << "m" << std::hex << CTX->ParentThread->ThreadManager.TID;
return {ss.str(), HandledPacketType::TYPE_ACK};
}
if (match("QStartNoAckMode")) {
@@ -952,9 +959,14 @@ void GdbServer::GdbServerLoop() {
}
}
}
static void* ThreadHandler(void *Arg) {
FEXCore::GdbServer *This = reinterpret_cast<FEXCore::GdbServer*>(Arg);
This->GdbServerLoop();
return nullptr;
}
void GdbServer::StartThread() {
gdbServerThread = std::thread(&GdbServer::GdbServerLoop, this);
gdbServerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
}
std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
+15 -4
View File
@@ -1,23 +1,34 @@
/*
$info$
tags: glue|gdbserver
$end_info$
*/
#pragma once
#include <mutex>
#include <thread>
#include "Interface/Context/Context.h"
#include "Common/NetStream.h"
#include <FEXCore/Utils/Threads.h>
#include <mutex>
namespace FEXCore {
class GdbServer {
public:
GdbServer(FEXCore::Context::Context *ctx);
// Public for threading
void GdbServerLoop();
private:
void Break(int signal);
std::unique_ptr<std::iostream> OpenSocket();
void StartThread();
void GdbServerLoop();
std::string ReadPacket(std::iostream &stream);
void SendPacket(std::ostream &stream, std::string packet);
@@ -50,14 +61,14 @@ private:
HandledPacketType readReg(std::string& packet);
FEXCore::Context::Context *CTX;
std::thread gdbServerThread;
std::unique_ptr<FEXCore::Threads::Thread> gdbServerThread;
std::unique_ptr<std::iostream> CommsStream;
std::mutex sendMutex;
bool SettingNoAckMode{false};
bool NoAckMode{false};
std::string ThreadString{};
uint32_t CurrentDebuggingThread{};
FEXCore::Config::Value<std::string> Filename{FEXCore::Config::CONFIG_APP_FILENAME, ""};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
};
}
@@ -1,563 +0,0 @@
#include "Common/MathUtils.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
namespace FEXCore::CPU {
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->State.RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
using namespace vixl;
using namespace vixl::aarch64;
#define STATE x28
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
class DispatchGenerator : public vixl::aarch64::Assembler {
public:
DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
uint64_t ThreadStopHandlerAddress;
uint64_t AbsoluteLoopTopAddress;
uint64_t ThreadPauseHandlerAddress;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
private:
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
std::stack<uint64_t> SignalFrames;
};
void DispatchGenerator::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
{x25, x26},
{x27, x28},
{x29, x30},
}};
for (auto &RegPair : CalleeSaved) {
stp(RegPair.first, RegPair.second, PairOffset);
}
// Additionally we need to store the lower 64bits of v8-v15
// Here's a fun thing, we can use two ST4 instructions to store everything
// We just need a single sub to sp before that
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v8, v9, v10, v11},
{v12, v13, v14, v15},
}};
uint32_t VectorSaveSize = sizeof(uint64_t) * 8;
sub(sp, sp, VectorSaveSize);
// SP supporting move
// We just saved x19 so it is safe
add(x19, sp, 0);
MemOperand QuadOffset(x19, 32, PostIndex);
for (auto &RegQuad : FPRs) {
st4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
}
void DispatchGenerator::PopCalleeSavedRegisters() {
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v12, v13, v14, v15},
{v8, v9, v10, v11},
}};
MemOperand QuadOffset(sp, 32, PostIndex);
for (auto &RegQuad : FPRs) {
ld4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
{x23, x24},
{x21, x22},
{x19, x20},
}};
for (auto &RegPair : CalleeSaved) {
ldp(RegPair.first, RegPair.second, PairOffset);
}
}
void DispatchGenerator::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
}
}
}
DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: vixl::aarch64::Assembler(MAX_DISPATCHER_CODE_SIZE, vixl::aarch64::PositionDependentCode)
, CTX {ctx}
, State {Thread} {
SetAllowAssembler(true);
auto Buffer = GetBuffer();
DispatchPtr = Buffer->GetOffsetAddress<CPUBackend::AsmDispatch>(GetCursorOffset());
// while (!Thread->State.RunningEvents.ShouldStop.load()) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Push all the register we need to save
PushCalleeSavedRegisters();
// Push our memory base to the correct register
// Move our thread pointer to the correct register
// This is passed in to parameter 0 (x0)
mov(STATE, x0);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)));
Label Exit;
Label LoopTop;
Label NoBlock;
Label ThreadPauseHandler;
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, State.rip)));
auto RipReg = x2;
// Mask the address by the virtual address size so we can check for aliases
LoadConstant(x3, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, x3);
{
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
LoadConstant(x0, Thread->LookupCache->GetPagePointer());
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
// Load the pointer from the offset
ldr(x0, MemOperand(x0, x1, Shift::LSL, 3));
// If page pointer is zero then we have no block
cbz(x0, &NoBlock);
// Steal the page offset
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x1, &NoBlock);
// If we've made it here then we have a real compiled block
{
mov(x0, STATE);
blr(x1);
}
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, CTX)));
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
}
else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
{
bind(&Exit);
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
PopCalleeSavedRegisters();
// Return from the function
// LR is set to the correct return location now
ret();
}
// Need to create the block
{
bind(&NoBlock);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, CTX)));
mov(x1, STATE);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
LoadConstant(x3, Ptr.Data);
// X2 contains our guest RIP
blr(x3); // { CTX, ThreadState, RIP}
b(&LoopTop);
}
{
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
// Call our sleep handler
LoadConstant(x0, reinterpret_cast<uintptr_t>(CTX));
mov(x1, STATE);
LoadConstant(x2, reinterpret_cast<uint64_t>(SleepThread));
blr(x2);
// XXX: Unsupported atm
//PauseReturnInstruction = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
//// Fault to start running again
//hlt(0);
}
{
CallbackPtr = Buffer->GetOffsetAddress<CPUBackend::JITCallback>(GetCursorOffset());
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
// First thing we need to move the thread state pointer back in to our register
mov(STATE, x0);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.rip)));
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
FinalizeCode();
uint64_t CodeEnd = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), CodeEnd - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
}
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
constexpr uint32_t FPR_MAGIC = 0x46508001U;
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
struct ContextBackup {
// Host State
uint64_t GPRs[31];
uint64_t PrevSP;
uint64_t PrevPC;
uint64_t PState;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
};
void DispatchGenerator::StoreThreadState(int Signal, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = _mcontext->sp;
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ContextBackup);
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
memcpy(&Context->GPRs[0], &_mcontext->regs[0], 31 * sizeof(uint64_t));
Context->PrevSP = _mcontext->sp;
Context->PrevPC = _mcontext->pc;
Context->PState = _mcontext->pstate;
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
Context->FPSR = HostState->FPSR;
Context->FPCR = HostState->FPCR;
memcpy(&Context->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, &State->State, sizeof(FEXCore::Core::CPUState));
// Set the new SP
_mcontext->sp = NewSP;
}
void DispatchGenerator::RestoreThreadState(void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint64_t OldSP = _mcontext->sp;
uintptr_t NewSP = OldSP;
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(&State->State, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Context->FPRs[0], 32 * sizeof(__uint128_t));
Context->FPCR = HostState->FPCR;
Context->FPSR = HostState->FPSR;
// Restore GPRs and other state
_mcontext->pstate = Context->PState;
_mcontext->pc = Context->PrevPC;
_mcontext->sp = Context->PrevSP;
memcpy(&_mcontext->regs[0], &Context->GPRs[0], 31 * sizeof(uint64_t));
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool DispatchGenerator::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->pc = AbsoluteLoopTopAddress;
// Set x28 (which is our state register) to point to our guest thread data
_mcontext->regs[28 /* STATE */] = reinterpret_cast<uint64_t>(State);
State->State.State.gregs[X86State::REG_RDI] = Signal;
uint64_t OldGuestSP = State->State.State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
if (GuestAction->sa_flags & SA_SIGINFO) {
// XXX: siginfo_t(RSI), ucontext (RDX)
State->State.State.gregs[X86State::REG_RSI] = 0;
State->State.State.gregs[X86State::REG_RDX] = 0;
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
State->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
return true;
}
bool DispatchGenerator::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = State->SignalReason.load();
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->pc = ThreadPauseHandlerAddress;
// Set our state register to point to our guest thread data
_mcontext->regs[28 /* STATE */] = reinterpret_cast<uint64_t>(State);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the JIT and get out safely
_mcontext->sp = State->State.ReturningStackLocation;
// Set the new PC
_mcontext->pc = ThreadStopHandlerAddress;
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
// TODO: Implement this. It is missing from the dispatcher
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = nullptr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
DispatchGenerator *Gen = Generator;
return Gen->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
}
bool InterpreterCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
DispatchGenerator *Gen = Generator;
return Gen->HandleSignalPause(Signal, info, ucontext);
}
void InterpreterCore::DeleteAsmDispatch() {
delete Generator;
}
}
@@ -2,13 +2,15 @@
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::CPU {
class DispatchGenerator;
class X86DispatchGenerator;
class Arm64DispatchGenerator;
#define DESTMAP_AS_MAP 0
#if DESTMAP_AS_MAP
@@ -29,7 +31,6 @@ public:
bool NeedsOpDispatch() override { return true; }
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
void DeleteAsmDispatch();
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
@@ -38,8 +39,6 @@ private:
FEXCore::Core::InternalThreadState *State;
uint32_t AllocateTmpSpace(size_t Size);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
template<typename Res>
Res GetDest(void* SSAData, IR::OrderedNodeWrapper Op);
@@ -47,7 +46,7 @@ private:
template<typename Res>
Res GetSrc(void* SSAData, IR::OrderedNodeWrapper Src);
DispatchGenerator *Generator{};
Dispatcher *Dispatcher{};
};
}
@@ -2,9 +2,8 @@
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#ifdef _M_ARM_64
#include "Interface/Core/ArchHelpers/Arm64.h"
#endif
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/DebugData.h"
#include "Interface/Core/InternalThreadState.h"
@@ -22,62 +21,64 @@
#include <cmath>
#include <limits>
#include <vector>
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
#include "InterpreterOps.h"
namespace FEXCore::CPU {
static void InterpreterExecution(FEXCore::Core::InternalThreadState *Thread) {
auto LocalEntry = Thread->LocalIRCache.find(Thread->State.State.rip);
static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
auto LocalEntry = Thread->LocalIRCache.find(Thread->CurrentFrame->State.rip);
InterpreterOps::InterpretIR(Thread, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
bool InterpreterCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
#ifdef _M_ARM_64
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint32_t *PC = (uint32_t*)_mcontext->pc;
uint32_t Instr = PC[0];
if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x: PC: %p Instruction: 0x%08x\n", Op, PC, PC[0]);
return false;
}
}
constexpr bool is_arm64 = true;
#else
constexpr bool is_arm64 = false;
#endif
if constexpr (is_arm64) {
uint32_t *PC = reinterpret_cast<uint32_t*>(ArchHelpers::Context::GetPc(ucontext));
uint32_t Instr = PC[0];
if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x: PC: %p Instruction: 0x%08x\n", Op, PC, PC[0]);
return false;
}
}
}
return false;
}
@@ -91,7 +92,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->HandleSignalPause(Signal, info, ucontext);
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
});
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
@@ -101,7 +102,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
@@ -112,7 +113,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
InterpreterCore::~InterpreterCore() {
DeleteAsmDispatch();
delete Dispatcher;
}
@@ -400,6 +400,7 @@ static void StopThread(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->StopThread(Thread);
LogMan::Msg::A("unreachable");
__builtin_unreachable();
}
[[noreturn]]
@@ -407,6 +408,7 @@ static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->SignalThread(Thread, FEXCore::Core::SIGNALEVENT_RETURN);
LogMan::Msg::A("unreachable");
__builtin_unreachable();
}
template<IR::IROps Op>
@@ -703,7 +705,7 @@ struct OpHandlers<IR::OP_F80BCDLOAD> {
template<>
struct OpHandlers<IR::OP_F80LOADFCW> {
static void handle(uint16_t NewFCW) {
auto PC = (NewFCW >> 8) & 3;
switch(PC) {
case 0: extF80_roundingPrecision = 32; break;
@@ -711,7 +713,7 @@ struct OpHandlers<IR::OP_F80LOADFCW> {
case 3: extF80_roundingPrecision = 80; break;
case 1: LogMan::Msg::A("Invalid x87 precision mode, %d", PC);
}
auto RC = (NewFCW >> 10) & 3;
switch(RC) {
case 0:
@@ -825,8 +827,6 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
break;
}
case IR::OP_F80CVT: {
auto Op = IROp->C<IR::IROp_F80CVT>();
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&OpHandlers<IR::OP_F80CVT>::handle4);
@@ -862,7 +862,7 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
decltype(&OpHandlers<IR::OP_F80CMP>::handle<0>) handlers[] = { &OpHandlers<IR::OP_F80CMP>::handle<0>, &OpHandlers<IR::OP_F80CMP>::handle<1>, &OpHandlers<IR::OP_F80CMP>::handle<2>, &OpHandlers<IR::OP_F80CMP>::handle<3>, &OpHandlers<IR::OP_F80CMP>::handle<4>, &OpHandlers<IR::OP_F80CMP>::handle<5>, &OpHandlers<IR::OP_F80CMP>::handle<6>, &OpHandlers<IR::OP_F80CMP>::handle<7> };
*Info = GetFallbackInfo(handlers[Op->Flags]);
@@ -988,7 +988,6 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
}
case IR::OP_REMOVECODEENTRY: {
auto Op = IROp->C<IR::IROp_RemoveCodeEntry>();
Thread->CTX->RemoveCodeEntry(Thread, CurrentIR->GetHeader()->Entry);
break;
}
@@ -1016,7 +1015,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
}
case IR::OP_EXITFUNCTION: {
auto Op = IROp->C<IR::IROp_ExitFunction>();
uintptr_t* ContextPtr = reinterpret_cast<uintptr_t*>(&Thread->State.State.rip);
uintptr_t* ContextPtr = reinterpret_cast<uintptr_t*>(Thread->CurrentFrame);
void *Data = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(SSAData, Op->Header.Args[0]);
@@ -1083,7 +1082,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
Args.Argument[j] = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[j]);
}
uint64_t Res = FEXCore::Context::HandleSyscall(Thread->CTX->SyscallHandler, Thread, &Args);
uint64_t Res = FEXCore::Context::HandleSyscall(Thread->CTX->SyscallHandler, Thread->CurrentFrame, &Args);
GD = Res;
break;
}
@@ -1098,8 +1097,9 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
auto Op = IROp->C<IR::IROp_CPUID>();
uint64_t *DstPtr = GetDest<uint64_t*>(SSAData, WrapperOp);
uint64_t Arg = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]);
uint64_t Leaf = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]);
auto Results = Thread->CTX->CPUID.RunFunction(Arg);
auto Results = Thread->CTX->CPUID.RunFunction(Arg, Leaf);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
break;
}
@@ -1222,7 +1222,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
case IR::OP_LOADCONTEXT: {
auto Op = IROp->C<IR::IROp_LoadContext>();
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += Op->Offset;
#define LOAD_CTX(x, y) \
case x: { \
@@ -1249,7 +1249,8 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
@@ -1277,7 +1278,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
case IR::OP_STORECONTEXT: {
auto Op = IROp->C<IR::IROp_StoreContext>();
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += Op->Offset;
void *Data = reinterpret_cast<void*>(ContextPtr);
@@ -1289,7 +1290,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
@@ -1366,7 +1367,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
case IR::OP_LOADFLAG: {
auto Op = IROp->C<IR::IROp_LoadFlag>();
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += offsetof(FEXCore::Core::CPUState, flags[0]);
ContextPtr += Op->Flag;
uint8_t const *Data = reinterpret_cast<uint8_t const*>(ContextPtr);
@@ -1377,7 +1378,7 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
auto Op = IROp->C<IR::IROp_StoreFlag>();
uint8_t Arg = *GetSrc<uint8_t*>(SSAData, Op->Header.Args[0]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(&Thread->State.State);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Thread->CurrentFrame);
ContextPtr += offsetof(FEXCore::Core::CPUState, flags[0]);
ContextPtr += Op->Flag;
uint8_t *Data = reinterpret_cast<uint8_t*>(ContextPtr);
@@ -5157,4 +5158,4 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEX
}
}
}
}
@@ -1,470 +0,0 @@
#include "Common/MathUtils.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
class DispatchGenerator : public Xbyak::CodeGenerator {
public:
DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
uint64_t ThreadStopHandlerAddress;
uint64_t AbsoluteLoopTopAddress;
uint64_t ThreadPauseHandlerAddress;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
private:
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
};
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->State.RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE)
, CTX {ctx}
, State {Thread} {
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
// while (!Thread->State.RunningEvents.ShouldStop.load()) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Bunch of exit state stuff
// x86-64 ABI has the stack aligned when /call/ happens
// Which means the destination has a misaligned stack at that point
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
mov(STATE, rdi);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)], rsp);
Label LoopTop;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
L(LoopTop);
AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
mov(r13, Thread->LookupCache->GetPagePointer());
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
// Load page pointer
mov(rdi, qword [r13 + rax * 8]);
cmp(rdi, 0);
je(NoBlock);
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
cmp(rcx, rdx);
jne(NoBlock);
// Load the block pointer
mov(rax, qword [rdi + rax]);
cmp(rax, 0);
je(NoBlock);
// Real block if we made it here
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
{
L(ExitBlock);
ThreadStopHandlerAddress = getCurr<uint64_t>();
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
// Block creation
{
L(NoBlock);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
mov(rax, Ptr.Data);
call(rax);
// rdx already contains RIP here
jmp(LoopTop);
}
{
// Pause handler
ThreadPauseHandlerAddress = getCurr<uint64_t>();
L(ThreadPauseHandler);
mov(rdi, reinterpret_cast<uintptr_t>(CTX));
mov(rsi, STATE);
mov(rax, reinterpret_cast<uint64_t>(SleepThread));
call(rax);
// XXX: Unsupported atm
// uint64_t PauseReturnInstruction = getCurr<uint64_t>();
// ud2();
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
// First thing we need to move the thread state pointer back in to our register
mov(STATE, rdi);
// XXX: XMM?
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
mov(rax, CTX->X86CodeGen.CallbackReturn);
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])]);
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.rip)], rsi);
// Back to the loop top now
jmp(LoopTop);
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
// rsi = rsp
mov(rsp, rsi);
// Now jump back to the thunk
// XXX: XMM?
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
ready();
}
struct ContextBackup {
uint64_t StoredCookie;
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[NGREG];
_libc_fpstate FPRState;
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
};
void DispatchGenerator::StoreThreadState(int Signal, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = _mcontext->gregs[REG_RSP];
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ContextBackup);
// We need to back up behind the host's red zone
// We do this on the guest side as well
NewSP -= 128;
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
Context->StoredCookie = 0x4142434445464748ULL;
// Copy the GPRs
memcpy(&Context->GPRs[0], &_mcontext->gregs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(&Context->FPRState, _mcontext->fpregs, sizeof(_libc_fpstate));
// XXX: Save 256bit and 512bit AVX register state
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, &State->State, sizeof(FEXCore::Core::CPUState));
// Set the new SP
_mcontext->gregs[REG_RSP] = NewSP;
SignalFrames.push(NewSP);
}
void DispatchGenerator::RestoreThreadState(void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uintptr_t NewSP = OldSP;
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
if (Context->StoredCookie != 0x4142434445464748ULL) {
LogMan::Msg::D("COOKIE WAS NOT CORRECT!\n");
exit(-1);
}
// First thing, reset the guest state
memcpy(&State->State, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Context->GPRs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Context->FPRState, sizeof(_libc_fpstate));
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool DispatchGenerator::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = AbsoluteLoopTopAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(State);
uint64_t OldGuestSP = State->State.State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
State->State.State.gregs[X86State::REG_RDI] = Signal;
if (GuestAction->sa_flags & SA_SIGINFO) {
// XXX: siginfo_t(RSI), ucontext (RDX)
State->State.State.gregs[X86State::REG_RSI] = 0;
State->State.State.gregs[X86State::REG_RDX] = 0;
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
State->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
return true;
}
bool DispatchGenerator::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = State->SignalReason.load();
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadPauseHandlerAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(State);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the core and get out safely
_mcontext->gregs[REG_RSP] = State->State.ReturningStackLocation;
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadStopHandlerAddress;
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Generator->ReturnPtr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
DispatchGenerator *Gen = Generator;
return Gen->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
}
bool InterpreterCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
DispatchGenerator *Gen = Generator;
return Gen->HandleSignalPause(Signal, info, ucontext);
}
void InterpreterCore::DeleteAsmDispatch() {
delete Generator;
}
}
+13 -25
View File
@@ -1,3 +1,8 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -29,7 +34,7 @@ static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -398,7 +403,6 @@ DEF_OP(Xor) {
DEF_OP(Lshl) {
auto Op = IROp->C<IR::IROp_Lshl>();
uint8_t OpSize = IROp->Size;
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
lsl(GRS(Node), GRS(Op->Header.Args[0].ID()), (unsigned int)Const);
@@ -409,7 +413,6 @@ DEF_OP(Lshl) {
DEF_OP(Lshr) {
auto Op = IROp->C<IR::IROp_Lshr>();
uint8_t OpSize = IROp->Size;
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
lsr(GRS(Node), GRS(Op->Header.Args[0].ID()), (unsigned int)Const);
@@ -524,14 +527,10 @@ DEF_OP(LDiv) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LDIV);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -571,14 +570,10 @@ DEF_OP(LUDiv) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LUDIV);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LUDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -628,14 +623,10 @@ DEF_OP(LRem) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LREM);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -682,14 +673,11 @@ DEF_OP(LURem) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LUREM);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LUREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
@@ -920,8 +908,8 @@ Condition MapSelectCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:;
case FEXCore::IR::COND_VC:;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
default:
@@ -945,7 +933,7 @@ DEF_OP(Select) {
} else {
LogMan::Msg::A("Select: Expected GPR or FPR");
}
auto cc = MapSelectCC(Op->Cond);
uint64_t const_true, const_false;
@@ -1022,7 +1010,7 @@ DEF_OP(FCmp) {
fcmp(GetSrc(Op->Header.Args[0].ID()).D(), GetSrc(Op->Header.Args[1].ID()).D());
}
auto Dst = GetReg<RA_64>(Node);
bool set = false;
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ)) {
@@ -1057,8 +1045,8 @@ DEF_OP(FCmp) {
#undef DEF_OP
void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
@@ -865,8 +871,8 @@ DEF_OP(AtomicFetchXor) {
}
#undef DEF_OP
void JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
@@ -8,7 +14,7 @@
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::D("Unimplemented");
}
@@ -35,7 +41,7 @@ DEF_OP(CallbackReturn) {
// spill back to CTX
SpillStaticRegs();
// First we must reset the stack
ResetStack();
@@ -46,9 +52,9 @@ DEF_OP(CallbackReturn) {
str(w2, MemOperand(x0));
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
add(x2, x2, 8);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
PopCalleeSavedRegisters();
@@ -67,7 +73,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ExitFunctionLinkerAddress};
Literal l_BranchHost{Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchGuest{NewRIP};
ldr(x0, &l_BranchHost);
@@ -77,9 +83,9 @@ DEF_OP(ExitFunction) {
place(&l_BranchGuest);
} else {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// L1 Cache
LoadConstant(x0, State->LookupCache->GetL1Pointer());
LoadConstant(x0, ThreadState->LookupCache->GetL1Pointer());
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -90,8 +96,8 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
LoadConstant(TMP1, AbsoluteLoopTopAddress);
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, State.rip)));
LoadConstant(TMP1, Dispatcher->AbsoluteLoopTopAddress);
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
}
@@ -131,8 +137,8 @@ Condition MapBranchCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:;
case FEXCore::IR::COND_VC:;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
default:
@@ -182,7 +188,7 @@ DEF_OP(CondJump) {
b(TrueTargetLabel, MapBranchCC(Op->Cond));
}
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
@@ -217,7 +223,7 @@ DEF_OP(Syscall) {
blr(x3);
add(sp, sp, SPOffset);
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs();
@@ -239,16 +245,12 @@ DEF_OP(Thunk) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
#if _M_X86_64
ERROR_AND_DIE("JIT: OP_THUNK not supported with arm simulator")
#else
auto thunkFn = State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
blr(x2);
#endif
PopDynamicRegsAndLR();
FillStaticRegs(); // load from ctx after ra64 refill
}
@@ -262,7 +264,7 @@ DEF_OP(ValidateCode) {
LoadConstant(GetReg<RA_64>(Node), 0);
LoadConstant(x0, IR->GetHeader()->Entry + Op->Offset);
LoadConstant(x1, 1);
while (len >= 8)
{
ldr(x2, MemOperand(x0, idx));
@@ -302,17 +304,16 @@ DEF_OP(ValidateCode) {
}
DEF_OP(RemoveCodeEntry) {
auto Op = IROp->C<IR::IROp_RemoveCodeEntry>();
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
PushDynamicRegsAndLR();
mov(x0, STATE);
LoadConstant(x1, IR->GetHeader()->Entry);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntry));
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -323,15 +324,17 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR();
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
LoadConstant(x0, reinterpret_cast<uint64_t>(&CTX->CPUID));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t);
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t, uint32_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
@@ -354,8 +357,8 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -193,8 +199,8 @@ DEF_OP(Vector_FToF) {
}
#undef DEF_OP
void JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_U, Float_FromGPR_U);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -83,8 +89,8 @@ DEF_OP(AESKeyGenAssist) {
}
#undef DEF_OP
void JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
@@ -1,18 +1,24 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
}
#undef DEF_OP
void JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
File diff suppressed because it is too large. Load diff
+16 -91
View File
@@ -1,9 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#pragma once
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
@@ -29,54 +36,16 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
// All but x29 are caller saved
const std::array<aarch64::Register, 16> SRA64 = {
x4, x5, x6, x7, x8, x9, x10, x11,
x12, x18, x17, x16, x15, x14, x13, x29
};
// All are callee saved
const std::array<aarch64::Register, 9> RA64 = {
x20, x21, x22, x23, x24, x25, x26, x27,
x19
};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA64Pair = {{
{x20, x21},
{x22, x23},
{x24, x25},
{x26, x27},
}};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA32Pair = {{
{w20, w21},
{w22, w23},
{w24, w25},
{w26, w27},
}};
// All are caller saved
const std::array<aarch64::VRegister, 16> SRAFPR = {
v16, v17, v18, v19, v20, v21, v22, v23,
v24, v25, v26, v27, v28, v29, v30, v31
};
// v8..v15 = (lower 64bits) Callee saved
const std::array<aarch64::VRegister, 12> RAFPR = {
/*v0, v1, v2, v3,*/v4, v5, v6, v7, // v0 ~ v3 are used as temps
v8, v9, v10, v11, v12, v13, v14, v15
};
class JITCore final : public CPUBackend, public vixl::aarch64::Assembler {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
explicit JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
explicit Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
~JITCore() override;
~Arm64JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
@@ -86,20 +55,18 @@ public:
void ClearCache() override;
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static CodeBuffer AllocateNewCodeBuffer(size_t Size);
CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
private:
Dispatcher *Dispatcher;
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
std::map<IR::OrderedNodeWrapper::NodeOffsetType, aarch64::Label> JumpTargets;
@@ -162,9 +129,6 @@ private:
#if DEBUG
vixl::aarch64::Decoder Decoder;
#endif
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
@@ -181,8 +145,6 @@ private:
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the codebuffer that our dispatcher lives in
CodeBuffer DispatcherCodeBuffer{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
@@ -190,38 +152,11 @@ private:
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true);
#if DEBUG
vixl::aarch64::Disassembler Disasm;
#endif
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread);
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
static uint64_t ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadState *Thread, uint64_t *record);
/**
* @name Dispatch Helper functions
* @{ */
uint64_t AbsoluteLoopTopAddressFillSRA{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t ThreadStopHandlerAddressSpillSRA{};
uint64_t ThreadStopHandlerAddress{};
uint64_t PauseReturnInstruction{};
uint32_t SignalHandlerRefCounter{};
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
/** @} */
static uint64_t ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
@@ -233,17 +168,7 @@ private:
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
uint32_t SpillSlots{};
void SpillStaticRegs();
void FillStaticRegs();
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
void ResetStack();
using OpHandler = void (JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
using OpHandler = void (Arm64JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -95,16 +101,15 @@ DEF_OP(StoreContext) {
DEF_OP(LoadRegister) {
auto Op = IROp->C<IR::IROp_LoadRegister>();
uint8_t OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.gregs[0])) / 8;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0])) / 8;
auto regOffs = Op->Offset & 7;
LogMan::Throw::A(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
switch(Op->Header.Size) {
case 1:
LogMan::Throw::A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
@@ -129,7 +134,7 @@ DEF_OP(LoadRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regOffs = Op->Offset & 15;
LogMan::Throw::A(regId < SRAFPR.size(), "out of range regId");
@@ -181,8 +186,6 @@ DEF_OP(LoadRegister) {
DEF_OP(StoreRegister) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
uint8_t OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = Op->Offset / 8 - 1;
@@ -191,7 +194,7 @@ DEF_OP(StoreRegister) {
LogMan::Throw::A(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
switch(Op->Header.Size) {
case 1:
LogMan::Throw::A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
@@ -202,7 +205,7 @@ DEF_OP(StoreRegister) {
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), 0, 16);
break;
case 4:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), 0, 32);
@@ -215,7 +218,7 @@ DEF_OP(StoreRegister) {
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regOffs = Op->Offset & 15;
LogMan::Throw::A(regId < SRAFPR.size(), "regId out of range");
@@ -530,7 +533,7 @@ DEF_OP(StoreFlag) {
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
}
MemOperand JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
if (Offset.IsInvalid()) {
return MemOperand(Base);
} else {
@@ -551,6 +554,7 @@ MemOperand JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Bas
}
}
}
__builtin_unreachable();
}
DEF_OP(LoadMem) {
@@ -792,8 +796,8 @@ DEF_OP(VStoreMemElement) {
}
#undef DEF_OP
void JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -32,18 +38,18 @@ DEF_OP(Break) {
case 4: { // HLT
// Time to quit
// Set our stack to the starting stack location
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
LoadConstant(TMP1, ThreadStopHandlerAddressSpillSRA);
LoadConstant(TMP1, Dispatcher->ThreadStopHandlerAddressSpillSRA);
br(TMP1);
break;
}
case 6: { // INT3
ResetStack();
LoadConstant(TMP1, ThreadPauseHandlerAddressSpillSRA);
LoadConstant(TMP1, Dispatcher->ThreadPauseHandlerAddressSpillSRA);
br(TMP1);
break;
}
@@ -52,7 +58,6 @@ DEF_OP(Break) {
}
DEF_OP(GetRoundingMode) {
auto Op = IROp->C<IR::IROp_GetRoundingMode>();
auto Dst = GetReg<RA_64>(Node);
mrs(Dst, FPCR);
lsr(Dst, Dst, 22);
@@ -111,8 +116,8 @@ DEF_OP(SetRoundingMode) {
}
#undef DEF_OP
void JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -68,8 +74,8 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VectorZero) {
uint8_t OpSize = IROp->Size;
switch (OpSize) {
@@ -1519,7 +1525,7 @@ DEF_OP(VSShrS) {
DEF_OP(VInsElement) {
auto Op = IROp->C<IR::IROp_VInsElement>();
auto reg = GetSrc(Op->Header.Args[0].ID());
if (GetDst(Node).GetCode() != reg.GetCode()) {
@@ -1554,7 +1560,7 @@ DEF_OP(VInsElement) {
DEF_OP(VInsScalarElement) {
auto Op = IROp->C<IR::IROp_VInsScalarElement>();
auto reg = GetSrc(Op->Header.Args[0].ID());
if (GetDst(Node).GetCode() != reg.GetCode()) {
@@ -2112,8 +2118,8 @@ DEF_OP(VTBL1) {
}
#undef DEF_OP
void JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
+2 -1
View File
@@ -11,5 +11,6 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
FEXCore::CPU::CPUBackend *CreateJITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
FEXCore::CPU::CPUBackend *CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
FEXCore::CPU::CPUBackend *CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
}
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -1168,8 +1174,8 @@ DEF_OP(FCmp) {
#undef DEF_OP
void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
@@ -548,8 +554,8 @@ DEF_OP(AtomicFetchXor) {
}
#undef DEF_OP
void JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -6,7 +12,7 @@
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::D("Unimplemented");
}
@@ -40,7 +46,7 @@ DEF_OP(CallbackReturn) {
sub(dword [rax], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])], 8);
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
// Now jump back to the thunk
// XXX: XMM?
@@ -75,12 +81,12 @@ DEF_OP(ExitFunction) {
jmp(qword[rax]);
L(l_BranchHost);
dq(ExitFunctionLinkerAddress);
dq(Dispatcher->ExitFunctionLinkerAddress);
L(l_BranchGuest);
dq(NewRIP);
} else {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, ThreadState->LookupCache->GetL1Pointer());
mov(rax, RipReg);
@@ -89,14 +95,14 @@ DEF_OP(ExitFunction) {
shl(rax, 4);
Xbyak::RegExp LookupBase = rcx + rax;
cmp(qword[LookupBase + 8], RipReg);
jne(FullLookup);
jmp(qword[LookupBase + 0]);
L(FullLookup);
mov(rax, AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.rip)], RipReg);
mov(rax, Dispatcher->AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(rax);
}
@@ -227,7 +233,7 @@ DEF_OP(Thunk) {
sub(rsp, 8); // Align
mov(rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
mov(rax, reinterpret_cast<uintptr_t>(thunkFn));
@@ -271,8 +277,6 @@ DEF_OP(ValidateCode) {
}
DEF_OP(RemoveCodeEntry) {
auto Op = IROp->C<IR::IROp_RemoveCodeEntry>();
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -286,7 +290,7 @@ DEF_OP(RemoveCodeEntry) {
mov(rsi, rax);
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntry));
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
call(rax);
if (NumPush & 1)
@@ -299,7 +303,7 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function);
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function, uint32_t Leaf);
union {
ClassPtrType ClassPtr;
uint64_t Raw;
@@ -316,6 +320,7 @@ DEF_OP(CPUID) {
// Result: RAX, RDX. 4xi32
mov (rsi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov (rdx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov (rdi, reinterpret_cast<uint64_t>(&CTX->CPUID));
auto NumPush = RA64.size();
@@ -341,8 +346,8 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -171,8 +177,8 @@ DEF_OP(Vector_FToF) {
}
#undef DEF_OP
void JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_U, Float_FromGPR_U);
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -35,8 +41,8 @@ DEF_OP(AESKeyGenAssist) {
}
#undef DEF_OP
void JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
@@ -14,8 +20,8 @@ DEF_OP(GetHostFlag) {
}
#undef DEF_OP
void JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
+68 -605
View File
@@ -1,5 +1,13 @@
/*
$info$
tags: backend|x86-64
desc: Main glue logic of the x86-64 splatter backend
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
@@ -38,308 +46,13 @@ void FreeCodeBuffer(CodeBuffer Buffer) {
}
namespace FEXCore::CPU {
struct ContextBackup {
uint64_t StoredCookie;
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[NGREG];
_libc_fpstate FPRState;
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
};
void JITCore::StoreThreadState(int Signal, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = _mcontext->gregs[REG_RSP];
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ContextBackup);
// We need to back up behind the host's red zone
// We do this on the guest side as well
NewSP -= 128;
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
Context->StoredCookie = 0x4142434445464748ULL;
// Copy the GPRs
memcpy(&Context->GPRs[0], &_mcontext->gregs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(&Context->FPRState, _mcontext->fpregs, sizeof(_libc_fpstate));
// XXX: Save 256bit and 512bit AVX register state
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, &ThreadState->State, sizeof(FEXCore::Core::CPUState));
// Set the new SP
_mcontext->gregs[REG_RSP] = NewSP;
SignalFrames.push(NewSP);
}
void JITCore::RestoreThreadState(void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uintptr_t NewSP = OldSP;
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
if (Context->StoredCookie != 0x4142434445464748ULL) {
LogMan::Msg::D("COOKIE WAS NOT CORRECT!\n");
exit(-1);
}
// First thing, reset the guest state
memcpy(&ThreadState->State, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Context->GPRs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Context->FPRState, sizeof(_libc_fpstate));
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool JITCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = AbsoluteLoopTopAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(ThreadState);
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
uint64_t OldGuestSP = ThreadState->State.State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
if (GuestAction->sa_flags & SA_SIGINFO) {
// Setup ucontext a bit
if (CTX->Config.Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::ucontext_t);
uint64_t UContextLocation = NewGuestSP;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
// We have extended float information
guest_uctx->uc_flags |= FEXCore::x86_64::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = &guest_uctx->__fpregs_mem;
#define COPY_REG(x) \
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x] = ThreadState->State.State.gregs[X86State::REG_##x];
COPY_REG(R8);
COPY_REG(R9);
COPY_REG(R10);
COPY_REG(R11);
COPY_REG(R12);
COPY_REG(R13);
COPY_REG(R14);
COPY_REG(R15);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
// Copy float registers
memcpy(guest_uctx->__fpregs_mem._st, ThreadState->State.State.mm, sizeof(ThreadState->State.State.mm));
memcpy(guest_uctx->__fpregs_mem._xmm, ThreadState->State.State.xmm, sizeof(ThreadState->State.State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = ThreadState->State.State.FCW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
(ThreadState->State.State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] << 11) |
(ThreadState->State.State.flags[FEXCore::X86State::X87FLAG_C0_LOC] << 8) |
(ThreadState->State.State.flags[FEXCore::X86State::X87FLAG_C1_LOC] << 9) |
(ThreadState->State.State.flags[FEXCore::X86State::X87FLAG_C2_LOC] << 10) |
(ThreadState->State.State.flags[FEXCore::X86State::X87FLAG_C3_LOC] << 14);
// Copy over signal stack information
guest_uctx->uc_stack.ss_flags = GuestStack->ss_flags;
guest_uctx->uc_stack.ss_sp = GuestStack->ss_sp;
guest_uctx->uc_stack.ss_size = GuestStack->ss_size;
// XXX: siginfo_t(RSI)
ThreadState->State.State.gregs[X86State::REG_RSI] = 0x4142434445460000;
ThreadState->State.State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
// XXX: 32bit Support
NewGuestSP -= sizeof(FEXCore::x86::ucontext_t);
uint64_t UContextLocation = 0; // NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::siginfo_t);
uint64_t SigInfoLocation = 0; // NewGuestSP;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = UContextLocation;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SigInfoLocation;
}
ThreadState->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
ThreadState->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
if (CTX->Config.Is64BitMode) {
ThreadState->State.State.gregs[X86State::REG_RDI] = Signal;
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
ThreadState->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
}
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
LogMan::Throw::A(CTX->X86CodeGen.SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
ThreadState->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
}
return true;
}
void JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
JITCore *Core = reinterpret_cast<JITCore*>(Original);
void X86JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
}
bool JITCore::HandleSIGILL(int Signal, void *info, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
if (_mcontext->gregs[REG_RIP] == ThreadSharedData.SignalHandlerReturnAddress) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
if (_mcontext->gregs[REG_RIP] == PauseReturnInstruction) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
return false;
}
bool JITCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = ThreadState->SignalReason.load();
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadPauseHandlerAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(ThreadState);
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the JIT and get out safely
_mcontext->gregs[REG_RSP] = ThreadState->State.ReturningStackLocation;
// Our ref counting doesn't matter anymore
SignalHandlerRefCounter = 0;
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadStopHandlerAddress;
ThreadState->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
void JITCore::PushRegs() {
void X86JITCore::PushRegs() {
for (auto &Xmm : RAXMM_x) {
sub(rsp, 16);
movaps(ptr[rsp], Xmm);
@@ -353,7 +66,7 @@ void JITCore::PushRegs() {
sub(rsp, 8); // Align
}
void JITCore::PopRegs() {
void X86JITCore::PopRegs() {
auto NumPush = RA64.size();
if (NumPush & 1)
@@ -367,7 +80,7 @@ void JITCore::PopRegs() {
}
}
void JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
auto Name = FEXCore::IR::GetName(IROp->Op);
@@ -574,17 +287,15 @@ void JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
}
void JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
JITCore::JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
: CodeGenerator(Buffer.Size, Buffer.Ptr, nullptr)
, CTX {ctx}
, ThreadState {Thread}
, InitialCodeBuffer {Buffer}
{
ThreadSharedData.SignalHandlerRefCounterPtr = &SignalHandlerRefCounter;
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetRAPass();
@@ -600,7 +311,7 @@ JITCore::JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadSt
}
for (uint32_t i = 0; i < FEXCore::IR::IROps::OP_LAST + 1; ++i) {
OpHandlers[i] = &JITCore::Op_Unhandled;
OpHandlers[i] = &X86JITCore::Op_Unhandled;
}
RegisterALUHandlers();
@@ -615,22 +326,32 @@ JITCore::JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadSt
RegisterEncryptionHandlers();
if (!CompileThread) {
CreateCustomDispatch(Thread);
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
Dispatcher = new X86Dispatcher(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
JITCore *Core = reinterpret_cast<JITCore*>(Thread->CPUBackend.get());
return Core->HandleSIGILL(Signal, info, ucontext);
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
});
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
JITCore *Core = reinterpret_cast<JITCore*>(Thread->CPUBackend.get());
return Core->HandleSignalPause(Signal, info, ucontext);
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
});
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
JITCore *Core = reinterpret_cast<JITCore*>(Thread->CPUBackend.get());
return Core->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
@@ -639,20 +360,17 @@ JITCore::JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadSt
}
}
JITCore::~JITCore() {
X86JITCore::~X86JITCore() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
if (DispatcherCodeBuffer.Ptr) {
// Dispatcher may not exist if this is a compile thread
FreeCodeBuffer(DispatcherCodeBuffer);
}
FreeCodeBuffer(InitialCodeBuffer);
}
void JITCore::ClearCache() {
void X86JITCore::ClearCache() {
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
@@ -686,13 +404,13 @@ void JITCore::ClearCache() {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(JITCore::INITIAL_CODE_SIZE);
auto NewCodeBuffer = AllocateNewCodeBuffer(X86JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
}
IR::PhysicalRegister JITCore::GetPhys(uint32_t Node) {
IR::PhysicalRegister X86JITCore::GetPhys(uint32_t Node) {
auto PhyReg = RAData->GetNodeRegister(Node);
LogMan::Throw::A(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa%d. Class: %d", Node, PhyReg.Class);
@@ -700,16 +418,16 @@ IR::PhysicalRegister JITCore::GetPhys(uint32_t Node) {
return PhyReg;
}
bool JITCore::IsFPR(uint32_t Node) {
bool X86JITCore::IsFPR(uint32_t Node) {
return RAData->GetNodeRegister(Node).Class == IR::FPRClass.Val;
}
bool JITCore::IsGPR(uint32_t Node) {
bool X86JITCore::IsGPR(uint32_t Node) {
return RAData->GetNodeRegister(Node).Class == IR::GPRClass.Val;
}
template<uint8_t RAType>
Xbyak::Reg JITCore::GetSrc(uint32_t Node) {
Xbyak::Reg X86JITCore::GetSrc(uint32_t Node) {
// rax, rcx, rdx, rsi, r8, r9,
// r10
// Callee Saved
@@ -728,24 +446,24 @@ Xbyak::Reg JITCore::GetSrc(uint32_t Node) {
}
template
Xbyak::Reg JITCore::GetSrc<JITCore::RA_64>(uint32_t Node);
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_64>(uint32_t Node);
template
Xbyak::Reg JITCore::GetSrc<JITCore::RA_32>(uint32_t Node);
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_32>(uint32_t Node);
template
Xbyak::Reg JITCore::GetSrc<JITCore::RA_16>(uint32_t Node);
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_16>(uint32_t Node);
template
Xbyak::Reg JITCore::GetSrc<JITCore::RA_8>(uint32_t Node);
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_8>(uint32_t Node);
Xbyak::Xmm JITCore::GetSrc(uint32_t Node) {
Xbyak::Xmm X86JITCore::GetSrc(uint32_t Node) {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
template<uint8_t RAType>
Xbyak::Reg JITCore::GetDst(uint32_t Node) {
Xbyak::Reg X86JITCore::GetDst(uint32_t Node) {
auto PhyReg = GetPhys(Node);
if (RAType == RA_64)
return RA64[PhyReg.Reg].cvt64();
@@ -760,19 +478,19 @@ Xbyak::Reg JITCore::GetDst(uint32_t Node) {
}
template
Xbyak::Reg JITCore::GetDst<JITCore::RA_64>(uint32_t Node);
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_64>(uint32_t Node);
template
Xbyak::Reg JITCore::GetDst<JITCore::RA_32>(uint32_t Node);
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_32>(uint32_t Node);
template
Xbyak::Reg JITCore::GetDst<JITCore::RA_16>(uint32_t Node);
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_16>(uint32_t Node);
template
Xbyak::Reg JITCore::GetDst<JITCore::RA_8>(uint32_t Node);
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_8>(uint32_t Node);
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair(uint32_t Node) {
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(uint32_t Node) {
auto PhyReg = GetPhys(Node);
if (RAType == RA_64)
return RA64Pair[PhyReg.Reg];
@@ -781,17 +499,17 @@ std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair(uint32_t Node) {
}
template
std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair<JITCore::RA_64>(uint32_t Node);
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_64>(uint32_t Node);
template
std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair<JITCore::RA_32>(uint32_t Node);
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_32>(uint32_t Node);
Xbyak::Xmm JITCore::GetDst(uint32_t Node) {
Xbyak::Xmm X86JITCore::GetDst(uint32_t Node) {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
bool JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
bool X86JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINECONSTANT) {
@@ -805,7 +523,7 @@ bool JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Va
}
}
bool JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
bool X86JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
@@ -819,7 +537,7 @@ bool JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint
}
}
std::tuple<JITCore::SetCC, JITCore::CMovCC, JITCore::JCC> JITCore::GetCC(IR::CondClassType cond) {
std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::GetCC(IR::CondClassType cond) {
switch (cond.Val) {
case FEXCore::IR::COND_EQ: return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
case FEXCore::IR::COND_NEQ: return { &CodeGenerator::setne, &CodeGenerator::cmovne, &CodeGenerator::jne };
@@ -852,12 +570,11 @@ std::tuple<JITCore::SetCC, JITCore::CMovCC, JITCore::JCC> JITCore::GetCC(IR::Con
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *X86JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->RAData = RAData;
auto HeaderOp = IR->GetHeader();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
@@ -874,13 +591,13 @@ void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(RunBlock);
// Else we need to pause now
mov(rax, ThreadPauseHandlerAddress);
mov(rax, Dispatcher->ThreadPauseHandlerAddress);
jmp(rax);
ud2();
@@ -1024,29 +741,18 @@ void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [
return Entry;
}
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->State.RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadState *Thread, uint64_t *record) {
uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->State.State.rip = GuestRip;
return core->AbsoluteLoopTopAddress;
Thread->CurrentFrame->State.rip = GuestRip;
return core->Dispatcher->AbsoluteLoopTopAddress;
}
auto LinkerAddress = core->ExitFunctionLinkerAddress;
auto LinkerAddress = core->Dispatcher->ExitFunctionLinkerAddress;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
@@ -1056,250 +762,7 @@ uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadS
return HostCode;
}
void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
DispatcherCodeBuffer = AllocateNewCodeBuffer(MAX_DISPATCHER_CODE_SIZE);
setNewBuffer(DispatcherCodeBuffer.Ptr, DispatcherCodeBuffer.Size);
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
// r10, r11
//
// Callee Saved
// rbx, rbp, r12, r13, r14, r15
//
// 1St Argument: rdi <ThreadState>
// XMM:
// All temp
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
// while (!Thread->State.RunningEvents.ShouldStop.load()) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Bunch of exit state stuff
// x86-64 ABI has the stack aligned when /call/ happens
// Which means the destination has a misaligned stack at that point
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
mov(STATE, rdi);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [STATE + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)], rsp);
Label LoopTop;
Label FullLookup;
Label NoBlock;
Label ThreadPauseHandler{};
L(LoopTop);
AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
// L1 Cache
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
jne(FullLookup);
jmp(qword[r13 + rax + 0]);
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
// Load page pointer
mov(rdi, qword [r13 + rax * 8]);
cmp(rdi, 0);
je(NoBlock);
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
cmp(rcx, rdx);
jne(NoBlock);
// Load the block pointer
mov(rax, qword [rdi + rax]);
cmp(rax, 0);
je(NoBlock);
// Update L1
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
mov(qword[r13 + rcx*8 + 8], rdx);
mov(qword[r13 + rcx*8 + 0], rax);
// Real block if we made it here
jmp(rax);
}
{
ThreadStopHandlerAddress = getCurr<uint64_t>();
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
// {rdi, rsi, rdx}
mov(rdi, (uintptr_t)this);
mov(rsi, STATE);
mov(rdx, rax); // rax is set at the block end
mov(rax, (uintptr_t)&ExitFunctionLink);
call(rax);
jmp(rax);
}
Label FallbackCore;
// Block creation
{
L(NoBlock);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
mov(rax, Ptr.Data);
call(rax);
// RAX contains nulptr or block ptr here
cmp(rax, 0);
je(FallbackCore);
// rdx already contains RIP here
jmp(LoopTop);
}
{
L(FallbackCore);
ud2();
}
{
// Signal return handler
ThreadSharedData.SignalHandlerReturnAddress = getCurr<uint64_t>();
ud2();
}
{
// Pause handler
ThreadPauseHandlerAddress = getCurr<uint64_t>();
L(ThreadPauseHandler);
mov(rdi, reinterpret_cast<uintptr_t>(CTX));
mov(rsi, STATE);
mov(rax, reinterpret_cast<uint64_t>(SleepThread));
call(rax);
PauseReturnInstruction = getCurr<uint64_t>();
ud2();
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
// First thing we need to move the thread state pointer back in to our register
mov(STATE, rdi);
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
add(dword [rax], 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
mov(rax, CTX->X86CodeGen.CallbackReturn);
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])]);
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.rip)], rsi);
// Back to the loop top now
jmp(LoopTop);
}
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(DispatcherCodeBuffer.Ptr, DispatcherCodeBuffer.Size, Name);
#endif
ready();
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
FEXCore::CPU::CPUBackend *CreateJITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return new JITCore(ctx, Thread, AllocateNewCodeBuffer(CompileThread ? JITCore::MAX_CODE_SIZE : JITCore::INITIAL_CODE_SIZE), CompileThread);
FEXCore::CPU::CPUBackend *CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return new X86JITCore(ctx, Thread, AllocateNewCodeBuffer(CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
}
}
@@ -1 +0,0 @@
+18 -31
View File
@@ -1,11 +1,18 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#pragma once
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JIT.h"
#include "Common/MathUtils.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
@@ -54,10 +61,10 @@ const std::array<std::pair<Xbyak::Reg, Xbyak::Reg>, 4> RA64Pair = {{ {rsi, r8},
const std::array<Xbyak::Reg, 11> RAXMM = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
class JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
~JITCore() override;
explicit X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
~X86JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
@@ -69,10 +76,6 @@ public:
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
private:
@@ -80,6 +83,7 @@ private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
FEXCore::CPU::Dispatcher *Dispatcher;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, Label> JumpTargets;
Xbyak::util::Cpu Features{};
@@ -128,7 +132,6 @@ private:
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr);
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value);
void CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread);
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
@@ -136,13 +139,11 @@ private:
bool GetSamplingData {true};
#endif
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 1;
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
static uint64_t ExitFunctionLink(JITCore* code, FEXCore::Core::InternalThreadState *Thread, uint64_t *record);
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
@@ -153,20 +154,9 @@ private:
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the codebuffer that our dispatcher lives in
CodeBuffer DispatcherCodeBuffer{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t ThreadStopHandlerAddress{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t PauseReturnInstruction{};
uint32_t SignalHandlerRefCounter{};
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
@@ -175,17 +165,14 @@ private:
CompilerSharedData ThreadSharedData;
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
uint32_t SpillSlots{};
using SetCC = void (JITCore::*)(const Operand& op);
using CMovCC = void (JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (JITCore::*)(const Label& label, LabelType type);
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (X86JITCore::*)(const Label& label, LabelType type);
std::tuple<SetCC, CMovCC, JCC> GetCC(IR::CondClassType cond);
using OpHandler = void (JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
using OpHandler = void (X86JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -210,7 +197,7 @@ private:
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -5,7 +11,7 @@
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -419,7 +425,7 @@ DEF_OP(StoreFlag) {
mov(byte [STATE + (offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag)], al);
}
Xbyak::RegExp JITCore::GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
Xbyak::RegExp X86JITCore::GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
if (Offset.IsInvalid()) {
return Base;
} else {
@@ -568,8 +574,8 @@ DEF_OP(VStoreMemElement) {
}
#undef DEF_OP
void JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, Unhandled); // SRA specific, not supported on this backend
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -6,7 +12,7 @@ static void PrintValue(uint64_t Value) {
LogMan::Msg::D("Value: 0x%lx", Value);
}
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -34,10 +40,10 @@ DEF_OP(Break) {
case 4: { // HLT
// Time to quit
// Set our stack to the starting stack location
mov(rsp, qword [STATE + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)]);
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadStopHandlerAddress);
mov(TMP1, Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
break;
}
@@ -50,16 +56,16 @@ DEF_OP(Break) {
}
// This jump target needs to be a constant offset here
mov(TMP1, ThreadPauseHandlerAddress);
mov(TMP1, Dispatcher->ThreadPauseHandlerAddress);
jmp(TMP1);
}
else {
// If we don't have a gdb server attached then....crash?
// Treat this case like HLT
mov(rsp, qword [STATE + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)]);
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadStopHandlerAddress);
mov(TMP1, Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
}
break;
@@ -125,8 +131,8 @@ DEF_OP(Print) {
}
#undef DEF_OP
void JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -67,8 +73,8 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VectorZero) {
auto Dst = GetDst(Node);
vpxor(Dst, Dst, Dst);
@@ -1869,8 +1875,8 @@ DEF_OP(VTBL1) {
}
#undef DEF_OP
void JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
@@ -1,3 +1,10 @@
/*
$info$
tags: glue|block-database
desc: Stores information about blocks, and provides C++ implementations to lookup the blocks
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/LookupCache.h"
+150 -51
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: frontend|x86-to-ir, opcodes|dispatcher-implementations
desc: Handles x86/64 ops to IR, no-pf opt, local-flags opt
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -275,21 +282,58 @@ void OpDispatchBuilder::SecondaryALUOp(OpcodeArgs) {
#undef OPD
// X86 basic ALU ops just do the operation between the destination and a single source
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto ALUOp = _Add(Dest, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
uint8_t Size = GetDstSize(Op);
OrderedNode *Result{};
OrderedNode *Dest{};
OrderedNode *Result = ALUOp;
if (DestIsLockedMem(Op)) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
switch (IROp) {
case FEXCore::IR::IROps::OP_ADD: {
Dest = _AtomicFetchAdd(DestMem, Src, Size);
Result = _Add(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_SUB: {
Dest = _AtomicFetchSub(DestMem, Src, Size);
Result = _Sub(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_OR: {
Dest = _AtomicFetchOr(DestMem, Src, Size);
Result = _Or(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_AND: {
Dest = _AtomicFetchAnd(DestMem, Src, Size);
Result = _And(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_XOR: {
Dest = _AtomicFetchXor(DestMem, Src, Size);
Result = _Xor(Dest, Src);
break;
}
default: LogMan::Msg::A("Unknown Atomic IR Op: %d", IROp); break;
}
}
else {
Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto ALUOp = _Add(Dest, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
StoreResult(GPRClass, Op, Result, -1);
// Store result masks, but we need to
if (RequiresMask && Size < 4) {
Result = _Bfe(Size, Size * 8, 0, ALUOp);
Result = ALUOp;
StoreResult(GPRClass, Op, Result, -1);
// Store result masks, but we need to
if (RequiresMask && Size < 4) {
Result = _Bfe(Size, Size * 8, 0, ALUOp);
}
}
// Flags set
@@ -317,43 +361,59 @@ void OpDispatchBuilder::SecondaryALUOp(OpcodeArgs) {
template<uint32_t SrcIndex>
void OpDispatchBuilder::ADCOp(OpcodeArgs) {
LogMan::Throw::A(!DestIsLockedMem(Op), "Can't handle LOCK on ADC\n");
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
uint8_t Size = GetDstSize(Op);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
auto ALUOp = _Add(Src, CF);
auto ALUOp = _Add(Dest, Src);
auto Carry = _Add(ALUOp, CF);
OrderedNode *Result{};
OrderedNode *Before{};
if (DestIsLockedMem(Op)) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
Before = _AtomicFetchAdd(DestMem, ALUOp, Size);
Result = _Add(Before, ALUOp);
}
else {
Before = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
Result = _Add(Before, ALUOp);
StoreResult(GPRClass, Op, Result, -1);
}
OrderedNode *Result = Carry;
StoreResult(GPRClass, Op, Result, -1);
auto Size = GetDstSize(Op);
if (Size < 4)
Result = _Bfe(Size, Size * 8, 0, Carry);
GenerateFlags_ADC(Op, Result, Dest, Src, CF);
Result = _Bfe(Size, Size * 8, 0, Result);
GenerateFlags_ADC(Op, Result, Before, Src, CF);
}
template<uint32_t SrcIndex>
void OpDispatchBuilder::SBBOp(OpcodeArgs) {
LogMan::Throw::A(!DestIsLockedMem(Op), "Can't handle LOCK on SBB\n");
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[SrcIndex], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto Size = GetDstSize(Op);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
auto ALUOp = _Add(Src, CF);
auto ALUOp = _Sub(Dest, Src);
auto Carry = _Sub(ALUOp, CF);
OrderedNode *Result = Carry;
StoreResult(GPRClass, Op, Result, -1);
auto Size = GetDstSize(Op);
if (Size < 4) {
Result = _Bfe(Size, Size * 8, 0, Carry);
OrderedNode *Result{};
OrderedNode *Before{};
if (DestIsLockedMem(Op)) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
Before = _AtomicFetchSub(DestMem, ALUOp, Size);
Result = _Sub(Before, ALUOp);
}
GenerateFlags_SBB(Op, Result, Dest, Src, CF);
else {
Before = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
Result = _Sub(Before, ALUOp);
StoreResult(GPRClass, Op, Result, -1);
}
if (Size < 4) {
Result = _Bfe(Size, Size * 8, 0, Result);
}
GenerateFlags_SBB(Op, Result, Before, Src, CF);
}
void OpDispatchBuilder::PUSHOp(OpcodeArgs) {
@@ -948,8 +1008,6 @@ void OpDispatchBuilder::CondJUMPOp(OpcodeArgs) {
// Fallback
{
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
auto CondJump = _CondJump(SrcCond);
// Taking branch block
@@ -1010,7 +1068,6 @@ void OpDispatchBuilder::CondJUMPRCXOp(OpcodeArgs) {
auto CurrentBlock = GetCurrentBlock();
{
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
auto CondJump = _CondJump(CondReg, {COND_EQ});
// Taking branch block
@@ -1087,7 +1144,6 @@ void OpDispatchBuilder::LoopOp(OpcodeArgs) {
auto FalseBlock = JumpTargets.find(Op->PC + Op->InstSize);
{
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
auto CondJump = _CondJump(SrcCond);
// Taking branch block
@@ -1127,7 +1183,6 @@ void OpDispatchBuilder::LoopOp(OpcodeArgs) {
}
void OpDispatchBuilder::JUMPOp(OpcodeArgs) {
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
BlockSetRIP = true;
// This is just an unconditional relative literal jump
@@ -1302,6 +1357,7 @@ void OpDispatchBuilder::XCHGOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
if (DestIsMem(Op)) {
HandledLock = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Dest = AppendSegmentOffset(Dest, Op->Flags);
@@ -1514,8 +1570,10 @@ void OpDispatchBuilder::MOVOffsetOp(OpcodeArgs) {
void OpDispatchBuilder::CPUIDOp(OpcodeArgs) {
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _CPUID(Src);
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Leaf = _LoadContext(4, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RCX]), GPRClass);
auto Res = _CPUID(Src, Leaf);
OrderedNode *Result_Lower = _ExtractElementPair(Res, 0);
OrderedNode *Result_Upper = _ExtractElementPair(Res, 1);
@@ -2408,7 +2466,7 @@ void OpDispatchBuilder::BTOp(OpcodeArgs) {
else {
// Load the address to the memory location
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Dest = AppendSegmentOffset(Dest, Op->Flags);
// Get the bit selection from the src
OrderedNode *BitSelect = _Bfe(3, 0, Src);
@@ -2472,6 +2530,7 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
else {
// Load the address to the memory location
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Dest = AppendSegmentOffset(Dest, Op->Flags);
// Get the bit selection from the src
OrderedNode *BitSelect = _Bfe(3, 0, Src);
@@ -2489,6 +2548,7 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
BitMask = _Not(BitMask);
if (DestIsLockedMem(Op)) {
HandledLock = true;
// XXX: Technically this can optimize to an AArch64 ldclralb
// We don't current support this IR op though
Result = _AtomicFetchAnd(MemoryLocation, BitMask, 1);
@@ -2549,7 +2609,7 @@ void OpDispatchBuilder::BTSOp(OpcodeArgs) {
else {
// Load the address to the memory location
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Dest = AppendSegmentOffset(Dest, Op->Flags);
// Get the bit selection from the src
OrderedNode *BitSelect = _Bfe(3, 0, Src);
@@ -2565,6 +2625,7 @@ void OpDispatchBuilder::BTSOp(OpcodeArgs) {
OrderedNode *BitMask = _Lshl(_Constant(1), BitSelect);
if (DestIsLockedMem(Op)) {
HandledLock = true;
Result = _AtomicFetchOr(MemoryLocation, BitMask, 1);
// Now shift in to the correct bit location
Result = _Lshr(Result, BitSelect);
@@ -2623,6 +2684,7 @@ void OpDispatchBuilder::BTCOp(OpcodeArgs) {
else {
// Load the address to the memory location
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Dest = AppendSegmentOffset(Dest, Op->Flags);
// Get the bit selection from the src
OrderedNode *BitSelect = _Bfe(3, 0, Src);
@@ -2638,6 +2700,7 @@ void OpDispatchBuilder::BTCOp(OpcodeArgs) {
OrderedNode *BitMask = _Lshl(_Constant(1), BitSelect);
if (DestIsLockedMem(Op)) {
HandledLock = true;
Result = _AtomicFetchXor(MemoryLocation, BitMask, 1);
// Now shift in to the correct bit location
Result = _Lshr(Result, BitSelect);
@@ -2777,7 +2840,6 @@ void OpDispatchBuilder::MULOp(OpcodeArgs) {
}
void OpDispatchBuilder::NOTOp(OpcodeArgs) {
LogMan::Throw::A(!DestIsLockedMem(Op), "Can't handle LOCK on NOT\n");
uint8_t Size = GetSrcSize(Op);
OrderedNode *MaskConst{};
if (Size == 8) {
@@ -2787,9 +2849,17 @@ void OpDispatchBuilder::NOTOp(OpcodeArgs) {
MaskConst = _Constant((1ULL << (Size * 8)) - 1);
}
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
Src = _Xor(Src, MaskConst);
StoreResult(GPRClass, Op, Src, -1);
if (DestIsLockedMem(Op)) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
_AtomicXor(DestMem, MaskConst, Size);
}
else {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
Src = _Xor(Src, MaskConst);
StoreResult(GPRClass, Op, Src, -1);
}
}
void OpDispatchBuilder::XADDOp(OpcodeArgs) {
@@ -2815,6 +2885,8 @@ void OpDispatchBuilder::XADDOp(OpcodeArgs) {
GenerateFlags_ADD(Op, Result, Dest, Src);
}
else {
HandledLock = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
Dest = AppendSegmentOffset(Dest, Op->Flags);
auto Before = _AtomicFetchAdd(Dest, Src, GetSrcSize(Op));
StoreResult(GPRClass, Op, Op->Src[0], Before, -1);
Result = _Add(Before, Src); // Seperate result just for flags
@@ -2945,7 +3017,9 @@ void OpDispatchBuilder::INCOp(OpcodeArgs) {
bool IsLocked = DestIsLockedMem(Op);
if (IsLocked) {
HandledLock = true;
auto DestAddress = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestAddress = AppendSegmentOffset(DestAddress, Op->Flags);
Dest = _AtomicFetchAdd(DestAddress, OneConst, GetSrcSize(Op));
} else {
@@ -2974,7 +3048,9 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
bool IsLocked = DestIsLockedMem(Op);
if (IsLocked) {
HandledLock = true;
auto DestAddress = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestAddress = AppendSegmentOffset(DestAddress, Op->Flags);
Dest = _AtomicFetchSub(DestAddress, OneConst, GetSrcSize(Op));
} else {
@@ -4269,8 +4345,6 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
// 0x80064000
// 0x80064000
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
if (Op->Dest.TypeNone.Type == FEXCore::X86Tables::DecodedOperand::TYPE_GPR) {
OrderedNode *Src1{};
OrderedNode *Src1Lower{};
@@ -4295,14 +4369,14 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
// AKA if they match then don't touch RAX value
// Otherwise set it to the rm operand
OrderedNode *CASResult = _Select(FEXCore::IR::COND_EQ,
Src1Lower, Src3,
Src1Lower, Src3Lower,
Src3Lower, Src1Lower);
// Op1 = RAX == Op1 ? Op2 : Op1
// If they match then set the rm operand to the input
// else don't set the rm operand
OrderedNode *DestResult = _Select(FEXCore::IR::COND_EQ,
Src1Lower, Src3,
Src1Lower, Src3Lower,
Src2, Src1);
// Store in to GPR Dest
@@ -4311,7 +4385,7 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
// This allows us to only hit the ZEXT case on failure
OrderedNode *RAXResult = _Select(FEXCore::IR::COND_EQ,
CASResult, Src3Lower,
Src3, Src3Lower);
Src3, Src1Lower);
// When the size is 4 we need to make sure not zext the GPR when the comparison fails
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RAX]), RAXResult);
@@ -4331,6 +4405,8 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
GenerateFlags_SUB(Op, Result, Src3Lower, CASResult);
}
else {
HandledLock = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
OrderedNode *Src3{};
OrderedNode *Src3Lower{};
if (GPRSize == 8 && Size == 4) {
@@ -4380,6 +4456,7 @@ void OpDispatchBuilder::CMPXCHGPairOp(OpcodeArgs) {
// Unlike CMPXCHG, the destination can only be a memory location
auto Size = GetSrcSize(Op);
HandledLock = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
// If this is a memory location then we want the pointer to it
OrderedNode *Src1 = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
@@ -5751,7 +5828,9 @@ void OpDispatchBuilder::ALUOp(OpcodeArgs) {
OrderedNode *Dest{};
if (DestIsLockedMem(Op)) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
switch (IROp) {
case FEXCore::IR::IROps::OP_ADD: {
Dest = _AtomicFetchAdd(DestMem, Src, Size);
@@ -5844,6 +5923,7 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
}
case 0x0B:
Reason = 5;
break;
case 0xCC:
Reason = 6;
setRIP = true;
@@ -6787,11 +6867,20 @@ void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037
auto NewFCW = _Constant(16, 0x037);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
// Init FSW to 0
SetX87Top(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(_Constant(0));
// XXX: Add FTW support
}
template<size_t width, bool Integer, OpDispatchBuilder::FCOMIFlags whichflags, bool poptwice>
@@ -7027,6 +7116,7 @@ void OpDispatchBuilder::X87ATAN(OpcodeArgs) {
void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
@@ -7072,6 +7162,7 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
auto Size = GetDstSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
@@ -7200,6 +7291,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
auto Size = GetDstSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
OrderedNode *Top = GetX87Top();
{
@@ -7278,11 +7370,15 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
ST0Location = _Add(ST0Location, _Constant(8));
auto topBytes = _VExtractElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
// reset to default
FNINIT(Op);
}
void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
@@ -7450,6 +7546,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
// Saves 512bytes to the memory location provided
// Header changes depending on if REX.W is set or not
@@ -7550,6 +7647,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1, false);
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
@@ -9127,6 +9225,7 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::PABS<2>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::PABS<4>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::PABS<4>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVAPSOp},
{OPD(PF_38_66, 0x3B), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
+1 -1
View File
@@ -475,8 +475,8 @@ public:
OrderedNode *GetPackedRFLAG(bool Lower8);
void SetMultiblock(bool _Multiblock) { Multiblock = _Multiblock; }
bool GetMultiblock() { return Multiblock; }
bool HandledLock = false;
private:
bool DecodeFailure{false};
FEXCore::IR::IROp_IRHeader *Current_Header{};
+8 -1
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include <cstring>
@@ -28,7 +35,7 @@ X86GeneratedCode::~X86GeneratedCode() {
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
FEXCore::Config::Value<bool> Is64BitMode{FEXCore::Config::CONFIG_IS64BIT_MODE, 0};
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <FEXCore/Config/Config.h>
+7
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: frontend|x86-tables ~ Metadata that drives the frontend x86/64 decoding
tags: frontend|x86-tables
$end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/Context.h>
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
@@ -279,13 +285,13 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(BaseOps, BaseOpTable, sizeof(BaseOpTable) / sizeof(BaseOpTable[0]));
GenerateTable(BaseOps, BaseOpTable, std::size(BaseOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(BaseOps, BaseOpTable_64, sizeof(BaseOpTable_64) / sizeof(BaseOpTable_64[0]));
GenerateTable(BaseOps, BaseOpTable_64, std::size(BaseOpTable_64));
}
else {
GenerateTable(BaseOps, BaseOpTable_32, sizeof(BaseOpTable_32) / sizeof(BaseOpTable_32[0]));
GenerateTable(BaseOps, BaseOpTable_32, std::size(BaseOpTable_32));
}
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -38,6 +44,6 @@ void InitializeDDDTables() {
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
};
GenerateTable(DDDNowOps, DDDNowOpTable, sizeof(DDDNowOpTable) / sizeof(DDDNowOpTable[0]));
GenerateTable(DDDNowOps, DDDNowOpTable, std::size(DDDNowOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -20,6 +26,6 @@ void InitializeEVEXTables() {
{0xE7, 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
};
GenerateTable(EVEXTableOps, EVEXTable, sizeof(EVEXTable) / sizeof(EVEXTable[0]));
GenerateTable(EVEXTableOps, EVEXTable, std::size(EVEXTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -54,7 +60,7 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0x25), 1, X86InstInfo{"PMOVSXDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x28), 1, X86InstInfo{"PMULDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x29), 1, X86InstInfo{"PCMPEQQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x2B), 1, X86InstInfo{"PACKUSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x30), 1, X86InstInfo{"PMOVZXBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -92,6 +98,6 @@ void InitializeH0F38Tables() {
};
#undef OPD
GenerateTable(H0F38TableOps, H0F38Table, sizeof(H0F38Table) / sizeof(H0F38Table[0]));
GenerateTable(H0F38TableOps, H0F38Table, std::size(H0F38Table));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -48,10 +54,10 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(H0F3ATableOps, H0F3ATable, sizeof(H0F3ATable) / sizeof(H0F3ATable[0]));
GenerateTable(H0F3ATableOps, H0F3ATable, std::size(H0F3ATable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(H0F3ATableOps, H0F3ATable_64, sizeof(H0F3ATable_64) / sizeof(H0F3ATable_64[0]));
GenerateTable(H0F3ATableOps, H0F3ATable_64, std::size(H0F3ATable_64));
}
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -152,12 +158,12 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable, sizeof(PrimaryGroupOpTable) / sizeof(PrimaryGroupOpTable[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_64, sizeof(PrimaryGroupOpTable_64) / sizeof(PrimaryGroupOpTable_64[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
}
else {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_32, sizeof(PrimaryGroupOpTable_32) / sizeof(PrimaryGroupOpTable_32[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_32, std::size(PrimaryGroupOpTable_32));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -477,7 +483,7 @@ void InitializeSecondaryGroupTables() {
};
#undef OPD
GenerateTable(SecondInstGroupOps, SecondaryExtensionOpTable, sizeof(SecondaryExtensionOpTable) / sizeof(SecondaryExtensionOpTable[0]));
GenerateTable(SecondInstGroupOps, SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -46,6 +52,6 @@ void InitializeSecondaryModRMTables() {
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondModRMTableOps, SecondaryModRMExtensionOpTable, sizeof(SecondaryModRMExtensionOpTable) / sizeof(SecondaryModRMExtensionOpTable[0]));
GenerateTable(SecondModRMTableOps, SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -571,18 +577,18 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondBaseOps, TwoByteOpTable, sizeof(TwoByteOpTable) / sizeof(TwoByteOpTable[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable, std::size(TwoByteOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(SecondBaseOps, TwoByteOpTable_64, sizeof(TwoByteOpTable_64) / sizeof(TwoByteOpTable_64[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable_64, std::size(TwoByteOpTable_64));
}
else {
GenerateTable(SecondBaseOps, TwoByteOpTable_32, sizeof(TwoByteOpTable_32) / sizeof(TwoByteOpTable_32[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable_32, std::size(TwoByteOpTable_32));
}
GenerateTableWithCopy(RepModOps, RepModOpTable, sizeof(RepModOpTable) / sizeof(RepModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(RepNEModOps, RepNEModOpTable, sizeof(RepNEModOpTable) / sizeof(RepNEModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(OpSizeModOps, OpSizeModOpTable, sizeof(OpSizeModOpTable) / sizeof(OpSizeModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(RepModOps, RepModOpTable, std::size(RepModOpTable), SecondBaseOps);
GenerateTableWithCopy(RepNEModOps, RepNEModOpTable, std::size(RepNEModOpTable), SecondBaseOps);
GenerateTableWithCopy(OpSizeModOps, OpSizeModOpTable, std::size(OpSizeModOpTable), SecondBaseOps);
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -500,7 +506,7 @@ void InitializeVEXTables() {
};
#undef OPD
GenerateTable(VEXTableOps, VEXTable, sizeof(VEXTable) / sizeof(VEXTable[0]));
GenerateTable(VEXTableGroupOps, VEXGroupTable, sizeof(VEXGroupTable) / sizeof(VEXGroupTable[0]));
GenerateTable(VEXTableOps, VEXTable, std::size(VEXTable));
GenerateTable(VEXTableGroupOps, VEXGroupTable, std::size(VEXGroupTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#pragma once
#include <FEXCore/Debug/X86Tables.h>
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -254,6 +260,6 @@ void InitializeX87Tables() {
#undef OPD
#undef OPDReg
GenerateX87Table(X87Ops, X87OpTable, sizeof(X87OpTable) / sizeof(X87OpTable[0]));
GenerateX87Table(X87Ops, X87OpTable, std::size(X87OpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -119,7 +125,7 @@ void InitializeXOPTables() {
};
#undef OPD
GenerateTable(XOPTableOps, XOPTable, sizeof(XOPTable) / sizeof(XOPTable[0]));
GenerateTable(XOPTableGroupOps, XOPGroupTable, sizeof(XOPGroupTable) / sizeof(XOPGroupTable[0]));
GenerateTable(XOPTableOps, XOPTable, std::size(XOPTable));
GenerateTable(XOPTableGroupOps, XOPGroupTable, std::size(XOPGroupTable));
}
}
+13 -7
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: glue|thunks ~ FEXCore side of thunks: Registration, Lookup
tags: glue|thunks
$end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include "Thunks.h"
@@ -23,7 +30,7 @@ static thread_local FEXCore::Core::InternalThreadState *Thread;
namespace FEXCore {
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
class ThunkHandler_impl final: public ThunkHandler {
@@ -41,9 +48,8 @@ namespace FEXCore {
Set arg0/1 to arg regs, use CTX::HandleCallback to handle the callback
*/
static void CallCallback(void *callback, void *arg0, void* arg1) {
Thread->State.State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->State.State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CTX->HandleCallback((uintptr_t)callback);
}
@@ -57,10 +63,10 @@ namespace FEXCore {
auto Name = Args->Name;
auto CallbackThunks = Args->CallbackThunks;
auto SOName = CTX->Config.ThunkLibsPath + "/" + (const char*)Name + "-host.so";
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
LogMan::Msg::D("Load lib: %s -> %s", Name, SOName.c_str());
auto Handle = dlopen(SOName.c_str(), RTLD_LOCAL | RTLD_NOW);
if (!Handle) {
@@ -78,7 +84,7 @@ namespace FEXCore {
LogMan::Msg::E("Load lib: failed to find export %s", InitSym.c_str());
return;
}
auto Exports = InitFN((void*)&CallCallback, CallbackThunks);
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: glue|thunks
$end_info$
*/
#pragma once
#include <FEXCore/IR/IR.h>
+5 -1
View File
@@ -1138,7 +1138,11 @@
"DestClass": "GPRPair",
"FixedDestSize": "8",
"NumElements": "2",
"SSAArgs": "1"
"SSAArgs": "2",
"SSANames": [
"Function",
"Leaf"
]
},
"Bfi": {
+9 -1
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: ir|dumper ~ IR -> Text
tags: ir|dumper
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
@@ -24,6 +31,7 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const*
*out << "#0x" << std::hex << Arg;
}
[[maybe_unused]]
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, const char* Arg) {
*out << Arg;
}
@@ -91,7 +99,7 @@ static void PrintArg(std::stringstream *out, IRListView const* IR, OrderedNodeWr
*out << "%ssa" << std::to_string(Arg.ID());
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(Arg.ID());
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(GPRFixed"; break;
+7
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: ir|emitter ~ C++ Functions to generate IR. See IR.json for spec.
tags: ir|emitter
$end_info$
*/
#include <FEXCore/IR/IREmitter.h>
namespace FEXCore::IR {
+7
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: ir|parser ~ Text -> IR
tags: ir|parser
$end_info$
*/
#include <string>
#include <vector>
#include <istream>
+9 -1
View File
@@ -1,3 +1,11 @@
/*
$info$
meta: ir|opts ~ IR to IR Optimization
tags: ir|opts
desc: Defines which passes are run, and runs them
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -7,7 +15,7 @@
namespace FEXCore::IR {
void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllocation) {
FEXCore::Config::Value<bool> DisablePasses{FEXCore::Config::CONFIG_DEBUG_DISABLE_OPTIMIZATION_PASSES, false};
FEX_CONFIG_OPT(DisablePasses, O0);
if (!DisablePasses()) {
InsertPass(CreateContextLoadStoreElimination());
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: ir|opts
$end_info$
*/
#pragma once
#include <FEXCore/IR/IntrusiveIRList.h>
+14 -7
View File
@@ -1,3 +1,11 @@
/*
$info$
tags: ir|opts
desc: ConstProp, ZExt elim, addressgen coalesce, const pooling, fcmp reduction, const inlining
$end_info$
*/
#if defined(_M_ARM_64)
//aarch64 heuristics
#include "aarch64/assembler-aarch64.h"
@@ -158,13 +166,8 @@ bool ConstProp::Run(IREmitter *IREmit) {
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
auto Header = CurrentIR.GetHeader();
auto OriginalWriteCursor = IREmit->GetWriteCursor();
auto HeaderOp = CurrentIR.GetHeader();
{
// constants are pooled per block
@@ -197,7 +200,8 @@ bool ConstProp::Run(IREmitter *IREmit) {
auto SelectOp = SelectOpHdr->CW<IR::IROp_Select>();
// the value isn't used after the select otherwise
if (SelectOpHdr->Op == OP_SELECT && SelectOpNode->NumUses == 1
// make sure the sizes match
if (SelectOpHdr->Size == UnaryOpHdr->Size && SelectOpHdr->Op == OP_SELECT && SelectOpNode->NumUses == 1
&& IREmit->IsValueConstant(SelectOp->TrueVal)
&& IREmit->IsValueConstant(SelectOp->FalseVal)) {
@@ -743,8 +747,11 @@ bool ConstProp::Run(IREmitter *IREmit) {
}
}
}
break;
}
default: break;
default:
break;
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: ir|opts
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: Transforms ContextLoad/Store to temporaries, similar to mem2reg
$end_info$
*/
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -598,7 +605,8 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
}
}
else if (IROp->Op == OP_STORECONTEXTINDEXED ||
IROp->Op == OP_LOADCONTEXTINDEXED) {
IROp->Op == OP_LOADCONTEXTINDEXED ||
IROp->Op == OP_SYSCALL) {
// We can't track through these
ResetClassificationAccesses(&LocalInfo);
}
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: Cross block store-after-store elimination
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -28,7 +35,7 @@ bool IsFullGPR(uint32_t Offset, uint8_t Size) {
return false;
if (Offset & 7)
return false;
if (Offset < 8 || Offset >= (17 * 8))
return false;
@@ -36,7 +43,7 @@ bool IsFullGPR(uint32_t Offset, uint8_t Size) {
}
bool IsGPR(uint32_t Offset) {
if (Offset < 8 || Offset >= (17 * 8))
return false;
@@ -48,7 +55,7 @@ uint32_t GPRBit(uint32_t Offset) {
return 0;
}
return 1 << ((Offset - 8)/8);
return 1 << ((Offset - 8)/8);
}
struct FPRInfo {
@@ -58,9 +65,9 @@ struct FPRInfo {
};
bool IsFPR(uint32_t Offset) {
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto end = offsetof(FEXCore::Core::ThreadState, State.xmm[17][0]);
auto begin = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto end = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[17][0]);
if (Offset < begin || Offset >= end)
return false;
@@ -84,7 +91,7 @@ uint64_t FPRBit(uint32_t Offset, uint32_t Size) {
return 0;
}
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto begin = offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0]);
auto regn = (Offset - begin)/16;
auto bitn = regn * 3;
@@ -138,7 +145,7 @@ bool DeadStoreElimination::Run(IREmitter *IREmit) {
//// Flags ////
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
auto& BlockInfo = InfoMap[BlockNode];
BlockInfo.flag.writes |= 1UL << Op->Flag;
@@ -203,7 +210,7 @@ bool DeadStoreElimination::Run(IREmitter *IREmit) {
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
auto CodeBlock = BlockIROp->C<IROp_CodeBlock>();
auto IROp = CurrentIR.GetNode(CurrentIR.GetNode(CodeBlock->Last)->Header.Previous)->Op(CurrentIR.GetData());
if (IROp->Op == OP_JUMP) {
@@ -220,8 +227,8 @@ bool DeadStoreElimination::Run(IREmitter *IREmit) {
// Flags that are written by the next block can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
//// GPRs ////
// stores to remove are written by the next block but not read
@@ -258,7 +265,7 @@ bool DeadStoreElimination::Run(IREmitter *IREmit) {
// Flags that are written by the next blocks can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
//// GPRs ////
// stores to remove are written by the next blocks but not read
@@ -300,7 +307,7 @@ bool DeadStoreElimination::Run(IREmitter *IREmit) {
}
} else if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->C<IR::IROp_StoreContext>();
auto& BlockInfo = InfoMap[BlockNode];
//// GPRs ////
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: Sorts the ssa storage in memory, needed for RA and others
$end_info$
*/
#include "Common/MathUtils.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: Sanity checking pass
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/Context/Context.h"
@@ -56,7 +63,7 @@ bool IRValidation::Run(IREmitter *IREmit) {
}
NodeIsLive.Set(1); // IRHEADER
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
@@ -271,8 +278,8 @@ bool IRValidation::Run(IREmitter *IREmit) {
if (HadWarning) {
Out << "Warnings:" << std::endl << Warnings.str() << std::endl;
}
fprintf(stderr, "%s", Out.str().c_str());
LogMan::Msg::E("%s", Out.str().c_str());
}
return false;
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: Sanity checking pass
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -1,3 +1,10 @@
/*
$info$
tags: ir|opts
desc: This is not used right now, possibly broken
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
Loaded 100 of 263 files, more files were not shown because too many files have changed in this diff. Show more