Compare commits

..
229 Commits
Author SHA1 Message Date
Ryan Houdek 27072d2853 Docs: Update for release FEX-2110 2021-10-08 16:29:35 -07:00
Ryan Houdek 0dc8e23342 Merge pull request #1301 from Sonicadvance1/thunk_versioning
Thunks: Support versioned libraries
2021-10-08 03:01:04 -07:00
Ryan Houdek b102714d5c Thunks: Support versioned libraries
We can't expect users to have development libraries installed.
Load the versioned libraries if they exist instead

Also load them in global namespace, which is required for getting
symbols.

Behaviour on x86-64 host seems sporatic here, not sure if unintended
feature.
Doesn't quite behave the same on AArch64 host
2021-10-08 02:41:37 -07:00
Ryan Houdek 72125c9ccf Merge pull request #1300 from Sonicadvance1/thunksdb_global_file
Thunks: Install a global thunksDB for our current thunks
2021-10-08 02:02:28 -07:00
Ryan Houdek 9a07b550f3 Merge pull request #1298 from Sonicadvance1/libgl_thunks
Thunks: Adds a few missing libGL thunk functions
2021-10-08 02:02:17 -07:00
Ryan Houdek 351412a3e3 Thunks: Install a global thunksDB for our current thunks
This covers the x86_64 definitions, in the future we can have the 32-bit
versions in here as well.
2021-10-07 20:20:34 -07:00
Ryan Houdek 842ab169ce Thunks: Adds a few missing libGL thunk functions
These should really be autogenerated by we aren't there yet.
For now this misses some function aliases which fixes Mangohud when
using GL thunking.
2021-10-07 19:15:57 -07:00
Ryan Houdek b3efb1d2d7 Merge pull request #1296 from Sonicadvance1/vulkan_thunks
Vulkan thunks
2021-10-07 12:49:08 -07:00
Ryan Houdek 0d0ce38050 Thunks: Disable malloc libraries
Keeping this in the commit history to go back to later.
While these are required for static-pie builds to work, with the glibc
bug we can't use those yet.
Disable for now since it is is unnecessary.
2021-10-06 11:53:48 -07:00
Ryan Houdek cb21f52f93 Thunks: Have the ThunkHandler load the FEX malloc symbol libraries
If the thunk configuration is enabled then preemptively load the
libraries. Since they will need to be loaded for any thunking library.

This is because we need to ALWAYS share the allocators to the thunks.
2021-10-05 23:51:26 -07:00
Ryan Houdek 14b0cc5af4 Thunks: Wires up all the new thunks to the generators 2021-10-05 23:51:26 -07:00
Ryan Houdek c0fc4c4623 Thunks: Adds libvulkan
This is specifically the device loader rather than the libvulkan loader
library.
This is meant to override what is provided in the ICD files, not the
loader.

There's no real need to replace the loader, it's quite complex
2021-10-05 23:51:26 -07:00
Ryan Houdek c3230f6a91 Thunks: Adds libdrm 2021-10-05 23:51:26 -07:00
Ryan Houdek 4a3bd7cd13 Thunks: Adds libxshmfence 2021-10-05 23:51:26 -07:00
Ryan Houdek f98627da32 Thunks: Adds libxcb_xfixes 2021-10-05 23:51:26 -07:00
Ryan Houdek f3c20f2743 Thunks: Adds libxcb_sync 2021-10-05 23:51:26 -07:00
Ryan Houdek ec6140fde0 Thunks: Adds libxcb_shm 2021-10-05 23:51:25 -07:00
Ryan Houdek c5490821fc Thunks: Adds libxcb_randr 2021-10-05 23:51:25 -07:00
Ryan Houdek c0aad64578 Thunks: Adds libxcb_present 2021-10-05 23:51:25 -07:00
Ryan Houdek 7381240fb6 Thunks: Adds libxcb_glx 2021-10-05 23:51:25 -07:00
Ryan Houdek 3ea8e4864d Thunks: Adds libxcb_dri3 2021-10-05 23:51:25 -07:00
Ryan Houdek 7bf8d09391 Thunks: Adds libxcb_dri2 2021-10-05 23:51:25 -07:00
Ryan Houdek 798c6772b7 Thunks: Adds libxcb 2021-10-05 23:51:25 -07:00
Ryan Houdek 29debcda3d Thunks: Adds FEX malloc libraries
These are required to expose FEX's allocators to the guest.
Which is required when we are compiled with jemalloc, otherwise
the libraries will crash on allocations.

This needs to be done in stages otherwise glibc will load the library
and as it is setting up symbols, replace malloc, which isn't set
currently and cause a crash
2021-10-02 18:12:07 -07:00
Ryan Houdek 39c1751215 Thunks: Adds a few more features to Generator python file
This will be necessary for Vulkan
2021-10-02 18:05:03 -07:00
Ryan Houdek 9451bc5273 Thunks: Pass Vulkan XML to the thunks generators 2021-10-02 18:03:24 -07:00
Ryan Houdek e2b24f7f59 Thunks: Support thunk init function on host
Allows library to have an additional initialization function
2021-10-02 17:54:30 -07:00
Ryan Houdek 4589876ebc Adds Vulkan-Docs repo to externals 2021-10-02 17:09:23 -07:00
Ryan Houdek e0b878f1fd Merge pull request #1289 from Sonicadvance1/thunks_debugging_changes
Thunks: Some minor X related thunk changes
2021-10-02 11:13:22 -07:00
Ryan Houdek 073224ffea Merge pull request #1287 from Sonicadvance1/expand_thunk_Xext
Thunks: Expands Xext thunked functions listo
2021-10-02 11:13:14 -07:00
Ryan Houdek 00511c16a4 Merge pull request #1286 from Sonicadvance1/expand_thunk_asound
Thunks: Expands what asound thunking supports
2021-10-02 11:13:04 -07:00
Ryan Houdek 66c7fdb6de Merge pull request #1293 from Sonicadvance1/thunks_database
Thunks: Adds a new Thunks database config file
2021-10-02 11:11:55 -07:00
Ryan Houdek a48ed376d8 Merge pull request #1291 from Sonicadvance1/thunk_file_search
Thunks: Makes file searches a bit easier
2021-10-02 11:10:08 -07:00
Ryan Houdek 98bee5a6dd Merge pull request #1290 from Sonicadvance1/thunk_init_constructor
Thunks: Adds init helpers with function call
2021-10-02 11:09:23 -07:00
Ryan Houdek 8d13261ab6 Merge pull request #1292 from Sonicadvance1/thunks_script
Thunks: ThunkHelpers script improvements
2021-10-02 11:08:23 -07:00
Ryan Houdek 2b757a9b9a Merge pull request #1284 from Sonicadvance1/fix_struct_match
StructVerifier: Fix struct match and minor fixes
2021-10-02 11:07:51 -07:00
Ryan Houdek 08f56540d2 Merge pull request #1282 from Sonicadvance1/InterpreterLoadStores
Interpreter: Changes basic loadstores to sized accesses
2021-10-02 11:07:38 -07:00
Ryan Houdek fb01e8bf28 Merge pull request #1283 from Sonicadvance1/minor_core_cleanup
Core: Minor documentation and code splitting
2021-10-02 11:06:56 -07:00
Ryan Houdek d5f9f43ab0 Merge pull request #1281 from Sonicadvance1/SupportHostEnv
Adds support for setting host environment variables from config
2021-10-02 11:06:22 -07:00
Ryan Houdek 9d460e807c Merge pull request #1288 from Sonicadvance1/libclang_definition_extract
Scripts: Adds a new script for extracting function definitions
2021-10-02 11:05:36 -07:00
Ryan Houdek 185f265f5c Thunks: Adds a new Thunks database config file
This adds a new `ThunksDB.json` file to the config folder.
This file lets users describe thunks in a meaningful way without
duplicating it amongst multiple configuration files.

eg:
```
{
  "DB": {
    "GL": {
      "Library" : "libGL-guest.so",
      "Depends": [
        "X11"
      ],
      "Overlay": [
        "/usr/lib/x86_64-linux-gnu/libGL.so",
        "/usr/lib/x86_64-linux-gnu/libGL.so.1",
        "/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
        "/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
        "/lib/x86_64-linux-gnu/libGL.so",
        "/lib/x86_64-linux-gnu/libGL.so.1",
        "/lib/x86_64-linux-gnu/libGL.so.1.2.0",
        "/lib/x86_64-linux-gnu/libGL.so.1.7.0"
      ]
    },
    "X11": {
      "Library": "libX11-guest.so",
      "Overlay": [
        "/usr/lib/x86_64-linux-gnu/libX11.so.6",
        "/lib/x86_64-linux-gnu/libX11.so.6"
      ]
    }
  }
}
```

This file lets the user describe the library with an nicer name, in this instance
`GL` instead of `libGL-guest.so`.
It also tracks depedencies, like how GL currently has a hard dependency on X11.
This allows the loader to automatically enable the dependencies if described.
The `Overlays` array is like the regular Thunks config file but now in this DB file.

With the DB file now describing the libraries, this allows us to then stick a lighter
description inside of the Thunk Config file.

```
{
 "ThunksDB": {
   "GL": 1
 }
}
```

With this example Thunk config file (Which can be configured per application), There is
a new property of name `ThunksDB`.
All this takes is key:value pairs which describe the user friendly library name and an Integer
to state if the thunk should be enabled or not.
This allows very quick toggling of thunks directly inside of the configuration files rather than
breaking the configuration to disable it.
2021-10-02 11:03:08 -07:00
Ryan Houdek 591fc001cc Merge pull request #1294 from neobrain/fix_jitsymbols_build
Build fix for ENABLE_JITSYMBOLS
2021-10-01 10:25:50 -07:00
Tony Wasserka d71add78ae Build fix for ENABLE_JITSYMBOLS 2021-10-01 17:20:26 +02:00
Ryan Houdek d93edb2da8 Merge pull request #1285 from Sonicadvance1/thunk_config_fixes
Thunks: Minor fixes to the config and loading
2021-09-30 22:25:33 -07:00
Ryan Houdek 2360f9cec1 Thunks: ThunkHelpers script improvements
Allows having a library filename be different from the name.
This will fix a quirk in future thunks where the library name and
filename don't match

Also allow custom callback unpacks. A future thunk will need this
2021-09-30 19:13:28 -07:00
Ryan Houdek 10d596314e Thunks: Makes file searches a bit easier
Instead of searching inside the thunk folder for Guest and Host files
with a library name attached to it, also search for ones That are just
named `Host.cpp` and `Guest.cpp`.

Makes quickly pounding out a bunch of thunks significantly less tedious.
2021-09-30 18:48:32 -07:00
Ryan Houdek cab0cf6a6b Thunks: Adds init helpers with function call
This will be used in the future.
2021-09-30 18:43:17 -07:00
Ryan Houdek 6162a8c7f4 Thunks: Some minor X related thunk changes
Makes it print to stderr instead of stdout.
Also sets the mutex symbols to something that can be debugged.
Found something linking to them but not actually using them.
2021-09-30 18:39:22 -07:00
Ryan Houdek 1b0d2bbf9f Scripts: Adds a new script for extracting function definitions
This is very useful for extracting function definitions for thunks.
Keep it in upstream to not get lost.

Sometimes it can munge a definition but it is usually fine.
2021-09-30 18:35:43 -07:00
Ryan Houdek 121023fb72 Thunks: Expands Xext thunked functions listo
Adds in the function definitions from extutil and Xlibint.

Xlibint has some particularly nasty functions that aren't supported.
2021-09-30 18:23:12 -07:00
Ryan Houdek c30cb87b01 Thunks: Expands what asound thunking supports
These are pulled from headers using an automated script.
Almost all functions are easy drop in, only a handful are commented out.

Works in every game that I've tested.
2021-09-30 18:20:20 -07:00
Ryan Houdek d8edbba71c Thunks: Minor fixes to the config and loading
There was a hard upper limit to 128 json elements, this has now been
removed.
The debugging text for the thunk overlay is now disabled. It can get
very spammy but it is useful for debugging purposes.

Fixes an issue where thunks would be entirely disabled if a RootFS isn't
set.
This meant debugging in a chroot or x86_64 host was breaking rootfs
2021-09-30 18:15:41 -07:00
Ryan Houdek 7c553f3508 Linux: Minor fixes to 32-bit epoll
epoll_pwait has a sigsetsize argument, which was unused in this case but
it caused strace to look ugly.
On epoll_ctl, don't write back the resulting event. The event isn't
written and depending on how the guest allocated the object, we could
have a 4byte overwrite.
2021-09-30 18:10:51 -07:00
Ryan Houdek 5b2d944886 Linux: Minor struct verifier and types fixes
fex-match annotation was actually not doing anything due to python typo.
Fixes the minor warnings that cropped up. Nothing actually broken
2021-09-30 18:08:05 -07:00
Ryan Houdek 9ab7de56ef Core: Minor documentation and code splitting
This will change slightly in the future. Clean this up and document.
2021-09-30 18:01:57 -07:00
Ryan Houdek fc46cb9390 Interpreter: Changes basic loadstores to sized accesses
This makes debugging a bit more clear as to what is happening on crash
2021-09-30 17:56:36 -07:00
Ryan Houdek 8cf4b285bf Adds support for setting host environment variables from config
This can be useful for edge case environment variable setting
2021-09-30 17:51:49 -07:00
Ryan Houdek 3744ec2a44 Merge pull request #1262 from phire/RAValidation
RA validation
2021-09-13 01:39:24 -07:00
Scott Mansell 486c62f77c RAValidation: Improve comments 2021-09-13 20:25:33 +12:00
Ryan Houdek 327c4d550a Merge pull request #1277 from phire/aotir_use_after_free
AOTIR: fix use after free
2021-09-12 23:36:55 -07:00
Scott Mansell f14b73689f AOTIR: fix use after free
These variables are owned by other parts of the code, and AOTIR
should not be freeing/deleting them.
2021-09-13 17:54:26 +12:00
Scott Mansell 93de38d13c FillRegister: Keep refrence to original SSA
This allows the RA Validation pass to verify that a spill slot contains
the correct SSA.

I was originally planning to do a IR equlivent test between the IR
before and after register allocation, mostly to catch this type of error.
But this approach was much faster to implement and gives 90%
of the benefits.
2021-09-13 15:43:26 +12:00
Scott Mansell 3107898f06 Fix ReplaceUsesWithAfter:
There were two issues:
 1. The OrderedNode *After overload created the wrong iterator
 2. Despite being named After, they were actually inclusive

Kept ReplaceAllUsesWithRange with the current inclusive behaviour and
adjusted the argument name to match
2021-09-13 15:43:26 +12:00
Ryan Houdek 21ff433999 Merge pull request #1276 from Sonicadvance1/implement_message_queue
Linux: Implements support for POSIX message queues on 32-bit
2021-09-12 17:46:03 -07:00
Ryan Houdek 217e4764c4 Linux: Implements support for POSIX message queues on 32-bit
These would have worked for 64-bit but wasn't working on 32-bit.
This now passes my unit test.

Relies on #1275 to be merged first.
Fixes #1260
2021-09-12 17:36:26 -07:00
Ryan Houdek f232dcebfc Merge pull request #1275 from Sonicadvance1/implement_timer_create
Linux: Implements support for timer_create
2021-09-12 17:34:06 -07:00
Ryan Houdek 76dd09ea60 Linux: Implements support for timer_create
Needed to fix rt_sigtimedwait to use the raw syscall.
Needed to have sigtimedwait and sigtimedwait_time64 parse siginfo_t
correctly.
glibc uses this to ensure it is sending the correct signal across from
their helper thread.

Needed to correctly parse sigval and sigevent for 32-bit.

Passes my unit test for timer_create
2021-09-12 17:25:12 -07:00
Ryan Houdek 368095f96f Merge pull request #1274 from Sonicadvance1/fix_timex
Linux: Fixes timex definition for 32-bit syscalls
2021-09-12 17:14:36 -07:00
Scott Mansell 2cd844bbcb RAValidation: Validate Spill slots
This can prove that the ssa in the Spill slots are consistant based on
control flow.

It's one major blind spot is that it can't prove the spill slot actually
contains the correct ssa, if it has the same wrong ssa on every cfg
path.
2021-09-13 11:56:39 +12:00
Scott Mansell 55a9ee702b Register Allocator Validation
This is a validation pass that attempts to prove the output from the RA pass is valid.

The current design should be able to prove that the RA result is internally
consistent. That no-matter what control flow path you take thought the control
flow graph, the physical registers and spill slots will always contain a single
possible SSA value.

It also checks that the SSA values in the IR actually line up with the SSA value
in the physical register.
2021-09-13 11:56:29 +12:00
Scott Mansell cb57797550 PassManager: Lookup pass by name 2021-09-13 11:46:11 +12:00
Ryan Houdek 1b6cbd39a4 Linux: Fixes timex definition for 32-bit syscalls
Fixes #1251
2021-09-11 18:39:46 -07:00
Ryan Houdek d004fee7e4 Merge pull request #1273 from Sonicadvance1/softfloat_x86_debug
Softfloat: Allow forcing use of some x87 on x86 host
2021-09-11 18:08:41 -07:00
Ryan Houdek 1fa3afd0c1 Merge pull request #1272 from Sonicadvance1/ensure_long
OpcodeDispatcher: Ensure some x87 templates get passed long constants
2021-09-11 18:08:29 -07:00
Ryan Houdek ba174e5c32 Merge pull request #1271 from Sonicadvance1/fix_wine_working_dir
FileManager: Allow reading real root
2021-09-11 18:08:20 -07:00
Ryan Houdek e91061f7ff Merge pull request #1268 from Sonicadvance1/fix_rlimit_x32
Linux: Fixes rlimit syscalls for 32-bit
2021-09-11 18:08:11 -07:00
Ryan Houdek a56463a7f1 Merge pull request #1269 from Sonicadvance1/change_app_config_shortcut
FEXConfig: Change shortcut for opening application profile
2021-09-11 18:08:01 -07:00
Ryan Houdek 56bddd3a22 Merge pull request #1266 from Sonicadvance1/pressure_vessel_checks
Config: Check for container-manager and redirect
2021-09-11 18:07:50 -07:00
Ryan Houdek 829db6c30d gvisor: Updates getdents test behaviour
This now passes on x86-64 host but still fails on AArch64 because we
don't emulate the getdents syscall
2021-09-11 05:51:10 -07:00
Ryan Houdek 7855b58c73 Softfloat: Allow forcing use of some x87 on x86 host
This is useful for testing the ops that are emulated using different
precision.
Haven't found anything that changes behaviour but useful to keep around
2021-09-11 05:42:18 -07:00
Ryan Houdek 2b91108255 OpcodeDispatcher: Ensure some x87 templates get passed long constants
Just to ensure these don't get truncated and intent
2021-09-11 05:41:14 -07:00
Ryan Houdek 593be950de FileManager: Allow reading real root
wine walks the file structure to ensure the cwdir is safe for use. It
does this with a combination of `getcwd` and `statx`.

Once it reachs `/` then in rootfs environments it would fail to find
directories we've deleted.
Thus making it impossible to find folders like `/mnt` and `/home`

Should be safe since it just gives the application a larger world view
2021-09-11 05:37:28 -07:00
Ryan Houdek 360c4a2060 Merge pull request #1267 from Sonicadvance1/flush_logs
FEXLoader: Flush log output
2021-09-11 05:34:33 -07:00
Ryan Houdek a5046e92cd Config: Check for container-manager and redirect
In the case of running inside of a container then we need to redirect
where we look for configuration.
Currently we only care about pressure vessel so we just redirect some
options to check inside of `/run/host/`

This will resolve an issue where installed thunks wouldn't be found.
2021-09-11 02:16:08 -07:00
Ryan Houdek 793b25f93b Thunks: Set default config option to our install location
These don't really need to be changed unless doing development
2021-09-11 02:16:08 -07:00
Ryan Houdek acfcfa127a FEXConfig: Change shortcut for opening application profile
CTRL+A is used for selecting all in a text box.
Don't override that, it's annoying
2021-09-11 02:13:00 -07:00
Ryan Houdek 35d09de8c6 Linux: Fixes rlimit syscalls for 32-bit
32-bit versions of these syscalls saturate on the upper limit.
Depending on which syscall it'll saturate signed or unsigned.

With set if the 32-bit value is UINT32_MAX then it'll saturate to the
maximum 64-bit value
2021-09-11 01:17:38 -07:00
Ryan Houdek 1f11e307ec FEXLoader: Flush log output
This was removed when we switched from FILE to raw fd.
This fixes an annoying issue where we would assert and not get any
output.
2021-09-11 00:53:49 -07:00
Ryan Houdek c9a62658e3 Merge pull request #1259 from Sonicadvance1/fix_fexmountdaemon_races
FEXMountDaemon: Fixes shutdown race conditions
2021-09-07 22:13:08 -07:00
Ryan Houdek 4ccd68af8a FEXMountDaemon: Fixes shutdown race conditions
This solves a problem where sometimes FEX would spin up a new process
while FEXMountDaemon was in the process of shutting down. Breaking
things on both sides.

Now the race conditions are squashed that I could see.

Also fixes one race where FEX is starting up and FEXMountDaemon is
spinning up. This case is where FEX managed to pull the lock file just
before it got deleted. Then sent the FEXMountDaemon a request to be
observed. With DGRAM sockets we would fire and forget. Use a STREAM with
a ack result so we know we can return.
2021-09-07 19:30:38 -07:00
Ryan Houdek cd4586b67d Docs: Update for release FEX-2109 2021-09-05 01:57:02 -07:00
Ryan Houdek 2c02dcac9f Merge pull request #1257 from CallumDev/caspair_fix_armv8
Fix unaligned CASPair on ARMv8.0
2021-09-05 01:35:43 -07:00
CallumDev dec512187e JIT: Remove nops in ARMv8.0 CASPair 2021-09-05 17:49:57 +09:30
Ryan Houdek dd34316562 Merge pull request #1256 from Sonicadvance1/stabilize_fexmountdaemon
FEXMountDaemon: Fixes dangling mounts problem
2021-09-05 01:13:01 -07:00
CallumDev 5f7532c569 Interpreter: Lower 4 byte CASPair to inline assembly 2021-09-05 17:30:18 +09:30
Ryan Houdek 0fa7af15b2 FEXMountDaemon: Fixes dangling mounts problem
The FEXMountDaemon no longer uses the inotify interface for refcounting
instances of FEX.
The inotify interface fails to send close events when an application
crashes. Which is either an API oversight or intentional choice.

Now to use FEXMountDaemon the FEX process must send the daemon a pipe
fd.
The FEXMountDaemon then uses the write end of the pipe to determine if
the read end of the pipe is still open. It does this using the epoll API
and ref counting how many pipes are still active.
This is possible since epoll will tell us if pipe status has changed to
error. Signalling to the write end that the read end has closed for
whatever reason.

Now we only use the "lock" file to remove races and tell the new
instances of FEX where the rootfs is mounted
2021-09-05 00:36:44 -07:00
CallumDev 8654d19f02 JIT: Fix unaligned CASPair on ARMv8.0 2021-09-05 16:54:53 +09:30
Ryan Houdek befd0aa9cf Merge pull request #1255 from Sonicadvance1/syscall_fixes
Syscall fixes
2021-09-04 18:39:33 -07:00
Ryan Houdek 96cef80a25 Linux: Adds missing Namespace handlers
Missed committing this at some point
2021-09-03 17:04:15 -07:00
Ryan Houdek fd3a88389b Linux: Fixes some 32-bit syscalls 2021-09-03 16:53:39 -07:00
Ryan Houdek d8350353d6 Linux: Creates a Types.h header that matches between architectures 2021-09-03 16:48:53 -07:00
Ryan Houdek 083d3a464a StructPackVerifier: Add a couple missing defines 2021-09-03 16:46:16 -07:00
Ryan Houdek 70931bf388 Merge pull request #1250 from Sonicadvance1/fix_signal_nodefer
Linux: Setup signal mask correctly to block signal-in-signal situations
2021-09-03 16:45:39 -07:00
Ryan Houdek 4f93259332 Linux: Setup signal mask correctly to block signal-in-signal situations
Due to how we emulate the guest signal handlers, we do the state setup
in the real host signal handler, then we jump out after state setup.
This was causing a situation where we were setting up the guest signal
handler state with the correct sa_mask.
Then after setting up the guest state we would leave the FEX signal
handler, restoring the signal mask to our original mask.

Instead now as we are setting up the guest state, we save our host
signal mask. Then on signal handler return we modify our host signal
mask to match what the guest wants.

Once we hit our sigreturn emulation we then restore the original signal
mask.

This looks to improve some stability problems regarding how wine uses
signals but it still doesn't fix the gvisor test sadly.
2021-09-02 20:21:10 -07:00
Ryan Houdek ad34cddbf2 Merge pull request #1249 from Sonicadvance1/fix_fexmountdaemon_messages
FEXMountDaemon: Fixes some minor issues
2021-09-02 16:32:41 -07:00
Ryan Houdek c74620f083 Merge pull request #1248 from Sonicadvance1/sigaltstack_ignore_onstack
Linux: sigaltstack ignore SS_ONSTACK
2021-09-02 16:32:24 -07:00
Ryan Houdek eefcde369a FEXMountDaemon: Fixes some minor issues
Now that we deparent FEXMountDaemon we can no longer use
PR_SET_PDEATHSIG.
Instead we rely on the ref counting and checking the pipe status to see
if the original parent has left us.
Further improvements that could be done in the future is that every user
of the mount point talks to the daemon to give it a pipe to check is
still live, since the ref counting sometimes is incorrect.
2021-09-02 15:42:55 -07:00
Ryan Houdek 005818177c unittests: Update posix tests known failures
This test is relying on legacy behaviour which is no longer true.
Keep running it but expect it to fail
2021-09-02 15:23:10 -07:00
Ryan Houdek 2ae47eae48 Linux: sigaltstack ignore SS_ONSTACK
This flag is ignored with sigaltstack.
Fixes an early assert in:
- Splice
- No Time to Explain Remastered
- Ittledew
- Hyperdrive Massacre
- English Country Tune
2021-09-02 15:15:20 -07:00
Ryan Houdek 1c4503e26a Merge pull request #1247 from Sonicadvance1/handle_xmm_state_32bit
Linux: Handle fpstate in the signal delegator correctly
2021-09-02 12:36:02 -07:00
Ryan Houdek 425ee98f81 Merge pull request #1246 from Sonicadvance1/itimer_32_fixes
Linux: Fixes 32-bit interval timers
2021-09-02 12:35:56 -07:00
Ryan Houdek 9945542375 SignalDelegator: Minor fix with SA_NODEFER
if a test application set NODEFER then changed the signal handler to
have one without it then we weren't correctly removing it
2021-09-02 03:31:58 -07:00
Ryan Houdek 0468bb4496 Linux: Handle fpstate in the signal delegator correctly
We were using the glibc context structure layout which doesn't match
what the kernel is doing.

Switches over to allocating fpstate independentally of the ucontext_t,
then pointing to it from uc_mcontext how we're supposed to.

glibc uses the __fpregs_mem region for other purposes.

This also fixes xmm state being stored on the 32-bit side, and also
fixes SIGALRM and SIGVTALRM siginfo overwriting structure data.
On 32-bit we can't just memcpy the siginfo_t over because the sizes
don't match. Throw a message instead.
2021-09-02 03:29:24 -07:00
Ryan Houdek e55e3a58b1 Linux: Fixes 32-bit interval timers
getitimer and setitimer use an itimerval struct with 32-bit members in
it.

Handles this case and now the timers work
2021-09-02 03:26:50 -07:00
Ryan Houdek 25e9585564 Merge pull request #1244 from CallumDev/fix_cas_armv8
Properly implement single CAS on ARMv8.0
2021-09-01 17:37:47 -07:00
CallumDev 9ab294b2f8 Interpreter: Templated AtomicCompareAndSwap 2021-09-01 21:02:07 +09:30
CallumDev d7f4fe7564 Interpreter: Fix x86 build 2021-09-01 19:24:58 +09:30
CallumDev e09219e5e5 Properly implement single CAS on ARMv8.0 2021-09-01 18:13:18 +09:30
Ryan Houdek c645d8683a Merge pull request #1243 from phire/24bit_assert
RA: Add max NoteCount assert
2021-08-31 06:43:13 -07:00
Scott Mansell eb9d3b11f2 RA: Add max NoteCount assert
The chance of hitting this is near zero, but still safer to have
an assert.
2021-09-01 01:32:41 +12:00
Ryan Houdek 7795078f7e Merge pull request #1242 from Sonicadvance1/deparent_fexmountdaemon
FEXMountDaemon: Early fork to deparent child
2021-08-30 23:07:19 -07:00
Ryan Houdek 0b564652d2 FEXMountDaemon: Early fork to deparent child
This allows FEXMountDaemon to remove FEXInterpreter as its parent.
Instead becoming the parent of whatever the current reaper process is.

Do it as early as possible this way FEXInterpreter won't get an
erroneous SIGCHLD.
2021-08-30 22:57:45 -07:00
Ryan Houdek 4f66d3e9dc Merge pull request #1240 from phire/extract_bucketlist
Move BucketList into it's own file
2021-08-30 19:21:12 -07:00
Scott Mansell bd7822bbe9 Move BucketList into it's own file 2021-08-31 14:06:07 +12:00
Ryan Houdek 0c484ac49c Merge pull request #1239 from Sonicadvance1/fix_ioctl_definition
Linux: Fix V3d and VC4 ioctl definitions
2021-08-30 19:02:46 -07:00
Ryan Houdek 5d21a1e6d6 Linux: Fix V3d and VC4 ioctl definitions
Oops, used the wrong definition

Also fix two struct definitions
2021-08-30 18:54:26 -07:00
Ryan Houdek d600b34b8e Merge pull request #1237 from Sonicadvance1/gvisor_fixes
Gvisor fixes
2021-08-30 18:08:47 -07:00
Ryan Houdek 4003ede7ee unittests: gvisor: Remove stale tmp files
Sometimes a test leaves a tmp file in the root of the rootfs.
This should be fixed but to ensure it stops happening, make sure to
delete it
2021-08-30 18:00:10 -07:00
Ryan Houdek 127d7c1a9e Merge pull request #1238 from Sonicadvance1/vc4_v3d_ioctl
Linux: x86: Initial V3D and VC4 ioctl emulation
2021-08-30 17:45:37 -07:00
Ryan Houdek f45de45568 unittests: Updates gvisor lists on changed behaviour 2021-08-30 17:43:12 -07:00
Ryan Houdek 5fdb66249b Linux: x86: Initial V3D and VC4 ioctl emulation
Untested but with how the structs are laid out, it likely just works
2021-08-30 17:35:58 -07:00
Ryan Houdek bb525e2291 unittests: Updates posix tests with expected failures
These are broken due to how sa_mask currently works. Will need to
resolve this.
2021-08-30 03:00:12 -07:00
Ryan Houdek de29c65585 Linux: Be more verbose about VFORK in code
With a comment claiming we don't support it yet since it causes problems
2021-08-30 02:27:54 -07:00
Ryan Houdek ef0f2246ac Linux: Fixes tkill syscall
There is no glibc wrapper for tkill and tgkill requires a tgid.
Kernel will reject us if we tried using -1 or 0 even though that is what
it does internally.
2021-08-30 02:26:46 -07:00
Ryan Houdek 76877e08bc Linux: Add some comments claiming shmctl and msqctl is incorrect on x86
These aren't correctly handled, so they will get garbage data for now
2021-08-30 02:25:48 -07:00
Ryan Houdek d2636f63f2 Linux: Return -EPERM on non-canonical TLS
If it lives outside of the canonical range then immediately reject with
-EPERM. Just like the official kernel.
2021-08-30 02:24:51 -07:00
Ryan Houdek c241c2f5e7 Linux: Report that we don't support RSEQ
This way the gvisor tests skip rather than just breaking
2021-08-30 02:24:19 -07:00
Ryan Houdek 99e24c284c Linux: Let an application correctly reset its signal handlers
If it is going back to SIG_IGN or SIG_DFL then let them unregister.
This is useful for when they are wanting to ignore a signal after a
while.
Or only capture a fault during a time, then afterwards want the
application to crash on error.

Additionally clears up some other minor logic which wasn't being used
anymore
2021-08-30 02:22:52 -07:00
Ryan Houdek bb3bd3faa1 Linux: Register guest signal delegators for all signals including SIGRT32
Had accidentally missed the last one.
2021-08-30 02:21:41 -07:00
Ryan Houdek 6b3a5470c4 Linux: Fixes tracking of /proc/self with dirfd
Previously we would miss this. Causing things like lscpu to leak state.
2021-08-30 02:20:39 -07:00
Ryan Houdek d241c925c0 Linux: Fixes /proc/self/cmdline arguments
With some changes in the frontend this had gotten out of sync.
2021-08-30 02:19:50 -07:00
Ryan Houdek 378dfcf164 Linux: Fixes semctl syscall on AArch64
Turns out the semid_ds struct is a different layout on x86-64 versus
AArch64.
This was causing some minor failures
2021-08-30 02:19:07 -07:00
Ryan Houdek 9fd558e173 Linux: Fixes signalfd
Was using the wrong signal mask for this.
We need to use the one passed in to the syscall, not the current active
mask
2021-08-30 02:16:26 -07:00
Ryan Houdek 3fbc3c347a Linux: Fixes the pread/pwrite family of syscalls
These syscall arguments are not laid out in a sane way.
Looks like they wanted to keep the interface the same for both x86-64
and 32-bit x86 so they split the offset argument in to two values.

Weirdly enough, even though these are 32-bit offsets on 32-bit x86; On
x86-64 these are still 64-bit. Which means the kernel weirdly allows you
to overlap the two 64-bit halves.

eg: `uint64_t Offset = (pos_high << 32) | pos_low;`
So you can have a 64-bit value that is the full range, but if you have
data in the lower 32-bits of pos_high then you corrupt the offset.
Additionally this allows you to interleave low and high if you want to
be obtuse.
2021-08-30 02:12:35 -07:00
Ryan Houdek fb0b03808f Linux: Fixes waitid syscall
The raw syscall has an rusage argument that can give you child process
rusage similar to wait4
2021-08-30 02:11:07 -07:00
Ryan Houdek a1aeb0d7ee Linux: Fixes execve with missing arguments
Most of the arguments to execve can be nullptr.
Fixes this from crashing
2021-08-30 02:09:51 -07:00
Ryan Houdek 4b6b7495df Signals: Fix incorrect SIGINFO check
Initially this was a small hack to work around sigqueueinfo sending over
siginfo_t. Now this is unnecessary and would crash if an application
used sigqueueinfo to a signal without the SIGINFO flag.
2021-08-30 02:06:55 -07:00
Ryan Houdek 29e30773ee Signals: On jumping to signal frame FPU state is reset
All registers are set to zero and setup to be able to do float
operations.
2021-08-30 02:05:56 -07:00
Ryan Houdek 3e05544e78 Merge pull request #1236 from Sonicadvance1/fexconfig_advanced
FEXConfig: Load application config and advanced tab
2021-08-29 13:05:19 -07:00
Ryan Houdek 61260c2a62 FEXConfig: Load application config and advanced tab
Loading a preexisting application config was impossible through the gui.
Adds a way to do so.

Also adds an advanced tab that just displays all the items.
Allows pruning of application config options to be fairly efficient
2021-08-29 12:30:39 -07:00
Ryan Houdek 109c42a629 Merge pull request #1235 from Sonicadvance1/fix_emulated_openat
EmulatedFiles: Fixes openat for emulated files not using FDCWD
2021-08-28 18:09:42 -07:00
Ryan Houdek 9140ba28b7 Merge pull request #1234 from Sonicadvance1/FEXBash_init
FEXBash: Allow creating a bash instance easily
2021-08-28 17:55:33 -07:00
Ryan Houdek 1c3be542f7 Merge pull request #1233 from Sonicadvance1/fix_execve_escape
Linux: Fixes accidental execve escape
2021-08-28 17:55:22 -07:00
Ryan Houdek 8a331202ee EmulatedFiles: Fixes openat for emulated files not using FDCWD
Fixes lscpu and other applications that open the directory first and
then files inside of that directory.
2021-08-28 17:54:05 -07:00
Ryan Houdek f72ecacd86 FEXBash: Allow creating a bash instance easily
If not passing in any arguments then just immediately create a bash
instance. Incredibly nice little helper
2021-08-28 16:54:25 -07:00
Ryan Houdek 057c1de69d Linux: Fixes accidental execve escape
If the binfmt_misc interpreter was installed then we were running execve
directly.
This allowed shebang programs to escape and see the host architecture
when we weren't planning on it.

Fixes `FEXBash steam` from complaining about missing packages.
2021-08-28 16:37:29 -07:00
Ryan Houdek 10793e89a9 Merge pull request #1232 from Sonicadvance1/iwyu_fixes
Massive amount of IWYU cleanup
2021-08-28 09:54:45 -07:00
Ryan Houdek f161e3bfb0 Massive amount of IWYU cleanup
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
2021-08-28 00:32:15 -07:00
Ryan Houdek c206942b59 Merge pull request #1231 from Sonicadvance1/Arm64Emitter_iwyu
Arm64Emitter: Resolves some IWYU warnings
2021-08-27 17:18:10 -07:00
Ryan Houdek 10922293c9 Merge pull request #1230 from Sonicadvance1/fix_stdio_log
FEXLoader: Fixes potential bug in log output to stdout/stderr
2021-08-27 17:09:04 -07:00
Ryan Houdek 38ce5876f8 Arm64Emitter: Resolves some IWYU warnings
Should help some build errors
2021-08-27 17:08:29 -07:00
Ryan Houdek de64db5852 FEXLoader: Fixes potential bug in log output to stdout/stderr
Comment in the file for why this can be a bug.

Noticed a game opening a file and our logs were ending up in their
files.
2021-08-27 15:23:17 -07:00
Ryan Houdek 097b3ad881 Merge pull request #1229 from Sonicadvance1/more_signal_splitting
SignalDelegator: More splitting and cleanup
2021-08-27 13:57:07 -07:00
Ryan Houdek c25ae9b2c3 Dispatcher: Changes SRA assert in to an error message instead
In some instances this is safe but we can't currently distinguish safe
or not.
2021-08-27 12:58:40 -07:00
Ryan Houdek b625437071 SignalDelegator: More splitting and cleanup
This is working towards getting the stack frame for guest signals being
pushed over to the Frontend.
Still some more work to do but this is the first step that can be split
up.
2021-08-27 12:46:41 -07:00
Ryan Houdek b0c8710b2c X86Enums: Adds some more defines around signals 2021-08-27 12:42:26 -07:00
Ryan Houdek dfce1dc476 SignalDispatcher: Minor fix for SRA
In the case of SRA plus a signal not using siginfo then we were not
spilling SRA registers.
In this case we would then shift the signal frame and fill the guest
context with garbage register data.
This would potentially cause some issues but haven't noticed anything
outside outside of my test applications.
2021-08-27 12:38:19 -07:00
Ryan Houdek 77db25fc24 Merge pull request #1227 from Sonicadvance1/initial_hangover
Hangover: Initial support for the syscall handling.
2021-08-26 01:49:59 -07:00
Ryan Houdek fa224a3557 Merge pull request #1226 from Sonicadvance1/fix_callback
Fixes Callback interface to take a thread argument
2021-08-26 01:41:16 -07:00
Ryan Houdek b373d0fcfe Merge pull request #1225 from Sonicadvance1/fix_libs
Fixes jemalloc library ordering
2021-08-26 01:40:08 -07:00
Ryan Houdek a3490aad5b Hangover: Initial support for the syscall handling.
Hangover currently abuses the syscall op for thunking purposes.
Very likely this will be changed in the future but for now make sure to
support that use case.
2021-08-26 01:34:13 -07:00
Ryan Houdek 304db72d6b Fixes Callback interface to take a thread argument
This was using the implicit thread TLS object. All users of this
have access to the thread object directly.

Use that instead. Fixes a subtle bug were the frontend could be trying
to do a callback and TLS sections weren't correctly set.
2021-08-26 01:23:17 -07:00
Ryan Houdek b15e0c5f6c Fixes jemalloc library ordering
FEXCore relies on jemalloc symbols if compiled with it.
Have FEXCore link to jemalloc instead of the frontend.

Fixes a missing symbol if someone loads libFEXCore

Additionally, stop trying to compile JEMalloc if not enabled
2021-08-26 01:21:56 -07:00
Ryan Houdek 97f413cfec Merge pull request #1224 from Sonicadvance1/expose_parent_thread
FEXCore: Return the ParentThread with InitCore
2021-08-25 22:47:37 -07:00
Ryan Houdek 6da3330646 Merge pull request #1223 from Sonicadvance1/move_config_to_fexcore
Config: Moves non-OS specific configuration loading to FEXCore
2021-08-25 22:47:30 -07:00
Ryan Houdek 5debdf8d57 Merge pull request #1222 from Sonicadvance1/minor_visibility_fixes
FEXCore: Minor symbol visibility fixes
2021-08-25 22:47:23 -07:00
Ryan Houdek 6483740553 FEXCore: Return the ParentThread with InitCore
This always returns success and for easier state management, just return
our parent thread.

Frontend needs full visibility of this state anyway for thread
management.
2021-08-24 23:20:44 -07:00
Ryan Houdek f131f07612 Config: Moves non-OS specific configuration loading to FEXCore
Puts the visibility of the main layer, application layers, and
environment in to FEXCore instead of FEX.
These layers aren't specific to FEX/FEXLoader and should live in
FEXCore.

Only the EmptyMapper remains in FEX, which should eventually move over
to FEXConfig since that is the only user.
2021-08-24 23:17:32 -07:00
Ryan Houdek f268a28caf FEXCore: Minor symbol visibility fixes
Noticed these missing
2021-08-24 22:35:38 -07:00
Ryan Houdek 6afc3ca13a Merge pull request #1219 from Sonicadvance1/fix_large_offset_syscalls
Linux: Fixes 32-bit syscalls that use 64-bit values
2021-08-23 22:42:08 -07:00
Ryan Houdek 35c664295d Merge pull request #1215 from Sonicadvance1/fix_arm64_signal_handling
Arm64: Fixes SRA spilling on signal
2021-08-23 00:47:10 -07:00
Ryan Houdek b7af5c641c Linux: Fixes 32-bit syscalls that use 64-bit values
A bunch of these were defined incorrectly. I tested a few of these
locally to ensure they were correct after the fact.

Shows that most 32-bit applications that we've encountered aren't
dealing with files larger than 4GB.
2021-08-23 00:44:04 -07:00
Ryan Houdek 6b87839437 Merge pull request #1216 from Sonicadvance1/fix_32bit_sigsegv
x86: Fixes siginfo_t si_addr for SIGBUS/SIGSEGV
2021-08-22 22:26:27 -07:00
Ryan Houdek 9563b5aa3a Merge pull request #1218 from Sonicadvance1/hotfix_nasm_fix
unittests: Hotfix for older nasm
2021-08-22 22:26:07 -07:00
Ryan Houdek 28b3bc3508 unittests: Hotfix for older nasm
Newer nasm takes the size specifier for LEA, older ones do not
2021-08-22 21:59:57 -07:00
Ryan Houdek 8099dfc830 Merge pull request #1210 from Sonicadvance1/proton_6.3_fixes
Proton 6.3 32-bit fixes
2021-08-22 21:32:26 -07:00
Ryan Houdek 926ddabbd1 Merge pull request #1211 from Sonicadvance1/implement_repne_strings
OpcodeDispatcher: Implements undocumented repne on string ops
2021-08-22 21:32:17 -07:00
Ryan Houdek af3af9d048 x86: Fixes siginfo_t si_addr for SIGBUS/SIGSEGV
si_addr is set to the address that is trying to be accessed, not the RIP
that is trying to access it.
We just need to copy our host value over for this.
SIGFPE and SIGILL we still don't have a good answer for.
2021-08-22 21:10:51 -07:00
Ryan Houdek 3b8f24d74f Arm64: Fixes SRA spilling on signal
We were checking the PC after we set the new location in the data
structure.
Didn't matter for x86-64 since it doesn't use SRA but it does matter for
ARM.

Now on signal while in JIT code it will spill correctly
2021-08-22 19:14:04 -07:00
Ryan Houdek 8931ddc382 unittests: Duplicates rep string unit tests with repne
These are exactly the same except using the other prefix.
Hardware tests confirm that this behave the same
2021-08-21 20:19:38 -07:00
Ryan Houdek dd8225be8a OpcodeDispatcher: Implements undocumented repne on string ops
MOVS, LODS, and STOS ops are only documented to support the REP prefix.
These instructions actually repeat correctly with the REPNE prefix as
well.
Behaviour is the same.

Fixes POD Gold
2021-08-21 20:16:25 -07:00
Ryan Houdek 02c10b9671 unittests: Adds tests for FXSave/FXRStor 2021-08-21 17:45:52 -07:00
Ryan Houdek 4d7455989c OpcodeDispatcher: Ensure FXSave/FXRStor doesn't store too many XMM registers in 32-bit
32-bit doesn't store XMM registers 8-15 since they don't exist there.
This technically falls under the "reserved" slot so applications can't
rely on them to not be written.
Saves us a bit of CPU overhead at the very least
2021-08-21 17:44:06 -07:00
Ryan Houdek 23fb4baf46 unittests: Adds unit tests for storing segment register sizes 2021-08-21 17:43:39 -07:00
Ryan Houdek 6e5fc5cdde Fixes segment register storing to memory
These were storing to memory as 32-bits but when storing to memory these
are only ever stored as 16-bit
2021-08-21 17:41:57 -07:00
Ryan Houdek 6a08587d2e Merge pull request #1209 from Sonicadvance1/wine_fixes2
32-bit wine fixes
2021-08-21 17:41:16 -07:00
Ryan Houdek 7bfa1c4838 Linux: Fixes 32-bit sigaltstack
Due to an incorrect pointer check, we were never setting the sigaltstack
on 32-bit applications.
Additionally the stack_t type didn't have the members in the correct
order.

This fixes wine 32-bit applications where wine sends an application
SIGUSR1 and expects the altstack to be used. This is because the
altstack has the thread's TEB region at the start of the stack.

Without this when the application was getting sent a SIGUSR1, it would
remain in the application stack, thus getting an invalid TEB and loading
an FS register with zero.
2021-08-20 23:55:56 -07:00
Ryan Houdek 84f42f6155 Dispatcher: Fixes alt stack check and redzone offset
There are two checks to the alt stack that the signal handler needs to
check.
First it needs to check if the signal handler was registered with teh
flag SA_ONSTACK.

Then it needs to check if the alt stack is actually enabled by not
having flag SS_DISABLE.

Additionally, 32-bit x86 doesn't have a redzone so stop offsetting by
128
2021-08-20 23:54:09 -07:00
Ryan Houdek e4a230c3ec OpcodeDispatcher: Fixes FTW saving/storing in FXSave/FXRStor
Missed this when implementing FTW
2021-08-20 23:53:13 -07:00
Ryan Houdek 09296fe73e unittests: Adds IRET unit tests 2021-08-20 23:52:42 -07:00
Ryan Houdek d2130e1df3 OpcodeDispatcher: Fixes 32-bit IRET
We were failing to store the updated ESP on 32-bit.
On 64-bit or CPL change (which we don't support) the stack pointer is
pulled from the frame.

This resolves an issue in wine's 32-bit loader where it was getting zero
for one of the context pointers since the stack pointer wasn't adjusted
2021-08-20 23:50:04 -07:00
Ryan Houdek 01c49dbb9c FEXLoader: Updates RanAsInterpreter check for FD exec
Only the interpreter can run when executed as FD.
This happens when executed with binfmt_misc and can resolve an issue if
someone sets up the hardlink incorrectly.
2021-08-20 23:48:42 -07:00
Ryan Houdek 6d60689ad4 FileManagement: Stop calling getpid for every file access
We can cache the pid result, we know every case in which the pid is
changed.
This makes watching strace significantly less annoying.
2021-08-20 23:47:48 -07:00
Ryan Houdek 63af80fce3 Merge pull request #1205 from Sonicadvance1/fix_jemalloc_missing_alias
Updates jemalloc to fix missing alias posix_memalign
2021-08-15 03:19:31 -07:00
Ryan Houdek 43052a5707 Merge pull request #1204 from Sonicadvance1/pressure_vessel_option_program
Adds new FEXGetConfig program
2021-08-15 03:19:24 -07:00
Ryan Houdek db5a26991f Updates jemalloc to fix missing alias posix_memalign
This was causing us to fail compiling in debug.
xbyak uses this.
2021-08-14 02:59:40 -07:00
Ryan Houdek c07b5e480b Adds new FEXGetConfig program
This is a simple program to get a few configuration options that are necessary
to expose for pressure-vessel
2021-08-13 20:53:30 -07:00
Ryan Houdek 7aae9b7e44 Merge pull request #1201 from Sonicadvance1/pressure-vessel_fixes
Linux: Implements support for clone with namespaces
2021-08-12 12:57:48 -07:00
Ryan Houdek e90892a20b unittests: Disable gvisor test that doesn't pass when namespaces are disabled
It's expecting EPERM (Which is what we used to return) but it is also
valid for the kernel to return EINVAL when not compiled with namespaces enabled
2021-08-10 19:13:32 -07:00
Ryan Houdek a46773a9ff unittests: Disables gvisor test that uses unsupported clone flags
It uses CLONE_VM which breaks our cloning
2021-08-10 18:39:40 -07:00
Ryan Houdek 740270c05f Linux: Implements support for clone with namespaces
This is very tricky to handle and it has a bunch of rough edges.
One of the major problems that we can't workaround is that if we receive a
clone flag that pthreads can't support with THREAD, then we are required to fall down
the pthreads code path.
This is because threads going down the clone path will break TLS and we don't have
a way to work around it currently.

So this adds a clone path, a clone3 path, and keeps the legacy path as well.
Which makes this fairly convoluted but it gets pressure-vessel working on x86-64 host.
It's a bit tricky to setup but it does work.

Still some work necessary to get pressure-vessel working on AArch64 host, but I'm working on that.
2021-08-10 18:31:04 -07:00
Ryan Houdek 83d20d8f34 Allocator: Leak the allocator object until we fix static initializer allocations
When we are running a 32-bit process we end up mixing VMA Region allocators which
can cause crashes on shutdown.
This is because if a statically initialized object allocates memory, and then frees that
memory in the atexit handler. There is a chance that if it had to reallocate memory during the
VMA allocator switch, that the atexit handler will try freeing memory from the FEX VMA region allocator
AFTER it has already been deallocated.

This can't be safely worked around with atexit handlers.
So until we resolve this issue, we HAVE to leak the allocator so it can safely clean up and then let the kernel
clean up after us
2021-08-10 18:23:39 -07:00
Ryan Houdek e388729403 Config: Changes string config default values to be string_view
These were causing static initializer construction all over the codebase.
Scope of these variables has also been changed over to the module instead of the header as well
2021-08-10 18:20:15 -07:00
Ryan Houdek e3fda9f232 Telemetry: Changes telemetry names map to use string_view
These don't need to be std::string and was causing allocations to occur in the static
initializer
2021-08-10 18:19:03 -07:00
Ryan Houdek f5940df822 RAPass: Removes static initialization of INVALID_REGCLASS
We were just using this as a reference and it was causing a static initializer
2021-08-10 18:18:07 -07:00
Ryan Houdek e6c4f9aad1 IRParser: Removes static initialization of map
This didn't need to be global
2021-08-10 18:16:49 -07:00
Ryan Houdek 8eb0df96f9 Merge pull request #1199 from MerryMage/GetCursorAddress
FEXCore: Use GetCursorAddress when able
2021-08-08 13:26:09 -07:00
Merry 55d981fcb0 FEXCore: Use GetCursorAddress when able 2021-08-08 21:06:06 +01:00
Ryan Houdek b120a8ea84 Merge pull request #1196 from Sonicadvance1/offline_telemetry
Implements support for offline *only* telemetry
2021-08-06 23:23:38 -07:00
Ryan Houdek 5d73ac3234 Merge pull request #1198 from Sonicadvance1/fix_binfmt_misc_arch
Arm64: Reimplements support for binfmt_misc without update-binfmts
2021-08-06 23:23:25 -07:00
Ryan Houdek b05adaeba3 Arm64: Reimplements support for binfmt_misc without update-binfmts
Arch doesn't have update-binfmts. Fall back to the classic approach
without it.
2021-08-06 23:14:25 -07:00
Ryan Houdek 0be16baebf Merge pull request #1197 from Sonicadvance1/rebase_skmp/no-sra
Rebase skmp/no sra
2021-08-06 23:09:13 -07:00
Ryan Houdek 8dd41e7fcd Merge pull request #1193 from Sonicadvance1/fix_readlinkat_self
Linux: Fixes readlinkat for self
2021-08-06 22:59:47 -07:00
Ryan Houdek 83cdf0f377 Fix no-sra crash 2021-08-06 22:57:33 -07:00
Stefanos Kornilios Mitsis Poiitidis b010ab42c7 JIT: Add an option to disable SRA 2021-08-06 22:52:58 -07:00
Ryan Houdek 1c1f40e5af Linux: Fixes readlinkat for self
Similar code to readlinkat. Needs to be correct for self otherwise we
return EINVAL which is unexpected since self should always be a symlink

Fixes a bug in running bwrap
2021-08-06 22:51:31 -07:00
Ryan Houdek c6c94570b4 Implements support for offline *only* telemetry
This information is only ever going to be offline. Will be useful for multiple reasons.

1) Searching for split lock usage in applications, which can be a programming bug.
  a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
2021-08-06 22:40:19 -07:00
Ryan Houdek cce3f365cc Merge pull request #1194 from Sonicadvance1/implement_pivot_root
Linux: Implements pivot_root syscall
2021-08-06 22:36:30 -07:00
Ryan Houdek 4641e44276 Linux: Implements pivot_root syscall
Somehow missed this one. Easy enough and matches between architectures.
Used by bubblewrap
2021-08-03 23:30:38 -07:00
317 changed files with 22678 additions and 2612 deletions

No files matched your search

+4
View File
@@ -42,3 +42,7 @@
[submodule "External/xxhash"]
path = External/xxhash
url = https://github.com/FEX-Emu/xxHash.git
[submodule "External/Vulkan-Docs"]
shallow = true
path = External/Vulkan-Docs
url = https://github.com/KhronosGroup/Vulkan-Docs.git
+24 -5
View File
@@ -16,6 +16,7 @@ option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_STATIC_PIE "Enables static-pie build" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -82,6 +83,11 @@ if (ENABLE_LLD)
link_libraries(${LD_OVERRIDE})
endif()
if (NOT ENABLE_OFFLINE_TELEMETRY)
# Disable FEX offline telemetry entirely if asked
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if (ENABLE_STATIC_PIE)
if (_M_ARM_64 AND ENABLE_LLD)
message (FATAL_ERROR "Static linking does not currently work with AArch64+LLD. Use GNU ld for now.")
@@ -219,6 +225,8 @@ endif()
if (ENABLE_JEMALLOC)
add_definitions(-DENABLE_JEMALLOC=1)
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
else()
message (STATUS
" jemalloc disabled!\n"
@@ -254,9 +262,6 @@ endif()
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
add_subdirectory(External/cpp-optparse/)
include_directories(External/cpp-optparse/)
@@ -411,6 +416,11 @@ add_subdirectory(Data/binfmts/)
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
# Install the ThunksDB file
install(
FILES ${CMAKE_CURRENT_SOURCE_DIR}/Data/ThunksDB.json
DESTINATION ${DATA_DIRECTORY}/)
if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
@@ -422,7 +432,10 @@ if (BUILD_THUNKS)
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
CMAKE_ARGS "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
@@ -440,7 +453,13 @@ if (BUILD_THUNKS)
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}" "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DX86_C_COMPILER:STRING=${X86_C_COMPILER}"
"-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
+287
View File
@@ -0,0 +1,287 @@
{
"DB": {
"GL": {
"Library" : "libGL-guest.so",
"Depends": [
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGL.so",
"/usr/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/lib/x86_64-linux-gnu/libGL.so",
"/lib/x86_64-linux-gnu/libGL.so.1",
"/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/lib/x86_64-linux-gnu/libGL.so.1.7.0"
]
},
"GLESv2": {
"Library": "libGLESv2-guest.so",
"Depends": [
"X11"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/lib/x86_64-linux-gnu/libGLESv2.so",
"/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libX11.so",
"/usr/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/lib/x86_64-linux-gnu/libX11.so",
"/lib/x86_64-linux-gnu/libX11.so.6",
"/lib/x86_64-linux-gnu/libX11.so.6.4.0"
]
},
"Vulkan-radeon": {
"Library": "libvulkan_radeon-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/lib/x86_64-linux-gnu/libvulkan_radeon.so"
],
"Comment": [
"Vulkan library relies on xcb, otherwise it crashes with jemalloc"
]
},
"Vulkan-lavapipe": {
"Library": "libvulkan_lvp-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/lib/x86_64-linux-gnu/libvulkan_lvp.so"
]
},
"Vulkan-freedreno": {
"Library": "libvulkan_freedreno-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/lib/x86_64-linux-gnu/libvulkan_freedreno.so"
]
},
"Vulkan-intel": {
"Library": "libvulkan_intel-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/lib/x86_64-linux-gnu/libvulkan_intel.so"
]
},
"Vulkan-panfrost": {
"Library": "libvulkan_panfrost-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/lib/x86_64-linux-gnu/libvulkan_panfrost.so"
]
},
"Vulkan-nvidia": {
"Library": "libvulkan_nvidia-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/lib/x86_64-linux-gnu/libGLX_nvidia.so.0"
],
"Comment": [
"Not currently wired up"
]
},
"Vulkan-virtio": {
"Library": "libvulkan_virtio-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/lib/x86_64-linux-gnu/libvulkan_virtio.so"
]
},
"xcb": {
"Library": "libxcb-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb.so",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/lib/x86_64-linux-gnu/libxcb.so",
"/lib/x86_64-linux-gnu/libxcb.so.1",
"/lib/x86_64-linux-gnu/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb_dri2-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb_dri3-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb_xfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb_shm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb_sync-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/lib/x86_64-linux-gnu/libxcb-sync.so",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb_randr-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb_present-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-present.so",
"/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb_glx-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libshmfence-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/lib/x86_64-linux-gnu/libxshmfence.so",
"/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libdrm.so",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/lib/x86_64-linux-gnu/libdrm.so",
"/lib/x86_64-linux-gnu/libdrm.so.2",
"/lib/x86_64-linux-gnu/libdrm.so.2.4.0"
]
},
"asound": {
"Library": "libasound-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libasound.so",
"/usr/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/lib/x86_64-linux-gnu/libasound.so",
"/lib/x86_64-linux-gnu/libasound.so.2",
"/lib/x86_64-linux-gnu/libasound.so.2.0.0"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXrender.so",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/lib/x86_64-linux-gnu/libXrender.so",
"/lib/x86_64-linux-gnu/libXrender.so.1",
"/lib/x86_64-linux-gnu/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXext.so",
"/usr/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/lib/x86_64-linux-gnu/libXext.so",
"/lib/x86_64-linux-gnu/libXext.so.6",
"/lib/x86_64-linux-gnu/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/lib/x86_64-linux-gnu/libXfixes.so",
"/lib/x86_64-linux-gnu/libXfixes.so.3",
"/lib/x86_64-linux-gnu/libXfixes.so.3.1.0"
]
},
"":{}
}
}
+16 -3
View File
@@ -86,6 +86,7 @@ set (SRCS
Interface/Core/OpcodeDispatcher/Vector.cpp
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/SignalDelegator.cpp
Interface/Core/X86Tables.cpp
Interface/Core/X86DebugInfo.cpp
Interface/Core/X86HelperGen.cpp
@@ -118,6 +119,7 @@ set (SRCS
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/PhiValidation.cpp
@@ -129,6 +131,7 @@ set (SRCS
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/LogManager.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
)
@@ -183,6 +186,16 @@ if (ENABLE_JITSYMBOLS)
list(APPEND DEFINES -DENABLE_JITSYMBOLS=1)
endif()
set (LIBS vixl dl fmt::fmt xxhash tiny-json)
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
endif()
# Generate config
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
${CMAKE_BINARY_DIR}/generated/Config/Config.json)
# Generate IR include file
set(OUTPUT_IR_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/IR")
set(OUTPUT_NAME "${OUTPUT_IR_FOLDER}/IRDefines.inc")
@@ -225,7 +238,7 @@ add_custom_target(IR_INC
set(OUTPUT_CONFIG_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/Config")
set(OUTPUT_CONFIG_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigValues.inl")
set(OUTPUT_CONFIG_OPTION_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigOptions.inl")
set(INPUT_CONFIG_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json")
set(INPUT_CONFIG_NAME "${CMAKE_BINARY_DIR}/generated/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
add_custom_target(CREATE_CONFIG_FOLDER ALL
@@ -269,7 +282,7 @@ function(AddObject Name Type)
add_dependencies(${Name} IR_INC)
add_dependencies(${Name} CONFIG_INC)
target_link_libraries(${Name} vixl dl fmt::fmt xxhash)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
@@ -310,7 +323,7 @@ endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} vixl dl fmt::fmt xxhash)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
+1 -5
View File
@@ -1,11 +1,7 @@
#include "NetStream.h"
#include <cstring>
#include <sys/types.h>
#include <sys/socket.h>
#include <stdio.h>
#include <unistd.h>
int NetStream::NetBuf::flushBuffer(const char *buffer, size_t size) {
@@ -29,7 +25,7 @@ std::streamsize NetStream::NetBuf::xsputn(const char* buffer, std::streamsize si
// Check if the string fits neatly in our buffer
if (size <= buf_remaining) {
std::memcpy(pptr(), buffer, size);
::memcpy(pptr(), buffer, size);
pbump(size);
return size;
}
+1
View File
@@ -2,6 +2,7 @@
#include <array>
#include <iostream>
#include <iterator>
#include <string.h>
class NetStream : public std::iostream {
+33 -1
View File
@@ -3,12 +3,44 @@
#include <cstdlib>
#include <filesystem>
#include <sys/stat.h>
#include <memory>
#include <pwd.h>
#include <system_error>
#include <unistd.h>
namespace FEXCore::Paths {
std::unique_ptr<std::string> CachePath;
std::unique_ptr<std::string> EntryCache;
char const* FindUserHomeThroughUID() {
auto passwd = getpwuid(geteuid());
if (passwd) {
return passwd->pw_dir;
}
return nullptr;
}
const char *GetHomeDirectory() {
char const *HomeDir = getenv("HOME");
// Try to get home directory from uid
if (!HomeDir) {
HomeDir = FindUserHomeThroughUID();
}
// try the PWD
if (!HomeDir) {
HomeDir = getenv("PWD");
}
// Still doesn't exit? You get local
if (!HomeDir) {
HomeDir = ".";
}
return HomeDir;
}
void InitializePaths() {
CachePath = std::make_unique<std::string>();
EntryCache = std::make_unique<std::string>();
+3
View File
@@ -4,6 +4,9 @@
namespace FEXCore::Paths {
void InitializePaths();
void ShutdownPaths();
const char *GetHomeDirectory();
std::string GetCachePath();
std::string GetEntryCachePath();
}
+26
View File
@@ -16,9 +16,19 @@ extern "C" {
struct X80SoftFloat {
#ifdef _M_X86_64
// Define this to push some operations to x87
// Only useful to see if precision loss is killing something
// #define DEBUG_X86_FLOAT
#ifdef DEBUG_X86_FLOAT
#define BIGFLOAT long double
#define BIGFLOATSIZE 10
#else
#define BIGFLOAT __float128
#define BIGFLOATSIZE 16
#endif
#elif defined(_M_ARM_64)
#define BIGFLOAT long double
#define BIGFLOATSIZE 16
#else
#error No 128bit float for this target!
#endif
@@ -170,8 +180,14 @@ struct X80SoftFloat {
}
operator BIGFLOAT() const {
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(*this);
return FEXCore::BitCast<BIGFLOAT>(Result);
#else
BIGFLOAT result{};
memcpy(&result, this, sizeof(result));
return result;
#endif
}
operator int16_t() const {
@@ -217,6 +233,12 @@ struct X80SoftFloat {
*this = ui64_to_extF80(rhs);
}
#if BIGFLOATSIZE == 10
void operator=(const long double rhs) {
memcpy(this, &rhs, sizeof(rhs));
}
#endif
operator void*() {
return reinterpret_cast<void*>(this);
}
@@ -236,7 +258,11 @@ struct X80SoftFloat {
}
X80SoftFloat(BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(FEXCore::BitCast<float128_t>(rhs));
#else
*this = FEXCore::BitCast<long double>(rhs);
#endif
}
X80SoftFloat(const int16_t rhs) {
+365 -47
View File
@@ -1,43 +1,164 @@
#include "Common/StringConv.h"
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Context/Context.h"
#include "Common/Paths.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <assert.h>
#include <cstdlib>
#include <filesystem>
#include <pwd.h>
#include <fstream>
#include <functional>
#include <map>
#include <memory>
#include <list>
#include <optional>
#include <stddef.h>
#include <stdint.h>
#include <string>
#include <string_view>
#include <sys/sysinfo.h>
#include <unistd.h>
#include <system_error>
#include <type_traits>
#include <unordered_map>
#include <utility>
#include <vector>
#include <tiny-json.h>
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Config {
char const* FindUserHomeThroughUID() {
auto passwd = getpwuid(geteuid());
if (passwd) {
return passwd->pw_dir;
}
return nullptr;
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#include <FEXCore/Config/ConfigValues.inl>
}
static bool LoadConfigFile(std::vector<char> &Data, const std::string &Config) {
std::fstream ConfigFile;
ConfigFile.open(Config, std::ios::in);
if (!ConfigFile.is_open()) {
return false;
}
const char *GetHomeDirectory() {
char const *HomeDir = getenv("HOME");
if (!ConfigFile.seekg(0, std::fstream::end)) {
LogMan::Msg::D("Couldn't load configuration file: Seek end");
return false;
}
// Try to get home directory from uid
if (!HomeDir) {
HomeDir = FindUserHomeThroughUID();
auto FileSize = ConfigFile.tellg();
if (ConfigFile.fail()) {
LogMan::Msg::D("Couldn't load configuration file: tellg");
return false;
}
if (!ConfigFile.seekg(0, std::fstream::beg)) {
LogMan::Msg::D("Couldn't load configuration file: Seek beginning");
return false;
}
if (FileSize > 0) {
Data.resize(FileSize);
if (!ConfigFile.read(&Data.at(0), FileSize)) {
// Probably means permissions aren't set. Just early exit
return false;
}
ConfigFile.close();
}
else {
return false;
}
return true;
}
namespace JSON {
struct JsonAllocator {
jsonPool_t PoolObject;
std::unique_ptr<std::list<json_t>> json_objects;
};
static_assert(offsetof(JsonAllocator, PoolObject) == 0, "This needs to be at offset zero");
json_t* PoolInit(jsonPool_t* Pool) {
JsonAllocator* alloc = reinterpret_cast<JsonAllocator*>(Pool);
alloc->json_objects = std::make_unique<std::list<json_t>>();
return &*alloc->json_objects->emplace(alloc->json_objects->end());
}
json_t* PoolAlloc(jsonPool_t* Pool) {
JsonAllocator* alloc = reinterpret_cast<JsonAllocator*>(Pool);
return &*alloc->json_objects->emplace(alloc->json_objects->end());
}
static void LoadJSonConfig(const std::string &Config, std::function<void(const char *Name, const char *ConfigSring)> Func) {
std::vector<char> Data;
if (!LoadConfigFile(Data, Config)) {
return;
}
// try the PWD
if (!HomeDir) {
HomeDir = getenv("PWD");
JsonAllocator Pool {
.PoolObject = {
.init = PoolInit,
.alloc = PoolAlloc,
},
};
json_t const *json = json_createWithPool(&Data.at(0), &Pool.PoolObject);
if (!json) {
LogMan::Msg::E("Couldn't create json");
return;
}
// Still doesn't exit? You get local
if (!HomeDir) {
HomeDir = ".";
json_t const* ConfigList = json_getProperty(json, "Config");
if (!ConfigList) {
LogMan::Msg::E("Couldn't get config list");
return;
}
return HomeDir;
for (json_t const* ConfigItem = json_getChild(ConfigList);
ConfigItem != nullptr;
ConfigItem = json_getSibling(ConfigItem)) {
const char* ConfigName = json_getName(ConfigItem);
const char* ConfigString = json_getValue(ConfigItem);
if (!ConfigName) {
LogMan::Msg::E("Couldn't get config name");
return;
}
if (!ConfigString) {
LogMan::Msg::E("Couldn't get ConfigString for '%s'", ConfigName);
return;
}
Func(ConfigName, ConfigString);
}
}
}
std::string GetDataDirectory() {
std::string DataDir{};
char const *HomeDir = Paths::GetHomeDirectory();
char const *DataXDG = getenv("XDG_DATA_HOME");
char const *DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (DataOverride) {
// Data override will override the complete directory
DataDir = DataOverride;
}
else {
DataDir = DataXDG ?: HomeDir;
DataDir += "/.fex-emu/";
}
return DataDir;
}
std::string GetConfigDirectory(bool Global) {
@@ -46,7 +167,7 @@ namespace FEXCore::Config {
ConfigDir = GLOBAL_DATA_DIRECTORY;
}
else {
char const *HomeDir = GetHomeDirectory();
char const *HomeDir = Paths::GetHomeDirectory();
char const *ConfigXDG = getenv("XDG_CONFIG_HOME");
char const *ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (ConfigOverride) {
@@ -109,23 +230,6 @@ namespace FEXCore::Config {
return ConfigFile;
}
std::string GetDataDirectory() {
std::string DataDir{};
char const *HomeDir = GetHomeDirectory();
char const *DataXDG = getenv("XDG_DATA_HOME");
char const *DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (DataOverride) {
// Data override will override the complete directory
DataDir = DataOverride;
}
else {
DataDir = DataXDG ?: HomeDir;
DataDir += "/.fex-emu/";
}
return DataDir;
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
}
@@ -225,7 +329,8 @@ namespace FEXCore::Config {
void MetaLayer::MergeConfigMap(const LayerOptions &Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto &it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV ||
it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
MergeEnvironmentVariables(it.first, it.second);
}
else {
@@ -253,7 +358,7 @@ namespace FEXCore::Config {
}
}
std::string ExpandPath(std::string PathName) {
std::string ExpandPath(std::string const &ContainerPrefix, std::string PathName) {
if (PathName.empty()) {
return {};
}
@@ -279,6 +384,69 @@ namespace FEXCore::Config {
return Path;
}
}
else {
// If the containerprefix and pathname isn't empty
// Then we check if the pathname exists in our current namespace
// If the path DOESN'T exist but DOES exist with the prefix applied
// then redirect to the prefix
//
// This might not be expected behaviour for some edge cases but since
// all paths aren't mounted inside the container, then it'll be fine
//
// Main catch case for this is the default thunk install folders
// HostThunks: $CMAKE_INSTALL_PREFIX/lib/fex-emu/HostThunks/
// GuestThunks: $CMAKE_INSTALL_PREFIX/share/fex-emu/GuestThunks/
if (!ContainerPrefix.empty() && !PathName.empty()) {
if (!std::filesystem::exists(PathName)) {
auto ContainerPath = ContainerPrefix + PathName;
if (std::filesystem::exists(ContainerPath)) {
return ContainerPath;
}
}
}
}
return {};
}
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (LoadConfigFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
return "/run/host/";
}
}
}
return {};
}
@@ -294,8 +462,9 @@ namespace FEXCore::Config {
}
}
auto ExpandPathIfExists = [](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(PathName);
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
}
@@ -303,7 +472,7 @@ namespace FEXCore::Config {
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(PathName());
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
@@ -327,7 +496,7 @@ namespace FEXCore::Config {
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKCONFIG)) {
FEX_CONFIG_OPT(PathName, THUNKCONFIG);
auto ExpandedString = ExpandPath(PathName());
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, ExpandedString);
@@ -417,6 +586,17 @@ namespace FEXCore::Config {
}
}
template<>
std::string Value<std::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
}
else {
return std::string(Default);
}
}
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template int8_t Value<int8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int8_t Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
@@ -442,5 +622,143 @@ namespace FEXCore::Config {
}
}
template void Value<std::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, std::list<std::string> *List);
// Application loaders
class MainLoader final : public FEXCore::Config::OptionMapper {
public:
explicit MainLoader();
explicit MainLoader(std::string ConfigFile);
void Load() override;
private:
std::string Config;
};
class AppLoader final : public FEXCore::Config::OptionMapper {
public:
explicit AppLoader(const std::string& Filename, bool Global);
void Load();
private:
std::string Config;
};
class EnvLoader final : public FEXCore::Config::Layer {
public:
explicit EnvLoader(char *const _envp[]);
void Load() override;
private:
char *const *envp;
};
static const std::map<std::string, FEXCore::Config::ConfigOption, std::less<>> ConfigLookup = {{
#define OPT_BASE(type, group, enum, json, default) {#json, FEXCore::Config::ConfigOption::CONFIG_##enum},
#include <FEXCore/Config/ConfigValues.inl>
}};
static const std::vector<std::pair<const char*, FEXCore::Config::ConfigOption>> EnvConfigLookup = {{
#define OPT_BASE(type, group, enum, json, default) {"FEX_" #enum, FEXCore::Config::ConfigOption::CONFIG_##enum},
#include <FEXCore/Config/ConfigValues.inl>
}};
OptionMapper::OptionMapper(FEXCore::Config::LayerType Layer)
: FEXCore::Config::Layer(Layer) {
}
void OptionMapper::MapNameToOption(const char *ConfigName, const char *ConfigString) {
auto it = ConfigLookup.find(ConfigName);
if (it != ConfigLookup.end()) {
Set(it->second, ConfigString);
}
}
MainLoader::MainLoader()
: FEXCore::Config::OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, Config{FEXCore::Config::GetConfigFileLocation()} {
}
MainLoader::MainLoader(std::string ConfigFile)
: FEXCore::Config::OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, Config{std::move(ConfigFile)} {
}
void MainLoader::Load() {
JSON::LoadJSonConfig(Config, [this](const char *Name, const char *ConfigString) {
MapNameToOption(Name, ConfigString);
});
}
AppLoader::AppLoader(const std::string& Filename, bool Global)
: FEXCore::Config::OptionMapper(Global ? FEXCore::Config::LayerType::LAYER_GLOBAL_APP : FEXCore::Config::LayerType::LAYER_LOCAL_APP) {
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
// Immediately load so we can reload the meta layer
Load();
}
void AppLoader::Load() {
JSON::LoadJSonConfig(Config, [this](const char *Name, const char *ConfigString) {
MapNameToOption(Name, ConfigString);
});
}
EnvLoader::EnvLoader(char *const _envp[])
: FEXCore::Config::Layer(FEXCore::Config::LayerType::LAYER_ENVIRONMENT)
, envp {_envp} {
}
void EnvLoader::Load() {
std::unordered_map<std::string_view, std::string_view> EnvMap;
for(const char *const *pvar=envp; pvar && *pvar; pvar++) {
std::string_view Var(*pvar);
size_t pos = Var.rfind('=');
if (std::string::npos == pos)
continue;
std::string_view Ident = Var.substr(0,pos);
std::string_view Value = Var.substr(pos+1);
EnvMap[Ident]=Value;
}
std::function GetVar = [=](const std::string_view id) -> std::optional<std::string_view> {
if (EnvMap.find(id) != EnvMap.end())
return EnvMap.at(id);
// If envp[] was empty, search using std::getenv()
const char* vs = std::getenv(id.data());
if (vs) {
return vs;
}
else {
return std::nullopt;
}
};
std::optional<std::string_view> Value;
for (auto &it : EnvConfigLookup) {
if ((Value = GetVar(it.first)).has_value()) {
Set(it.second, std::string(*Value));
}
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(std::string const *File) {
if (File) {
return std::make_unique<FEXCore::Config::MainLoader>(*File);
}
else {
return std::make_unique<FEXCore::Config::MainLoader>();
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, bool Global) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Global);
}
std::unique_ptr<FEXCore::Config::Layer> CreateEnvironmentLayer(char *const _envp[]) {
return std::make_unique<FEXCore::Config::EnvLoader>(_envp);
}
}
@@ -58,7 +58,7 @@
},
"ThunkHostLibs": {
"Type": "str",
"Default": "",
"Default": "@CMAKE_INSTALL_PREFIX@/lib/fex-emu/HostThunks/",
"ShortArg": "t",
"Desc": [
"Folder to find the host-side thunking libraries."
@@ -66,7 +66,7 @@
},
"ThunkGuestLibs": {
"Type": "str",
"Default": "",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks/",
"ShortArg": "j",
"Desc": [
"Folder to find the guest-side thunking libraries."
@@ -94,6 +94,16 @@
"Desc": [
"Adds an environment variable to the emulated environment."
]
},
"HostEnv": {
"Type": "strarray",
"Default": "",
"ShortArg": "H",
"Desc": [
"Adds an environment variable to the host environment.",
"This can be useful for setting environment variables that thunks can pick up.",
"Typically isn't necessary since the guest libc isn't thunked. But is possible."
]
}
},
"Debug": {
@@ -137,6 +147,13 @@
"Disables optimizations passes for debugging."
]
},
"SRA": {
"Type": "bool",
"Default": "true",
"Desc": [
"Set to false to disable Static Register Allocation"
]
},
"Force32BitAllocator": {
"Type": "bool",
"Default": "false",
+18 -5
View File
@@ -4,9 +4,18 @@
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "FEXCore/Debug/InternalThreadState.h"
#include <string.h>
#include <utility>
namespace FEXCore::HLE {
class SyscallVisitor;
}
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
@@ -34,7 +43,7 @@ namespace FEXCore::Context {
delete CTX;
}
bool InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader) {
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader) {
return CTX->InitCore(Loader);
}
@@ -101,8 +110,8 @@ namespace FEXCore::Context {
void RegisterExternalSyscallVisitor(FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t Syscall, [[maybe_unused]] FEXCore::HLE::SyscallVisitor *Visitor) {
}
void HandleCallback(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->HandleCallback(RIP);
void HandleCallback(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
CTX->HandleCallback(Thread, RIP);
}
void RegisterHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
@@ -117,6 +126,10 @@ namespace FEXCore::Context {
return CTX->CreateThread(NewThreadState, ParentTID);
}
void ExecutionThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->ExecutionThread(Thread);
}
void InitializeThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->InitializeThread(Thread);
}
+90 -19
View File
@@ -1,48 +1,51 @@
#pragma once
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Event.h>
#include <stdint.h>
#ifdef ENABLE_JITSYMBOLS
#include <Common/JITSymbols.h>
#endif
#include <atomic>
#include <condition_variable>
#include <functional>
#include <istream>
#include <map>
#include <memory>
#include <mutex>
#include <optional>
#include <ostream>
#include <set>
#include <shared_mutex>
#include <stddef.h>
#include <string>
#include <unordered_map>
#include <queue>
#include <vector>
namespace FEXCore {
class CodeLoader;
class ThunkHandler;
class BlockSamplingData;
class GdbServer;
class SiganlDelegator;
namespace CPU {
class Arm64JITCore;
class X86JITCore;
}
namespace HLE {
struct SyscallArguments;
class SyscallHandler;
}
}
namespace FEXCore::IR {
class RegisterAllocationPass;
class RegisterAllocationData;
class IRListView;
namespace Validation {
@@ -121,7 +124,9 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(DumpIR, DUMPIR);
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
} Config;
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
@@ -183,7 +188,7 @@ namespace FEXCore::Context {
Context();
~Context();
bool InitCore(FEXCore::CodeLoader *Loader);
FEXCore::Core::InternalThreadState* InitCore(FEXCore::CodeLoader *Loader);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
@@ -199,7 +204,7 @@ namespace FEXCore::Context {
bool GetGdbServerStatus() const { return DebugServer != nullptr; }
void StartGdbServer();
void StopGdbServer();
void HandleCallback(uint64_t RIP);
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
@@ -240,23 +245,80 @@ namespace FEXCore::Context {
};
[[nodiscard]] CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
bool LoadAOTIRCache(int streamfd);
void FinalizeAOTIRCache();
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer);
// Used for thread creation from syscalls
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
* @param CompileThread Is this for the compile service or not?
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
* This is exposed because the CompileService needs to initialize compilers while copying data from
* the paired InternalThreadState that it is compiling code for
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread);
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* OS thread Creation:
* - Thread = CreateThread(NewState, PPID);
* - InitializeThread(Thread);
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
void InitializeThreadData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
void CleanupAfterFork(FEXCore::Core::InternalThreadState *ExceptForThread);
std::vector<FEXCore::Core::InternalThreadState*>* GetThreads() { return &Threads; }
@@ -277,6 +339,15 @@ namespace FEXCore::Context {
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
private:
/**
* @brief Does some final thread initialization
*
* @param Thread The internal FEX thread state object
*
* InitCore and CreateThread both call this to finish up thread object initialization
*/
void InitializeThreadData(FEXCore::Core::InternalThreadState *Thread);
void WaitForIdleWithTimeout();
void NotifyPause();
+186 -27
View File
@@ -2,6 +2,7 @@
#include "Interface/Core/ArchHelpers/MContext.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <atomic>
#include <stdint.h>
@@ -9,6 +10,9 @@
#include <signal.h>
namespace FEXCore::ArchHelpers::Arm64 {
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock, TYPE_HAS_SPLIT_LOCKS);
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock16B, TYPE_16BYTE_SPLIT);
static __uint128_t LoadAcquire128(uint64_t Addr) {
__uint128_t Result{};
uint64_t Lower;
@@ -60,22 +64,12 @@ static bool StoreCAS8(uint8_t &Expected, uint8_t Val, uint64_t Addr) {
return Atom->compare_exchange_strong(Expected, Val);
}
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
static bool RunCASPAL(void *_ucontext, void *_info, uint32_t Size, uint32_t DesiredReg1, uint32_t DesiredReg2, uint32_t ExpectedReg1, uint32_t ExpectedReg2, uint32_t AddressReg) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = (Instr >> 30) & 1;
uint32_t DesiredReg1 = Instr & 0b11111;
uint32_t DesiredReg2 = DesiredReg1 + 1;
uint32_t ExpectedReg1 = (Instr >> 16) & 0b11111;
uint32_t ExpectedReg2 = ExpectedReg1 + 1;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
//Bus_ADRALN check happens in HandleCASPAL and HandleCASPAL_ARMv8
if (Size == 0) {
// 32bit
@@ -94,8 +88,15 @@ bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
// Both cross-cacheline and cross 16byte both need dual CAS loops that can tear
// ARMv8.4 LSE2 solves all atomic issues except cross-cacheline
// Check for Split lock across a cacheline
if ((Addr & 63) > 56) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) > 8) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
uint64_t Alignment = Addr & 0b111;
Addr &= ~0b111ULL;
uint64_t AddrUpper = Addr + 8;
@@ -248,6 +249,87 @@ bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = (Instr >> 30) & 1;
uint32_t DesiredReg1 = Instr & 0b11111;
uint32_t DesiredReg2 = DesiredReg1 + 1;
uint32_t ExpectedReg1 = (Instr >> 16) & 0b11111;
uint32_t ExpectedReg2 = ExpectedReg1 + 1;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
return RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, ExpectedReg1, ExpectedReg2, AddressReg);
}
uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return 0;
}
// caspair
// [1] ldaxp(TMP2.W(), TMP3.W(), MemOperand(MemSrc)); <-- DataReg & AddrReg
// [2] cmp(TMP2.W(), Expected.first.W()); <-- ExpectedReg1
// [3] ccmp(TMP3.W(), Expected.second.W(), NoFlag, Condition::eq); <-- ExpectedREg2
// [4] b(&LoopNotExpected, Condition::ne);
// [5] stlxp(TMP2.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc)); <-- DesiredReg
// [6] cbnz(TMP2.W(), &LoopTop);
// [7] mov(Dst.first.W(), Expected.first.W());
// [8] mov(Dst.second.W(), Expected.second.W());
// [9] b(&LoopExpected);
// [10] mov(Dst.first.W(), TMP2.W());
// [11] mov(Dst.second.W(), TMP3.W());
// [12] clrex();
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Size = (Instr >> 30) & 1;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
uint32_t ExpectedReg1{};
uint32_t ExpectedReg2{};
uint32_t DesiredReg1{};
uint32_t DesiredReg2{};
if(Size != 0) { //Only 32-bit pairs
return 0;
}
for(int i = 1; i < 10; i++) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
ExpectedReg1 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::CCMP_MASK) == FEXCore::ArchHelpers::Arm64::CCMP_INST) {
ExpectedReg2 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) {
DesiredReg1 = (NextInstr & 0x1F);
DesiredReg2 = (NextInstr >> 10) & 0x1F;
}
}
//mov expected into the temp registers used by JIT
mcontext->regs[DataReg] = mcontext->regs[ExpectedReg1];
mcontext->regs[DataReg2] = mcontext->regs[ExpectedReg2];
if(RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, DataReg, DataReg2, AddrReg)) {
return 9 * sizeof(uint32_t); // skip to mov + clrex
} else {
return 0;
}
}
uint16_t DoLoad16(uint64_t Addr) {
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) == 15) {
@@ -430,9 +512,16 @@ uint16_t DoCAS16(
uint64_t Addr,
CASExpectedFn<uint16_t> ExpectedFunction,
CASDesiredFn<uint16_t> DesiredFunction) {
if ((Addr & 63) == 63) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
// 16 bit
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) == 15) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
// Address crosses over 16byte or 64byte threshold
// Need a dual 8bit CAS loop
uint64_t AddrUpper = Addr + 1;
@@ -706,9 +795,16 @@ uint32_t DoCAS32(
uint64_t Addr,
CASExpectedFn<uint32_t> ExpectedFunction,
CASDesiredFn<uint32_t> DesiredFunction) {
if ((Addr & 63) > 60) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
// 32 bit
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) > 12) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
// Address crosses over 16byte threshold
// Needs dual 4 byte CAS loop
uint64_t Alignment = Addr & 0b11;
@@ -936,9 +1032,16 @@ uint64_t DoCAS64(
uint64_t Addr,
CASExpectedFn<uint64_t> ExpectedFunction,
CASDesiredFn<uint64_t> DesiredFunction) {
if ((Addr & 63) > 56) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
// 64bit
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) > 8) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
uint64_t Alignment = Addr & 0b111;
Addr &= ~0b111ULL;
uint64_t AddrUpper = Addr + 8;
@@ -1091,21 +1194,10 @@ uint64_t DoCAS64(
}
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
static bool RunCASAL(void *_ucontext, void *_info, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = 1 << (Instr >> 30);
uint32_t DesiredReg = Instr & 0b11111;
uint32_t ExpectedReg = (Instr >> 16) & 0b11111;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = mcontext->regs[AddressReg];
// Cross-cacheline CAS doesn't work on ARM
@@ -1185,6 +1277,23 @@ bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = 1 << (Instr >> 30);
uint32_t DesiredReg = Instr & 0b11111;
uint32_t ExpectedReg = (Instr >> 16) & 0b11111;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
return RunCASAL(_ucontext, _info, Size, DesiredReg, ExpectedReg, AddressReg);
}
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1527,6 +1636,53 @@ bool HandleAtomicLoad128(void *_ucontext, void *_info, uint32_t Instr) {
return true;
}
static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
{
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
// ARMv8.0 CAS
// [1] ldaxrb(TMP2.W(), MemOperand(MemSrc))
// [2] cmp (TMP2.W(), Expected.W())
// [3] b
// [4] stlxrb(TMP3.W(), Desired.W(), MemOperand(MemSrc)
// [5] cbnz
// [6] mov
// [7] b
// [8] mov (.., TMP2.W());
// [9] clrex
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Instr = PC[0];
uint32_t Size = 1 << (Instr >> 30);
uint32_t AddressReg = GetRnReg(Instr);
uint32_t ResultReg = GetRdReg(Instr); //TMP2
uint32_t DesiredReg = 0;
uint32_t ExpectedReg = 0;
for (size_t i = 1; i < 6; ++i) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
#endif
DesiredReg = GetRdReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
ExpectedReg = GetRmReg(NextInstr);
}
}
//set up CASAL by doing mov(TMP2, Expected)
mcontext->regs[ResultReg] = mcontext->regs[ExpectedReg];
if(RunCASAL(_ucontext, _info, Size, DesiredReg, ResultReg, AddressReg)) {
return 7 * sizeof(uint32_t); //jump to mov to allocated register
} else {
return 0;
}
}
uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1611,6 +1767,9 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
return HandleCAS_NoAtomics(_ucontext, _info); //ARMv8.0 CAS
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::AND_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_AND;
DataSourceReg = GetRmReg(NextInstr);
@@ -30,9 +30,14 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t ALU_OP_MASK = 0x7F'00'00'00;
constexpr uint32_t ADD_INST = 0x0B'00'00'00;
constexpr uint32_t SUB_INST = 0x4B'00'00'00;
constexpr uint32_t CMP_INST = 0x6B'00'00'00;
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t CCMP_MASK = 0x7F'E0'0C'10;
constexpr uint32_t CCMP_INST = 0x7A'40'00'00;
enum ExclusiveAtomicPairType {
TYPE_SWAP,
TYPE_ADD,
@@ -77,6 +82,7 @@ namespace FEXCore::ArchHelpers::Arm64 {
bool HandleAtomicLoad128(void *_ucontext, void *_info, uint32_t Instr);
uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr);
}
@@ -4,6 +4,11 @@
#include <FEXCore/Core/CoreState.h>
#include "aarch64/cpu-aarch64.h"
#include "cpu-features.h"
#include "aarch64/instructions-aarch64.h"
#include "utils-vixl.h"
#include <tuple>
namespace FEXCore::CPU {
#define STATE x28
@@ -143,22 +148,26 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
void Arm64Emitter::SpillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
}
void Arm64Emitter::FillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
}
@@ -222,7 +231,7 @@ void Arm64Emitter::ResetStack() {
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetBuffer()->GetOffsetAddress<uint64_t>(GetCursorOffset());
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
nop();
}
@@ -1,7 +1,16 @@
#pragma once
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "platform-vixl.h"
#include "FEXCore/Config/Config.h"
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <utility>
namespace FEXCore::CPU {
using namespace vixl;
@@ -72,6 +81,8 @@ protected:
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
}
@@ -1,6 +1,7 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include <FEXCore/Utils/LogManager.h>
#include <stdint.h>
namespace FEXCore::ArchHelpers::Arm64 {
@@ -23,4 +24,4 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
}
#endif
}
}
@@ -18,6 +18,7 @@ struct X86ContextBackup {
// RIP and RSP is stored in GPRs here
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
uint64_t sa_mask;
// Guest state
int Signal;
@@ -35,6 +36,7 @@ struct ArmContextBackup {
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
uint64_t sa_mask;
// Guest state
int Signal;
@@ -44,6 +46,11 @@ struct ArmContextBackup {
static constexpr int RedZoneSize = 0;
};
static inline ucontext_t* GetUContext(void* ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
return _context;
}
static inline mcontext_t* GetMContext(void* ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
return &_context->uc_mcontext;
@@ -102,6 +109,7 @@ using ContextBackup = ArmContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
memcpy(&Backup->GPRs[0], &_mcontext->regs[0], 31 * sizeof(uint64_t));
@@ -115,6 +123,9 @@ static inline void BackupContext(void* ucontext, T *Backup) {
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
@@ -123,6 +134,7 @@ static inline void BackupContext(void* ucontext, T *Backup) {
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
@@ -136,6 +148,9 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
ArchHelpers::Context::SetPc(ucontext, Backup->PrevPC);
ArchHelpers::Context::SetSp(ucontext, Backup->PrevSP);
memcpy(&_mcontext->regs[0], &Backup->GPRs[0], 31 * sizeof(uint64_t));
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
@@ -181,6 +196,7 @@ using ContextBackup = X86ContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
@@ -188,6 +204,9 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Copy the FPRState
memcpy(&Backup->FPRState, _mcontext->fpregs, sizeof(X86ContextBackup::FPRState));
// XXX: Save 256bit and 512bit AVX register state
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
@@ -196,12 +215,16 @@ static inline void BackupContext(void* ucontext, T *Backup) {
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Backup->GPRs[0], sizeof(X86ContextBackup::GPRs));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Backup->FPRState, sizeof(X86ContextBackup::FPRState));
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
@@ -209,4 +232,4 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
#endif
} // namespace FEXCore::ArchHelpers::Context
} // namespace FEXCore::ArchHelpers::Context
@@ -2,6 +2,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
#include <fstream>
#include <utility>
namespace FEXCore {
void BlockSamplingData::DumpBlockData() {
+2 -1
View File
@@ -1,6 +1,7 @@
#pragma once
#include <cstdint>
#include <unordered_map>
#include <stdint.h>
namespace FEXCore {
class BlockSamplingData {
+3
View File
@@ -5,8 +5,11 @@ desc: Handles presented capability bits for guest cpu
$end_info$
*/
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "git_version.h"
#include <cstring>
+3 -1
View File
@@ -4,7 +4,9 @@
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <utility>
namespace FEXCore {
namespace Context {
+14 -1
View File
@@ -1,8 +1,21 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include "FEXCore/HLE/Linux/ThreadManagement.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <memory>
#include <pthread.h>
#include <stdio.h>
namespace FEXCore {
static void* ThreadHandler(void *Arg) {
+6 -7
View File
@@ -1,23 +1,22 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <memory>
#include <thread>
#include <unordered_map>
#include <mutex>
#include <queue>
#include <stdint.h>
#include <vector>
namespace FEXCore {
namespace Context {
struct Context;
}
namespace Core {
struct InternalThreadState;
}
namespace IR {
class IRListView;
class RegisterAllocationData;
};
class CompileService final {
+79 -39
View File
@@ -7,42 +7,70 @@ desc: Glues Frontend, OpDispatcher and IR Opts & Compilation, LookupCache, Dispa
$end_info$
*/
#include "Common/MathUtils.h"
#include "Common/Paths.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/DebugData.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/GdbServer.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/Interpreter/InterpreterCore.h"
#include "Interface/Core/JIT/JITCore.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeLoader.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include "Interface/HLE/Thunks/Thunks.h"
#include "FEXCore/Utils/Allocator.h"
#include <xxhash.h>
#include <fstream>
#include <unistd.h>
#include <filesystem>
#include <algorithm>
#include <array>
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <cstdint>
#include <filesystem>
#include <functional>
#include <map>
#include <memory>
#include <mutex>
#include <queue>
#include <set>
#include <shared_mutex>
#include <signal.h>
#include <stdio.h>
#include <string.h>
#include <string>
#include <string_view>
#include <sstream>
#include <sys/mman.h>
#include <unistd.h>
#include <sys/stat.h>
#include "Interface/Core/GdbServer.h"
#include <sys/syscall.h>
#include <type_traits>
#include <unistd.h>
#include <unordered_map>
#include <utility>
#include <vector>
#include <xxhash.h>
namespace FEXCore::CPU {
bool CreateCPUCore(FEXCore::Context::Context *CTX) {
@@ -195,11 +223,7 @@ namespace FEXCore::Context {
}
}
bool Context::InitCore(FEXCore::CodeLoader *Loader) {
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
LocalLoader = Loader;
using namespace FEXCore::Core;
static FEXCore::Core::CPUState CreateDefaultCPUState() {
FEXCore::Core::CPUState NewThreadState{};
// Initialize default CPU state
@@ -217,7 +241,15 @@ namespace FEXCore::Context {
NewThreadState.flags[9] = 1;
NewThreadState.FCW = 0x37F;
NewThreadState.FTW = 0xFFFF;
return NewThreadState;
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
LocalLoader = Loader;
using namespace FEXCore::Core;
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
// We are the parent thread
@@ -228,8 +260,7 @@ namespace FEXCore::Context {
Thread->CurrentFrame->State.rip = StartingRIP = Loader->DefaultRIP();
InitializeThreadData(Thread);
return true;
return Thread;
}
void Context::StartGdbServer() {
@@ -243,8 +274,7 @@ namespace FEXCore::Context {
DebugServer.reset();
}
void Context::HandleCallback(uint64_t RIP) {
auto Thread = Core::ThreadData.Thread;
void Context::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
Thread->CPUBackend->CallbackPtr(Thread->CurrentFrame, RIP);
}
@@ -425,7 +455,13 @@ namespace FEXCore::Context {
auto IRHandler = [Thread](uint64_t Addr, IR::IREmitter *IR) -> void {
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(IR);
Core::LocalIREntry Entry = {Addr, 0ULL, decltype(Entry.IR)(IR->CreateIRCopy()), decltype(Entry.RAData)(Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->PullAllocationData() : nullptr), decltype(Entry.DebugData)(new Core::DebugData())};
Core::LocalIREntry Entry = {Addr, 0ULL,
decltype(Entry.IR)(IR->CreateIRCopy()),
decltype(Entry.RAData)(Thread->PassManager->HasPass("RA")
? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData()
: nullptr),
decltype(Entry.DebugData)(new Core::DebugData())
};
Thread->LocalIRCache.insert({Addr, std::move(Entry)});
};
@@ -456,6 +492,14 @@ namespace FEXCore::Context {
Thread->ThreadWaiting.Wait();
}
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = ::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
}
void Context::RunThread(FEXCore::Core::InternalThreadState *Thread) {
// Tell the thread to start executing
Thread->StartRunning.NotifyAll();
@@ -471,8 +515,10 @@ namespace FEXCore::Context {
Stop(false /* Ignore current thread */);
});
State->CTX = this;
#if _M_ARM_64
bool DoSRA = true;
bool DoSRA = State->CTX->Config.StaticRegisterAllocation;
#else
bool DoSRA = false;
#endif
@@ -482,8 +528,6 @@ namespace FEXCore::Context {
State->PassManager->RegisterSyscallHandler(SyscallHandler);
State->CTX = this;
// Create CPU backend
switch (Config.Core) {
case FEXCore::Config::CONFIG_INTERPRETER:
@@ -784,17 +828,17 @@ namespace FEXCore::Context {
Thread->PassManager->Run(Thread->OpDispatcher.get());
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->GetAllocationData() : nullptr);
IRDumper(Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->GetAllocationData() : nullptr);
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::I("IR 0x%lx:\n%s\n@@@@@\n", GuestRIP, out.str().c_str());
}
auto RAData = Thread->PassManager->GetRAPass() ? Thread->PassManager->GetRAPass()->PullAllocationData() : nullptr;
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
auto IRList = Thread->OpDispatcher->CreateIRCopy();
Thread->OpDispatcher->ResetWorkingList();
@@ -1091,7 +1135,9 @@ namespace FEXCore::Context {
if (NewBlock == 0) {
LogMan::Msg::E("CompileBlockJit: Failed to compile code %lX - aborting process", GuestRIP);
abort();
// Return similar behaviour of SIGILL abort
Frame->Thread->StatusCode = 128 + SIGILL;
Stop(false /* Ignore current thread */);
}
}
@@ -1190,8 +1236,6 @@ namespace FEXCore::Context {
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
delete IRList;
FEXCore::Allocator::free(RAData);
});
}
}
@@ -1226,11 +1270,7 @@ namespace FEXCore::Context {
Core::ThreadData.Thread = Thread;
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = ::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
InitializeThreadTLSData(Thread);
++IdleWaitRefCount;
@@ -2,17 +2,32 @@
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/X86HelperGen.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <array>
#include <bit>
#include <cmath>
#include <cstdint>
#include <memory>
#include <stddef.h>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/operands-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#ifdef ENABLE_JITSYMBOLS
#include <unistd.h>
#endif
namespace FEXCore::CPU {
@@ -27,8 +42,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
auto Buffer = GetBuffer();
DispatchPtr = Buffer->GetOffsetAddress<CPUBackend::AsmDispatch>(GetCursorOffset());
DispatchPtr = GetCursorAddress<CPUBackend::AsmDispatch>();
// while (true) {
// Ptr = FindBlock(RIP)
@@ -61,7 +75,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
AbsoluteLoopTopAddressFillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled) {
FillStaticRegs();
@@ -180,11 +194,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
PopCalleeSavedRegisters();
@@ -194,7 +208,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
ExitFunctionLinkerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
@@ -231,7 +245,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
SignalHandlerReturnAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
SignalHandlerReturnAddress = GetCursorAddress<uint64_t>();
// Now to get back to our old location we need to do a fault dance
// We can't use SIGTRAP here since gdb catches it and never gives it to the application!
@@ -239,12 +253,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
{
ThreadPauseHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
ThreadPauseHandlerAddress = GetCursorAddress<uint64_t>();
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
@@ -254,7 +268,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ldr(x2, &l_Sleep);
blr(x2);
PauseReturnInstruction = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
PauseReturnInstruction = GetCursorAddress<uint64_t>();
// Fault to start running again
hlt(0);
}
@@ -275,7 +289,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// On return to the thunk, the thunk can get whatever its return value is from the thread context depending on ABI handling on its end
// When the thunk itself returns, it'll do its regular return logic there
// void ReentrantCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
CallbackPtr = Buffer->GetOffsetAddress<CPUBackend::JITCallback>(GetCursorOffset());
CallbackPtr = GetCursorAddress<CPUBackend::JITCallback>();
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
@@ -324,7 +338,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
FinalizeCode();
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
End = GetCursorAddress<uint64_t>();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
@@ -3,7 +3,13 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::CPU {
@@ -15,4 +21,4 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
void SpillSRA(void *ucontext) override;
};
}
}
@@ -1,8 +1,23 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Common/MathUtils.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <atomic>
#include <condition_variable>
#include <bits/types/siginfo_t.h>
#include <signal.h>
#include <string.h>
namespace FEXCore::CPU {
@@ -65,10 +80,38 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
}
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
static uint32_t ConvertSignalToTrapNo(int Signal, siginfo_t *HostSigInfo) {
switch (Signal) {
case SIGSEGV:
if (HostSigInfo->si_code == SEGV_MAPERR ||
HostSigInfo->si_code == SEGV_ACCERR) {
// Protection fault
return X86State::X86_TRAPNO_PF;
}
break;
}
// Unknown mapping, fall back to old behaviour and just pass signal
return Signal;
}
static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
switch (Signal) {
case SIGSEGV:
if (HostSigInfo->si_code == SEGV_MAPERR ||
HostSigInfo->si_code == SEGV_ACCERR) {
// Protection fault
// Always a user fault for us
// XXX: PF_PROT and PF_WRITE
return X86State::X86_PF_USER;
}
break;
}
// Not a page fault issue
return 0;
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
@@ -79,70 +122,95 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
uint64_t OldPC = ArchHelpers::Context::GetPc(ucontext);
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
uint64_t OldGuestSP = Frame->State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
// Pulling from context here
bool Is64BitMode = CTX->Config.Is64BitMode;
uint64_t SignalReturn = CTX->X86CodeGen.SignalReturn;
// Spill the SRA regardless of signal handler type
// We are going to be returning to the top of the dispatcher which will fill again
// Otherwise we might load garbage
if (SRAEnabled) {
if (IsAddressInJITCode(OldPC, false)) {
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
} else {
if (!IsAddressInJITCode(OldPC, true)) {
// This is likely to cause issues but in some cases it isn't fatal
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
// Only throw a log message in this case
LogMan::Msg::E("Signals in dispatcher have unsynchronized context");
}
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
// altstack is only used if the signal handler was setup with SA_ONSTACK
if (GuestAction->sa_flags & SA_ONSTACK) {
// Additionally the altstack is only used if the enabled (SS_DISABLE flag is not set)
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
}
if (Is64BitMode) {
// Back up past the redzone, which is 128bytes
// 32-bit doesn't have a redzone
NewGuestSP -= 128;
}
// siginfo_t
siginfo_t *HostSigInfo = reinterpret_cast<siginfo_t*>(info);
if (GuestAction->sa_flags & SA_SIGINFO &&
!(HostSigInfo->si_code == SI_QUEUE || // If the siginfo comes from sigqueue or user then we don't need to check
HostSigInfo->si_code == SI_USER)) {
if (SRAEnabled) {
if (!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
} else {
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
}
}
if (GuestAction->sa_flags & SA_SIGINFO) {
// Setup ucontext a bit
if (CTX->Config.Is64BitMode) {
if (Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86_64::_libc_fpstate));
uint64_t FPStateLocation = NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86_64::ucontext_t);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86_64::ucontext_t));
uint64_t UContextLocation = NewGuestSP;
NewGuestSP -= sizeof(siginfo_t);
NewGuestSP = AlignDown(NewGuestSP, alignof(siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
siginfo_t *guest_siginfo = reinterpret_cast<siginfo_t*>(SigInfoLocation);
// We have extended float information
guest_uctx->uc_flags |= FEXCore::x86_64::UC_FP_XSTATE;
guest_uctx->uc_flags = FEXCore::x86_64::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = &guest_uctx->__fpregs_mem;
guest_uctx->uc_mcontext.fpregs = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(FPStateLocation);
FEXCore::x86_64::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(FPStateLocation);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CSGSFS] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = Signal;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_OLDMASK] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CR2] = 0;
@@ -167,15 +235,15 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#undef COPY_REG
// Copy float registers
memcpy(guest_uctx->__fpregs_mem._st, Frame->State.mm, sizeof(Frame->State.mm));
memcpy(guest_uctx->__fpregs_mem._xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
memcpy(fpstate->_st, Frame->State.mm, sizeof(Frame->State.mm));
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = Frame->State.FCW;
guest_uctx->__fpregs_mem.ftw = Frame->State.FTW;
fpstate->fcw = Frame->State.FCW;
fpstate->ftw = Frame->State.FTW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
fpstate->fsw =
(Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] << 11) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] << 8) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] << 9) |
@@ -196,27 +264,34 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
// XXX: 32bit Support
NewGuestSP -= sizeof(FEXCore::x86::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::_libc_fpstate));
uint64_t FPStateLocation = NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::ucontext_t);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::ucontext_t));
uint64_t UContextLocation = NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::siginfo_t);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
FEXCore::x86::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(UContextLocation);
FEXCore::x86::siginfo_t *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(SigInfoLocation);
// We have extended float information
guest_uctx->uc_flags |= FEXCore::x86::UC_FP_XSTATE;
guest_uctx->uc_flags = FEXCore::x86::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = static_cast<uint32_t>(reinterpret_cast<uint64_t>(&guest_uctx->__fpregs_mem));
guest_uctx->uc_mcontext.fpregs = static_cast<uint32_t>(FPStateLocation);
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(FPStateLocation);
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS] = Frame->State.gs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS] = Frame->State.fs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = Signal;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS] = Frame->State.cs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL] = 0;
@@ -236,20 +311,20 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
#undef COPY_REG
// Copy float registers
memcpy(guest_uctx->__fpregs_mem._st, Frame->State.mm, sizeof(Frame->State.mm));
if (0) {
// XXX: Handle XMM
// memcpy(guest_uctx->__fpregs_mem._xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
guest_uctx->__fpregs_mem.status = FEXCore::x86::fpstate_magic::MAGIC_XFPSTATE;
}
else {
guest_uctx->__fpregs_mem.status = FEXCore::x86::fpstate_magic::MAGIC_FPU;
for (size_t i = 0; i < 8; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&fpstate->_st[i], &Frame->State.mm[i], 10);
}
// Extended XMM state
fpstate->status = FEXCore::x86::fpstate_magic::MAGIC_XFPSTATE;
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = Frame->State.FCW;
guest_uctx->__fpregs_mem.ftw = Frame->State.FTW;
fpstate->fcw = Frame->State.FCW;
fpstate->ftw = Frame->State.FTW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
fpstate->fsw =
(Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] << 11) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] << 8) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] << 9) |
@@ -269,6 +344,12 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
switch (Signal) {
case SIGSEGV:
case SIGBUS:
// Macro expansion to get the si_addr
// This is the address trying to be accessed, not the RIP
guest_siginfo->_sifields._sigfault.addr = static_cast<uint32_t>(reinterpret_cast<uintptr_t>(HostSigInfo->si_addr));
break;
case SIGFPE:
case SIGILL:
// Macro expansion to get the si_addr
// Can't really give a real result here. Pull from the context for now
guest_siginfo->_sifields._sigfault.addr = Frame->State.rip;
@@ -280,10 +361,15 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_siginfo->_sifields._sigchld.utime = HostSigInfo->si_utime;
guest_siginfo->_sifields._sigchld.stime = HostSigInfo->si_stime;
break;
default:
// Hope for the best, most things just copy over
memcpy(&guest_siginfo->_sifields, &HostSigInfo->_sifields, sizeof(siginfo_t));
break;
case SIGALRM:
case SIGVTALRM:
guest_siginfo->_sifields._timer.tid = HostSigInfo->si_timerid;
guest_siginfo->_sifields._timer.overrun = HostSigInfo->si_overrun;
guest_siginfo->_sifields._timer.sigval.sival_int = HostSigInfo->si_int;
break;
default:
LogMan::Msg::E("Unhandled siginfo_t for signal: %d\n", Signal);
break;
}
NewGuestSP -= 4;
@@ -297,7 +383,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
if (!CTX->Config.Is64BitMode) {
if (!Is64BitMode) {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = Signal;
}
@@ -305,21 +391,29 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
if (CTX->Config.Is64BitMode) {
Frame->State.gregs[X86State::REG_RDI] = Signal;
if (Is64BitMode) {
Frame->State.gregs[FEXCore::X86State::REG_RDI] = Signal;
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
*(uint64_t*)NewGuestSP = SignalReturn;
Frame->State.gregs[FEXCore::X86State::REG_RSP] = NewGuestSP;
}
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
LOGMAN_THROW_A(CTX->X86CodeGen.SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
*(uint32_t*)NewGuestSP = SignalReturn;
LOGMAN_THROW_A(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[FEXCore::X86State::REG_RSP] = NewGuestSP;
}
// The guest starts its signal frame with a zero initialized FPU
// Set that up now. Little bit costly but it's a requirement
// This state will be restored on rt_sigreturn
memset(Frame->State.xmm, 0, sizeof(Frame->State.xmm));
memset(Frame->State.mm, 0, sizeof(Frame->State.mm));
Frame->State.FCW = 0x37F;
Frame->State.FTW = 0xFFFF;
return true;
}
@@ -365,9 +459,6 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
}
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
@@ -1,11 +1,24 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "Interface/Context/Context.h"
#include <bits/types/stack_t.h>
#include <cstdint>
#include <stddef.h>
#include <stack>
#include <tuple>
#include <vector>
namespace FEXCore {
struct GuestSigAction;
}
namespace FEXCore::Core {
struct CpuStateFrame;
struct InternalThreadState;
}
namespace FEXCore::CPU {
@@ -3,10 +3,21 @@
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Allocator.h>
#include <cmath>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <sys/mman.h>
#include "xbyak/xbyak.h"
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
@@ -2,11 +2,17 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Utils/Allocator.h>
#define XBYAK64
#include <xbyak/xbyak.h>
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
+9 -1
View File
@@ -7,14 +7,18 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/InternalThreadState.h"
#include <array>
#include <assert.h>
#include <algorithm>
#include <cstring>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <set>
#include <sys/mman.h>
@@ -696,6 +700,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return NormalOp(&X87Ops[X87Op], X87Op);
}
else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
uint16_t map_select = 1;
uint16_t pp = 0;
@@ -723,6 +728,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
if (LocalInfo->Type >= FEXCore::X86Tables::TYPE_VEX_GROUP_12 &&
LocalInfo->Type <= FEXCore::X86Tables::TYPE_VEX_GROUP_17) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
@@ -740,6 +746,8 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return NormalOp(LocalInfo, Op);
}
else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
FEXCORE_TELEMETRY_SET(EVEXOpTelem, 1);
/* uint8_t P1 = */ ReadByte();
/* uint8_t P2 = */ ReadByte();
/* uint8_t P3 = */ ReadByte();
+5 -2
View File
@@ -2,12 +2,12 @@
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <array>
#include <cstdint>
#include <utility>
#include <set>
#include <stack>
#include <stddef.h>
#include <vector>
namespace FEXCore::Context {
@@ -89,5 +89,8 @@ private:
};
const uint8_t *AdjustAddrForSpecialRegion(uint8_t const* _InstStream, uint64_t EntryPoint, uint64_t RIP);
FEXCORE_TELEMETRY_INIT(VEXOpTelem, TYPE_USES_VEX_OPS);
FEXCORE_TELEMETRY_INIT(EVEXOpTelem, TYPE_USES_EVEX_OPS);
};
}
+17 -6
View File
@@ -8,30 +8,41 @@ $end_info$
#include <cstdlib>
#include <cstdio>
#include <iomanip>
#include <iostream>
#include <sstream>
#include <string>
#include <memory>
#include <optional>
#include "Common/NetStream.h"
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <cstring>
#include <errno.h>
#include <fcntl.h>
#include <fmt/format.h>
#include <fstream>
#include <fmt/format.h>
#include <netdb.h>
#include <signal.h>
#include <stddef.h>
#include <string_view>
#include <sys/socket.h>
#include <sys/types.h>
#include <unistd.h>
#include <utility>
#include <vector>
#include "GdbServer.h"
#include <FEXCore/Core/CodeLoader.h>
#include <FEXCore/Core/X86Enums.h>
namespace FEXCore
{
+9 -6
View File
@@ -5,18 +5,21 @@ $end_info$
*/
#pragma once
#include <mutex>
#include <thread>
#include "Interface/Context/Context.h"
#include "Common/NetStream.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Threads.h>
#include <istream>
#include <memory>
#include <mutex>
#include <stdint.h>
#include <string>
namespace FEXCore {
namespace Context {
struct Context;
}
class GdbServer {
public:
GdbServer(FEXCore::Context::Context *ctx);
@@ -1,30 +1,31 @@
#include "Common/MathUtils.h"
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/DebugData.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include "Interface/HLE/Thunks/Thunks.h"
#include <atomic>
#include <cmath>
#include <limits>
#include <vector>
#include <memory>
#include <bits/types/stack_t.h>
#include <signal.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include "InterpreterOps.h"
namespace FEXCore::IR {
class IRListView;
class RegisterAllocationData;
}
namespace FEXCore::CPU {
class CPUBackend;
static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
@@ -117,7 +118,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -1,17 +1,16 @@
#include "Common/MathUtils.h"
#include "Common/SoftFloat.h"
#include "Common/SoftFloat-3e/softfloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "InterpreterOps.h"
#ifdef _M_ARM_64
#include "Interface/Core/ArchHelpers/Arm64.h"
#endif
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/DebugData.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
@@ -21,15 +20,18 @@
#include "Interface/HLE/Thunks/Thunks.h"
#include <alloca.h>
#include <algorithm>
#include <atomic>
#include <bit>
#include <cmath>
#include <cstdint>
#include <limits>
#include <vector>
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
#include <unistd.h>
#include <memory>
#include <stddef.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
namespace FEXCore::CPU {
@@ -92,6 +94,20 @@ static uint64_t AtomicFetchNeg(uint64_t *Addr) {
return Expected;
}
template<typename T>
static T AtomicCompareAndSwap(T expected, T desired, T *addr)
{
std::atomic<T> *Data = reinterpret_cast<std::atomic<T>*>(addr);
T Src1 = expected;
T Src2 = desired;
T Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
return Result ? Src1 : Expected;
}
#else
// Needs to match what the AArch64 JIT and unaligned signal handler expects
uint8_t AtomicFetchNeg(uint8_t *Addr) {
@@ -186,6 +202,141 @@ uint64_t AtomicFetchNeg(uint64_t *Addr) {
return Result;
}
template<typename T>
static T AtomicCompareAndSwap(T expected, T desired, T *addr);
template<>
uint8_t AtomicCompareAndSwap(uint8_t expected, uint8_t desired, uint8_t *addr) {
using Type = uint8_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxrb %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected], uxtb;
b.ne 2f;
stlxrb %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint16_t AtomicCompareAndSwap(uint16_t expected, uint16_t desired, uint16_t *addr) {
using Type = uint16_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxrh %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected], uxth;
b.ne 2f;
stlxrh %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint32_t AtomicCompareAndSwap(uint32_t expected, uint32_t desired, uint32_t *addr) {
using Type = uint32_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxr %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected];
b.ne 2f;
stlxr %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint64_t AtomicCompareAndSwap(uint64_t expected, uint64_t desired, uint64_t *addr) {
using Type = uint64_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxr %[Tmp], [%[Memory]];
cmp %[Tmp], %[Expected];
b.ne 2f;
stlxr %w[Tmp2], %[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %[Result], %[Expected];
b 3f;
2:
mov %[Result], %[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
#endif
namespace AES {
@@ -1515,14 +1666,11 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
// Size is the size of each pair element
switch (Size) {
case 4: {
std::atomic<uint64_t> *Data = *GetSrc<std::atomic<uint64_t> **>(SSAData, Op->Header.Args[2]);
uint64_t Src1 = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]);
uint64_t Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
GD = Result ? Src1 : Expected;
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(SSAData, Op->Header.Args[2])
);
break;
}
case 8: {
@@ -1592,7 +1740,32 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
}
}
memset(GDP, 0, 16);
memcpy(GDP, Data, Op->Size);
switch (OpSize) {
case 1: {
const uint8_t *D = (const uint8_t*)Data;
GD = *D;
break;
}
case 2: {
const uint16_t *D = (const uint16_t*)Data;
GD = *D;
break;
}
case 4: {
const uint32_t *D = (const uint32_t*)Data;
GD = *D;
break;
}
case 8: {
const uint64_t *D = (const uint64_t*)Data;
GD = *D;
break;
}
default:
memcpy(GDP, Data, Op->Size);
break;
}
break;
}
case IR::OP_VLOADMEMELEMENT: {
@@ -1619,7 +1792,29 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
case MEM_OFFSET_SXTW.Val: Data += (int32_t)Offset; break;
}
}
memcpy(Data, GetSrc<void*>(SSAData, Op->Value), Op->Size);
switch (OpSize) {
case 1: {
*reinterpret_cast<uint8_t*>(Data) = *GetSrc<uint8_t*>(SSAData, Op->Value);
break;
}
case 2: {
*reinterpret_cast<uint16_t*>(Data) = *GetSrc<uint16_t*>(SSAData, Op->Value);
break;
}
case 4: {
*reinterpret_cast<uint32_t*>(Data) = *GetSrc<uint32_t*>(SSAData, Op->Value);
break;
}
case 8: {
*reinterpret_cast<uint64_t*>(Data) = *GetSrc<uint64_t*>(SSAData, Op->Value);
break;
}
default:
memcpy(Data, GetSrc<void*>(SSAData, Op->Value), Op->Size);
break;
}
break;
}
case IR::OP_VSTOREMEMELEMENT: {
@@ -2198,46 +2393,35 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
auto Size = OpSize;
switch (Size) {
case 1: {
std::atomic<uint8_t> *Data = *GetSrc<std::atomic<uint8_t> **>(SSAData, Op->Header.Args[2]);
uint8_t Src1 = *GetSrc<uint8_t*>(SSAData, Op->Header.Args[0]);
uint8_t Src2 = *GetSrc<uint8_t*>(SSAData, Op->Header.Args[1]);
uint8_t Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
GD = Result ? Src1 : Expected;
GD = AtomicCompareAndSwap(
*GetSrc<uint8_t*>(SSAData, Op->Header.Args[0]),
*GetSrc<uint8_t*>(SSAData, Op->Header.Args[1]),
*GetSrc<uint8_t**>(SSAData, Op->Header.Args[2])
);
break;
}
case 2: {
std::atomic<uint16_t> *Data = *GetSrc<std::atomic<uint16_t> **>(SSAData, Op->Header.Args[2]);
uint16_t Src1 = *GetSrc<uint16_t*>(SSAData, Op->Header.Args[0]);
uint16_t Src2 = *GetSrc<uint16_t*>(SSAData, Op->Header.Args[1]);
uint16_t Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
GD = Result ? Src1 : Expected;
GD = AtomicCompareAndSwap(
*GetSrc<uint16_t*>(SSAData, Op->Header.Args[0]),
*GetSrc<uint16_t*>(SSAData, Op->Header.Args[1]),
*GetSrc<uint16_t**>(SSAData, Op->Header.Args[2])
);
break;
}
case 4: {
std::atomic<uint32_t> *Data = *GetSrc<std::atomic<uint32_t> **>(SSAData, Op->Header.Args[2]);
uint32_t Src1 = *GetSrc<uint32_t*>(SSAData, Op->Header.Args[0]);
uint32_t Src2 = *GetSrc<uint32_t*>(SSAData, Op->Header.Args[1]);
uint32_t Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
GD = Result ? Src1 : Expected;
GD = AtomicCompareAndSwap(
*GetSrc<uint32_t*>(SSAData, Op->Header.Args[0]),
*GetSrc<uint32_t*>(SSAData, Op->Header.Args[1]),
*GetSrc<uint32_t**>(SSAData, Op->Header.Args[2])
);
break;
}
case 8: {
std::atomic<uint64_t> *Data = *GetSrc<std::atomic<uint64_t> **>(SSAData, Op->Header.Args[2]);
uint64_t Src1 = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]);
uint64_t Expected = Src1;
bool Result = Data->compare_exchange_strong(Expected, Src2);
GD = Result ? Src1 : Expected;
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(SSAData, Op->Header.Args[2])
);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", Size); break;
@@ -1,9 +1,13 @@
#pragma once
#include <stdint.h>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::IR {
class IRListView;
struct IROp_Header;
}
namespace FEXCore::Core{
@@ -32,11 +36,11 @@ namespace FEXCore::CPU {
FallbackABI ABI;
void *fn;
};
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
};
};
};
@@ -44,15 +44,12 @@ DEF_OP(CASPair) {
aarch64::Label LoopNotExpected;
aarch64::Label LoopExpected;
bind(&LoopTop);
nop();
ldaxp(TMP2.W(), TMP3.W(), MemOperand(MemSrc));
nop();
cmp(TMP2.W(), Expected.first.W());
ccmp(TMP3.W(), Expected.second.W(), NoFlag, Condition::eq);
b(&LoopNotExpected, Condition::ne);
nop();
stlxp(TMP2.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc));
nop();
cbnz(TMP2.W(), &LoopTop);
mov(Dst.first.W(), Expected.first.W());
mov(Dst.second.W(), Expected.second.W());
@@ -73,15 +70,12 @@ DEF_OP(CASPair) {
aarch64::Label LoopNotExpected;
aarch64::Label LoopExpected;
bind(&LoopTop);
nop();
ldaxp(TMP2.X(), TMP3.X(), MemOperand(MemSrc));
nop();
cmp(TMP2.X(), Expected.first.X());
ccmp(TMP3.X(), Expected.second.X(), NoFlag, Condition::eq);
b(&LoopNotExpected, Condition::ne);
nop();
stlxp(TMP2.X(), Desired.first.X(), Desired.second.X(), MemOperand(MemSrc));
nop();
cbnz(TMP2.X(), &LoopTop);
mov(Dst.first.X(), Expected.first.X());
mov(Dst.second.X(), Expected.second.X());
+21 -32
View File
@@ -392,32 +392,22 @@ bool Arm64JITCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXP_MASK) == FEXCore::ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
// Convert to LDP
uint32_t LDP = 0b0010'1001'0100'0000'0000'0000'0000'0000;
LDP |= Size << 31;
LDP |= DataReg2 << 10;
LDP |= AddrReg << 5;
LDP |= DataReg;
PC[-1] = DMB;
PC[0] = LDP;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
//Should be compare and swap pair only. LDAXP not used elsewhere
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleCASPAL_ARMv8(ucontext, info, Instr);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) { // STLXP
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
// Convert to STP
uint32_t STP = 0b0010'1001'0000'0000'0000'0000'0000'0000;
STP |= Size << 31;
STP |= DataReg2 << 10;
STP |= AddrReg << 5;
STP |= DataReg;
PC[-1] = DMB;
PC[0] = STP;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
//Should not trigger - middle of an LDAXP/STAXP pair.
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
@@ -482,7 +472,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = true;
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
@@ -496,7 +486,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetRAPass();
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
#if DEBUG
Decoder.AppendVisitor(&Disasm)
@@ -561,7 +551,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -781,8 +771,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
// X1-X3 = Temp
// X4-r18 = RA
auto Buffer = GetBuffer();
auto GuestEntry = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
auto GuestEntry = GetCursorAddress<uint64_t>();
if (CTX->GetGdbServerStatus()) {
aarch64::Label RunBlock;
@@ -849,7 +838,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.push_back({Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()), 0, IR->GetID(BlockNode)});
DebugData->Subblocks.push_back({GetCursorAddress<uintptr_t>(), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
@@ -861,7 +850,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()) - DebugData->Subblocks.back().HostCodeStart;
DebugData->Subblocks.back().HostCodeSize = GetCursorAddress<uintptr_t>() - DebugData->Subblocks.back().HostCodeStart;
}
}
@@ -874,7 +863,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
FinalizeCode();
auto CodeEnd = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
auto CodeEnd = GetCursorAddress<uint64_t>();
CPU.EnsureIAndDCacheCoherency(reinterpret_cast<void*>(GuestEntry), CodeEnd - reinterpret_cast<uint64_t>(GuestEntry));
if (DebugData) {
@@ -5,7 +5,14 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stdint.h>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
@@ -5,7 +5,14 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stdint.h>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
@@ -4,14 +4,28 @@ tags: backend|x86-64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/HLE/Thunks/Thunks.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
@@ -5,7 +5,13 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -5,7 +5,12 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
@@ -5,7 +5,12 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
+25 -9
View File
@@ -8,20 +8,36 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <cmath>
#include <algorithm>
#include <array>
#include <bits/types/stack_t.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <signal.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <sys/mman.h>
#include <tuple>
#include <unordered_map>
#include <utility>
#include <vector>
#include <xbyak/xbyak.h>
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
@@ -304,7 +320,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
{
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetRAPass();
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AllocateRegisterSet(RegisterCount, RegisterClasses);
RAPass->AddRegisters(FEXCore::IR::GPRClass, NumGPRs);
@@ -360,7 +376,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -5,9 +5,15 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <cmath>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -4,8 +4,18 @@ tags: backend|x86-64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
@@ -5,7 +5,13 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <array>
#include <stdint.h>
#include <utility>
namespace FEXCore::CPU {
@@ -5,8 +5,14 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
+4 -3
View File
@@ -5,10 +5,11 @@ desc: Stores information about blocks, and provides C++ implementations to looku
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/LookupCache.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include <sys/mman.h>
+9 -1
View File
@@ -1,10 +1,18 @@
#pragma once
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <functional>
#include <map>
#include <stddef.h>
#include <utility>
#include <vector>
namespace FEXCore {
namespace Context {
struct Context;
}
class LookupCache {
public:
+68 -36
View File
@@ -7,16 +7,22 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <bit>
#include <climits>
#include <cstddef>
#include <cstdint>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <algorithm>
#include <array>
#include <cstdint>
#include <tuple>
namespace FEXCore::IR {
@@ -75,13 +81,25 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
};
static_assert(GPRIndexes_64.size() == GPRIndexes_32.size());
static std::array<uint64_t, SyscallArgs> GPRIndexes_Hangover = {
FEXCore::X86State::REG_RCX,
};
size_t NumArguments{};
const auto OSABI = CTX->SyscallHandler->GetOSABI();
if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX64) {
NumArguments = GPRIndexes_64.size();
GPRIndexes = &GPRIndexes_64;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX32) {
NumArguments = GPRIndexes_64.size();
GPRIndexes = &GPRIndexes_32;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
NumArguments = 1;
GPRIndexes = &GPRIndexes_Hangover;
}
else {
LogMan::Msg::D("Unhandled OSABI syscall");
}
@@ -91,16 +109,34 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), NewRIP);
const auto& GPRIndicesRef = *GPRIndexes;
auto SyscallOp = _Syscall(
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[0] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[1] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[2] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[3] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[4] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[5] * 8, GPRClass),
_LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[6] * 8, GPRClass));
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RAX]), SyscallOp);
OrderedNode *Arguments[SyscallArgs] {
InvalidNode,
InvalidNode,
InvalidNode,
InvalidNode,
InvalidNode,
InvalidNode,
InvalidNode,
};
for (size_t i = 0; i < NumArguments; ++i) {
Arguments[i] = _LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs) + GPRIndicesRef[i] * 8, GPRClass);
}
auto SyscallOp = _Syscall(
Arguments[0],
Arguments[1],
Arguments[2],
Arguments[3],
Arguments[4],
Arguments[5],
Arguments[6]);
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
// Hangover doesn't want us returning a result here
// syscall is being abused as a thunk for now.
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RAX]), SyscallOp);
}
}
void OpDispatchBuilder::ThunkOp(OpcodeArgs) {
@@ -208,17 +244,21 @@ void OpDispatchBuilder::IRETOp(OpcodeArgs) {
//eflags (lower 16 used)
auto eflags = _LoadMem(GPRClass, GPRSize, SP, GPRSize);
SetPackedRFLAG(false, eflags);
SP = _Add(SP, Constant);
if (CTX->Config.Is64BitMode) {
// RSP and SS only happen in 64-bit mode or if this is a CPL mode jump!
SP = _Add(SP, Constant);
// RSP
// FEX doesn't support a CPL mode switch, so don't need to worry about this on 32-bit
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RSP]), _LoadMem(GPRClass, GPRSize, SP, GPRSize));
SP = _Add(SP, Constant);
//ss
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, ss), _LoadMem(GPRClass, GPRSize, SP, GPRSize));
SP = _Add(SP, Constant);
}
else {
// Store the stack in 32-bit mode
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RSP]), SP);
}
_ExitFunction(NewRIP);
BlockSetRIP = true;
@@ -1545,7 +1585,14 @@ void OpDispatchBuilder::MOVSegOp(OpcodeArgs) {
DecodeFailure = true;
return;
}
StoreResult(GPRClass, Op, Segment, -1);
if (DestIsMem(Op)) {
// If the destination is memory then we always store 16-bits only
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Segment, 2, -1);
}
else {
// If the destination is a GPR then we follow register storing rules
StoreResult(GPRClass, Op, Segment, -1);
}
}
}
@@ -3113,11 +3160,6 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
}
void OpDispatchBuilder::STOSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX) {
LogMan::Msg::E("Invalid REPNE on STOS");
DecodeFailure = true;
return;
}
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::E("Can't handle adddress size");
DecodeFailure = true;
@@ -3126,7 +3168,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
const auto GPRSize = CTX->GetGPRSize();
const auto Size = GetSrcSize(Op);
const bool Repeat = (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX) != 0;
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) != 0;
if (!Repeat) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -3221,11 +3263,6 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
}
void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX) {
LogMan::Msg::E("Invalid REPNE on MOVS");
DecodeFailure = true;
return;
}
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::E("Can't handle adddress size");
DecodeFailure = true;
@@ -3242,7 +3279,7 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
auto DF = GetRFLAG(FEXCore::X86State::RFLAG_DF_LOC);
auto PtrDir = _Select(FEXCore::IR::COND_EQ, DF, _Constant(0), SizeConst, NegSizeConst);
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX) {
if (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) {
// Create all our blocks
auto LoopHead = CreateNewCodeBlockAfter(GetCurrentBlock());
auto LoopTail = CreateNewCodeBlockAfter(LoopHead);
@@ -3437,11 +3474,6 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
}
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX) {
LogMan::Msg::E("Invalid REPNE on LODS");
DecodeFailure = true;
return;
}
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::E("Can't handle adddress size");
DecodeFailure = true;
@@ -3450,7 +3482,7 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
const auto GPRSize = CTX->GetGPRSize();
const auto Size = GetSrcSize(Op);
const bool Repeat = (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX) != 0;
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) != 0;
if (!Repeat) {
OrderedNode *Dest_RSI = _LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RSI]), GPRClass);
+5 -3
View File
@@ -3,7 +3,8 @@
#include "Interface/Core/Frontend.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
@@ -12,9 +13,10 @@
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <functional>
#include <map>
#include <set>
#include <stddef.h>
#include <utility>
#include <vector>
namespace FEXCore::IR {
class Pass;
@@ -5,11 +5,16 @@ desc: Handles x86/64 Crypto instructions to IR
$end_info$
*/
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/X86Enums.h>
#include <stdint.h>
namespace FEXCore::IR {
class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
@@ -5,9 +5,17 @@ desc: Handles x86/64 flag generation
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <array>
#include <cstdint>
namespace FEXCore::IR {
constexpr std::array<uint32_t, 17> FlagOffsets = {
@@ -5,9 +5,20 @@ desc: Handles x86/64 Vector instructions to IR
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <bit>
#include <cstdint>
#include <stddef.h>
namespace FEXCore::IR {
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
@@ -1278,6 +1289,13 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
_StoreMem(GPRClass, 2, MemLocation, FSW, 2);
}
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
_StoreMem(GPRClass, 2, MemLocation, FTW, 2);
}
// BYTE | 0 1 | 2 3 | 4 | 5 | 6 7 | 8 9 | a b | c d | e f |
// ------------------------------------------
// 32 | ST0/MM0 | <R>
@@ -1296,14 +1314,14 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
// 240 | XMM5
// 256 | XMM6
// 272 | XMM7
// 288 | XMM8
// 304 | XMM9
// 320 | XMM10
// 336 | XMM11
// 352 | XMM12
// 368 | XMM13
// 384 | XMM14
// 400 | XMM15
// 288 | 64BitMode ? <R> : XMM8
// 304 | 64BitMode ? <R> : XMM9
// 320 | 64BitMode ? <R> : XMM10
// 336 | 64BitMode ? <R> : XMM11
// 352 | 64BitMode ? <R> : XMM12
// 368 | 64BitMode ? <R> : XMM13
// 384 | 64BitMode ? <R> : XMM14
// 400 | 64BitMode ? <R> : XMM15
// 416 | <R>
// 432 | <R>
// 448 | <R>
@@ -1328,7 +1346,9 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
_StoreMem(FPRClass, 16, MemLocation, MMReg, 16);
}
for (unsigned i = 0; i < 16; ++i) {
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *XMMReg = _LoadContext(16, offsetof(FEXCore::Core::CPUState, xmm[i]), FPRClass);
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
@@ -1363,12 +1383,21 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(C3);
}
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto NewFTW = _LoadMem(GPRClass, 2, MemLocation, 2);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
}
for (unsigned i = 0; i < 8; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 32));
auto MMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, mm[i]), MMReg);
}
for (unsigned i = 0; i < 16; ++i) {
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
auto XMMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, xmm[i]), XMMReg);
@@ -7,9 +7,18 @@ $end_info$
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
#include <stddef.h>
#include <stdint.h>
namespace FEXCore::IR {
class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
OrderedNode *OpDispatchBuilder::GetX87Top() {
@@ -134,17 +143,17 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
}
template
void OpDispatchBuilder::FLD_Const<0x8000'0000'0000'0000, 0b0'011'1111'1111'1111>(OpcodeArgs); // 1.0
void OpDispatchBuilder::FLD_Const<0x8000'0000'0000'0000ULL, 0b0'011'1111'1111'1111ULL>(OpcodeArgs); // 1.0
template
void OpDispatchBuilder::FLD_Const<0xD49A'784B'CD1B'8AFE, 0x4000>(OpcodeArgs); // log2l(10)
void OpDispatchBuilder::FLD_Const<0xD49A'784B'CD1B'8AFEULL, 0x4000ULL>(OpcodeArgs); // log2l(10)
template
void OpDispatchBuilder::FLD_Const<0xB8AA'3B29'5C17'F0BC, 0x3FFF>(OpcodeArgs); // log2l(e)
void OpDispatchBuilder::FLD_Const<0xB8AA'3B29'5C17'F0BCULL, 0x3FFFULL>(OpcodeArgs); // log2l(e)
template
void OpDispatchBuilder::FLD_Const<0xC90F'DAA2'2168'C235, 0x4000>(OpcodeArgs); // pi
void OpDispatchBuilder::FLD_Const<0xC90F'DAA2'2168'C235ULL, 0x4000ULL>(OpcodeArgs); // pi
template
void OpDispatchBuilder::FLD_Const<0x9A20'9A84'FBCF'F799, 0x3FFD>(OpcodeArgs); // log10l(2)
void OpDispatchBuilder::FLD_Const<0x9A20'9A84'FBCF'F799ULL, 0x3FFDULL>(OpcodeArgs); // log10l(2)
template
void OpDispatchBuilder::FLD_Const<0xB172'17F7'D1CF'79AC, 0x3FFE>(OpcodeArgs); // log(2)
void OpDispatchBuilder::FLD_Const<0xB172'17F7'D1CF'79ACULL, 0x3FFEULL>(OpcodeArgs); // log(2)
template
void OpDispatchBuilder::FLD_Const<0, 0>(OpcodeArgs); // 0.0
@@ -537,7 +546,7 @@ void OpDispatchBuilder::FCHS(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto low = _Constant(0);
auto high = _Constant(0b1'000'0000'0000'0000);
auto high = _Constant(0b1'000'0000'0000'0000ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
@@ -552,7 +561,7 @@ void OpDispatchBuilder::FABS(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto low = _Constant(~0ULL);
auto high = _Constant(0b0'111'1111'1111'1111);
auto high = _Constant(0b0'111'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
@@ -861,7 +870,7 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
OrderedNode *st1 = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
if (Plus1) {
auto low = _Constant(0x8000'0000'0000'0000);
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
@@ -884,8 +893,8 @@ void OpDispatchBuilder::X87TAN(OpcodeArgs) {
auto result = _F80TAN(a);
auto low = _Constant(0x8000'0000'0000'0000);
auto high = _Constant(0b0'011'1111'1111'1111);
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
@@ -0,0 +1,113 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/LogManager.h>
#include <unistd.h>
#include <signal.h>
namespace FEXCore {
struct ThreadState {
FEXCore::Core::InternalThreadState *Thread{};
};
thread_local ThreadState ThreadData{};
static bool IsSynchronous(int Signal) {
switch (Signal) {
case SIGBUS:
case SIGFPE:
case SIGILL:
case SIGSEGV:
case SIGTRAP:
return true;
default: break;
};
return false;
}
/**
* @brief Masks signals from the signal mask
*
* @param how Argument to sigmask. SIG_{BLOCK, SETMASK, UNBLOCK}
* @param Signal Which signal to set or -1 to sweep through them all
*/
static void MaskSignals(int how, int Signal = -1) {
// If we have a helper thread, we need to mask a significant amount of signals so the an errant thread doesn't receive a signal that it shouldn't
sigset_t SignalSet{};
sigemptyset(&SignalSet);
if (Signal == -1) {
for (int i = 0; i <= SignalDelegator::MAX_SIGNALS; ++i) {
// If it is a synchronous signal then don't ignore it
if (IsSynchronous(i)) {
continue;
}
// Add this signal to the ignore list
sigaddset(&SignalSet, i);
}
}
else {
sigaddset(&SignalSet, Signal);
}
// Be warned, a thread will inherit the signal mask if created from this thread
int Result = pthread_sigmask(how, &SignalSet, nullptr);
if (Result != 0) {
LogMan::Msg::E("Couldn't register thread to mask signals");
}
}
void SignalDelegator::MaskThreadSignals() {
MaskSignals(SIG_BLOCK);
}
FEXCore::Core::InternalThreadState *SignalDelegator::GetTLSThread() {
return ThreadData.Thread;
}
void SignalDelegator::RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) {
ThreadData.Thread = Thread;
RegisterFrontendTLSState(Thread);
}
void SignalDelegator::UninstallTLSState(FEXCore::Core::InternalThreadState *Thread) {
UninstallFrontendTLSState(Thread);
ThreadData.Thread = nullptr;
}
void SignalDelegator::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SetHostSignalHandler(Signal, Func, Required);
FrontendRegisterHostSignalHandler(Signal, Func, Required);
}
void SignalDelegator::RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SetFrontendHostSignalHandler(Signal, Func, Required);
FrontendRegisterFrontendHostSignalHandler(Signal, Func, Required);
}
void SignalDelegator::HandleSignal(int Signal, void *Info, void *UContext) {
// Let the host take first stab at handling the signal
auto Thread = GetTLSThread();
HostSignalHandler &Handler = HostHandlers[Signal];
if (!Thread) {
LogMan::Msg::E("[%d] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", ::gettid());
}
else {
if (Handler.Handler &&
Handler.Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
}
if (Handler.FrontendHandler &&
Handler.FrontendHandler(Thread, Signal, Info, UContext)) {
return;
}
// Now let the frontend handle the signal
// It's clearly a guest signal and this ends up being an OS specific issue
HandleGuestSignal(Thread, Signal, Info, UContext);
}
}
}
+4 -1
View File
@@ -6,12 +6,15 @@ $end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <cstdint>
#include <cstring>
#include <stdlib.h>
#include <vector>
#include <sys/mman.h>
#include <bits/mman-map-flags-generic.h>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
+1 -1
View File
@@ -5,8 +5,8 @@ $end_info$
*/
#pragma once
#include <FEXCore/Config/Config.h>
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
-6
View File
@@ -5,14 +5,8 @@ tags: frontend|x86-tables
$end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/X86Tables.h>
#include <array>
#include <cstdint>
#include <tuple>
#include <vector>
namespace FEXCore::X86Tables {
@@ -7,6 +7,9 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,10 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,10 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -4,9 +4,11 @@ tags: frontend|x86-tables
$end_info$
*/
#include <FEXCore/Core/Context.h>
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
#include <stdint.h>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,12 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Core/Context.h>
#include <iterator>
#include <stdint.h>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,11 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,11 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
#include <stdint.h>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,10 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -4,7 +4,12 @@ tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,9 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,9 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
@@ -6,6 +6,11 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Debug/X86Tables.h>
#include <iterator>
#include <stdint.h>
namespace FEXCore::X86Tables {
using namespace InstFlags;
+16 -13
View File
@@ -5,21 +5,24 @@ tags: glue|thunks
$end_info$
*/
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include "Thunks.h"
#include "stdio.h"
#include <dlfcn.h>
#include <string>
#include <map>
#include <array>
#include <Interface/Context/Context.h>
#include "Interface/Core/InternalThreadState.h"
#include "FEXCore/Core/X86Enums.h"
#include <mutex>
#include <malloc.h>
#include <map>
#include <memory>
#include <shared_mutex>
#include <stdint.h>
#include <string>
#include <utility>
struct LoadlibArgs {
const char *Name;
@@ -30,7 +33,6 @@ static thread_local FEXCore::Core::InternalThreadState *Thread;
namespace FEXCore {
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
class ThunkHandler_impl final: public ThunkHandler {
@@ -48,18 +50,17 @@ namespace FEXCore {
Set arg0/1 to arg regs, use CTX::HandleCallback to handle the callback
*/
static void CallCallback(void *callback, void *arg0, void* arg1) {
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CTX->HandleCallback((uintptr_t)callback);
Thread->CTX->HandleCallback(Thread, (uintptr_t)callback);
}
static void LoadLib(void *ArgsV) {
auto CTX = Thread->CTX;
auto Args = reinterpret_cast<LoadlibArgs*>(ArgsV);
auto CTX = Thread->CTX;
auto Name = Args->Name;
auto CallbackThunks = Args->CallbackThunks;
@@ -121,7 +122,9 @@ namespace FEXCore {
}
ThunkHandler_impl() {
}
~ThunkHandler_impl() {
}
};
+9 -2
View File
@@ -5,12 +5,19 @@ $end_info$
*/
#pragma once
#include <FEXCore/IR/IR.h>
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::IR {
struct SHA256Sum;
}
namespace FEXCore {
typedef void ThunkedFunction(void* ArgsRv);
@@ -22,4 +29,4 @@ namespace FEXCore {
static ThunkHandler* Create();
};
};
};
+8 -1
View File
@@ -605,10 +605,17 @@
"FillRegister": {
"Desc": ["Fills a register from a spill slot",
"Spill slots are register allocated and has live ranges calculated to handle slot calculation",
"```diff\n- !Don't use this op. It is for RA to handle spilling and filling!\n```"
"```diff\n- !Don't use this op. It is for RA to handle spilling and filling!\n```",
"",
"The OriginalValue SSA arg points at the original SSA value spilled, and only exists for",
"RA validation purposes"
],
"OpClass": "Memory",
"SSAArgs": "1",
"SSANames": [
"OriginalValue"
],
"HasDest": true,
"DestClass": "Complex",
"Args": [
+7 -2
View File
@@ -7,9 +7,14 @@ $end_info$
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/RegisterAllocationData.h>
#include <algorithm>
#include <array>
#include <ostream>
#include <stdint.h>
#include <string>
#include <string_view>
#include <iomanip>
namespace FEXCore::IR {
+12 -4
View File
@@ -5,7 +5,15 @@ tags: ir|emitter
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stdint.h>
#include <string.h>
#include <vector>
namespace FEXCore::IR {
void IREmitter::ResetWorkingList() {
@@ -18,12 +26,12 @@ void IREmitter::ResetWorkingList() {
CurrentCodeBlock = nullptr;
}
void IREmitter::ReplaceAllUsesWithRange(OrderedNode *Node, OrderedNode *NewNode, AllNodesIterator After, AllNodesIterator End) {
void IREmitter::ReplaceAllUsesWithRange(OrderedNode *Node, OrderedNode *NewNode, AllNodesIterator Begin, AllNodesIterator End) {
uintptr_t ListBegin = DualListData.ListBegin();
auto NodeId = Node->Wrapped(ListBegin).ID();
while (After != End) {
auto [RealNode, IROp] = After();
while (Begin != End) {
auto [RealNode, IROp] = Begin();
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
@@ -39,7 +47,7 @@ void IREmitter::ReplaceAllUsesWithRange(OrderedNode *Node, OrderedNode *NewNode,
}
}
++After;
++Begin;
}
}
+17 -10
View File
@@ -5,16 +5,24 @@ tags: ir|parser
$end_info$
*/
#include <string>
#include <vector>
#include <istream>
#include <unordered_map>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include <algorithm>
#include <array>
#include <cstdint>
#include <errno.h>
#include <memory>
#include <stdio.h>
#include <stdlib.h>
#include <string>
#include <string_view>
#include <utility>
#include <vector>
#include <istream>
#include <unordered_map>
namespace FEXCore::IR {
namespace {
@@ -70,8 +78,6 @@ std::string DecodeErrorToString(DecodeFailure Failure) {
return "Unknown Error";
}
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
class IRParser: public FEXCore::IR::IREmitter {
public:
template<typename Type>
@@ -299,9 +305,10 @@ class IRParser: public FEXCore::IR::IREmitter {
std::unordered_map<std::string, OrderedNode*> SSANameMapper;
std::vector<LineDefinition> Defs;
LineDefinition *CurrentDef{};
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
IRParser(std::istream *text) {
InitializeStaticTables();
InitializeNameMap();
std::string TmpLine;
while (!text->eof()) {
@@ -630,7 +637,7 @@ class IRParser: public FEXCore::IR::IREmitter {
return true;
}
void InitializeStaticTables() {
void InitializeNameMap() {
if (NameToOpMap.empty()) {
for (FEXCore::IR::IROps Op = FEXCore::IR::IROps::OP_DUMMY;
Op <= FEXCore::IR::IROps::OP_LAST;
+5 -3
View File
@@ -13,6 +13,7 @@ $end_info$
#include <FEXCore/Config/Config.h>
namespace FEXCore::IR {
class IREmitter;
void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllocation) {
FEX_CONFIG_OPT(DisablePasses, O0);
@@ -47,19 +48,20 @@ void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllo
// If the IR is compacted post-RA then the node indexing gets messed up and the backend isn't able to find the register assigned to a node
// Compact before IR, don't worry about RA generating spills/fills
CompactionPass = InsertPass(CreateIRCompaction());
InsertPass(CreateIRCompaction(), "Compaction");
}
void PassManager::AddDefaultValidationPasses() {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
InsertValidationPass(Validation::CreatePhiValidation());
InsertValidationPass(Validation::CreateIRValidation());
InsertValidationPass(Validation::CreateIRValidation(), "IRValidation");
InsertValidationPass(Validation::CreateRAValidation());
InsertValidationPass(Validation::CreateValueDominanceValidation());
#endif
}
void PassManager::InsertRegisterAllocationPass(bool OptimizeSRA) {
RAPass = InsertPass(IR::CreateRegisterAllocationPass(CompactionPass, OptimizeSRA));
InsertPass(IR::CreateRegisterAllocationPass(GetPass("Compaction"), OptimizeSRA), "RA");
}
bool PassManager::Run(IREmitter *IREmit) {
+26 -15
View File
@@ -7,11 +7,10 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
#include <functional>
#include <memory>
#include <utility>
#include <vector>
namespace FEXCore::HLE {
@@ -19,8 +18,8 @@ class SyscallHandler;
}
namespace FEXCore::IR {
class OpDispatchBuilder;
class SyscallOptimization;
class PassManager;
class IREmitter;
using ShouldExitHandler = std::function<void(void)>;
@@ -42,9 +41,14 @@ class PassManager final {
public:
void AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllocation);
void AddDefaultValidationPasses();
Pass* InsertPass(std::unique_ptr<Pass> Pass) {
Pass* InsertPass(std::unique_ptr<Pass> Pass, std::string Name = "") {
Pass->RegisterPassManager(this);
return Passes.emplace_back(std::move(Pass)).get();
auto PassPtr = Passes.emplace_back(std::move(Pass)).get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
return PassPtr;
}
void InsertRegisterAllocationPass(bool OptimizeSRA);
@@ -55,12 +59,17 @@ public:
ExitHandler = std::move(Handler);
}
bool HasRAPass() const {
return RAPass != nullptr;
bool HasPass(std::string Name) const {
return NameToPassMaping.contains(Name);
}
IR::RegisterAllocationPass *GetRAPass() {
return reinterpret_cast<IR::RegisterAllocationPass*>(RAPass);
template<typename T>
T* GetPass(std::string Name) {
return dynamic_cast<T*>(NameToPassMaping[Name]);
}
Pass* GetPass(std::string Name) {
return NameToPassMaping[Name];
}
void RegisterSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) {
@@ -72,16 +81,18 @@ protected:
FEXCore::HLE::SyscallHandler *SyscallHandler;
private:
Pass *RAPass{};
Pass *CompactionPass{};
std::vector<std::unique_ptr<Pass>> Passes;
std::unordered_map<std::string, Pass*> NameToPassMaping;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
std::vector<std::unique_ptr<Pass>> ValidationPasses;
void InsertValidationPass(std::unique_ptr<Pass> Pass) {
void InsertValidationPass(std::unique_ptr<Pass> Pass, std::string Name = "") {
Pass->RegisterPassManager(this);
ValidationPasses.emplace_back(std::move(Pass));
auto PassPtr = ValidationPasses.emplace_back(std::move(Pass)).get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
}
#endif
+1
View File
@@ -20,6 +20,7 @@ std::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass();
namespace Validation {
std::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
std::unique_ptr<FEXCore::IR::Pass> CreateRAValidation();
std::unique_ptr<FEXCore::IR::Pass> CreatePhiValidation();
std::unique_ptr<FEXCore::IR::Pass> CreateValueDominanceValidation();
}
+14 -1
View File
@@ -15,7 +15,20 @@ $end_info$
#endif
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <bit>
#include <cstdint>
#include <map>
#include <memory>
#include <string.h>
#include <tuple>
#include <unordered_map>
#include <utility>
namespace FEXCore::IR {
@@ -5,11 +5,12 @@ $end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <array>
#include <memory>
namespace FEXCore::IR {
@@ -7,9 +7,21 @@ $end_info$
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include <vector>
namespace {
struct ContextMemberClassification {
size_t Offset;
@@ -6,7 +6,17 @@ $end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <unordered_map>
namespace FEXCore::IR {
@@ -9,7 +9,16 @@ $end_info$
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <map>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <algorithm>
#include <memory>
#include <stdint.h>
#include <string.h>
#include <vector>
namespace FEXCore::IR {
+31 -26
View File
@@ -6,33 +6,26 @@ $end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Common/BitSet.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <memory>
#include <stddef.h>
#include <string>
#include <sstream>
namespace {
struct BlockInfo {
bool HasExit;
std::vector<FEXCore::IR::OrderedNode const*> Predecessors;
std::vector<FEXCore::IR::OrderedNode const*> Successors;
};
}
#include <unordered_map>
#include <utility>
#include <vector>
namespace FEXCore::IR::Validation {
class IRValidation final : public FEXCore::IR::Pass {
public:
~IRValidation();
bool Run(IREmitter *IREmit) override;
private:
BitSet<uint64_t> NodeIsLive;
size_t MaxNodes{};
};
IRValidation::~IRValidation() {
NodeIsLive.Free();
@@ -42,12 +35,14 @@ bool IRValidation::Run(IREmitter *IREmit) {
bool HadError = false;
bool HadWarning = false;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, BlockInfo> OffsetToBlockMap;
auto CurrentIR = IREmit->ViewIR();
std::ostringstream Errors;
std::ostringstream Warnings;
auto CurrentIR = IREmit->ViewIR();
OffsetToBlockMap.clear();
EntryBlock = nullptr;
if (CurrentIR.GetSSACount() > MaxNodes) {
NodeIsLive.Realloc(CurrentIR.GetSSACount());
}
@@ -60,8 +55,8 @@ bool IRValidation::Run(IREmitter *IREmit) {
#endif
IR::RegisterAllocationData * RAData{};
if (Manager->HasRAPass()) {
RAData = Manager->GetRAPass() ? Manager->GetRAPass()->GetAllocationData() : nullptr;
if (Manager->HasPass("RA")) {
RAData = Manager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData();
}
NodeIsLive.Set(1); // IRHEADER
@@ -70,10 +65,15 @@ bool IRValidation::Run(IREmitter *IREmit) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
if (!EntryBlock) {
EntryBlock = BlockNode;
}
uint32_t BlockID = CurrentIR.GetID(BlockNode);
BlockInfo *CurrentBlock = &OffsetToBlockMap.try_emplace(BlockID).first->second;
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
uint32_t ID = CurrentIR.GetID(CodeNode);
@@ -281,6 +281,11 @@ bool IRValidation::Run(IREmitter *IREmit) {
}
LogMan::Msg::EFmt("{}", Out.str());
LOGMAN_MSG_A("Encountered IR validation Error");
Errors.clear();
Warnings.clear();
}
return false;
@@ -0,0 +1,32 @@
#pragma once
#include "Common/BitSet.h"
#include <FEXCore/IR/IR.h>
namespace FEXCore::IR::Validation {
struct BlockInfo {
bool HasExit;
OrderedNode const *BlockNode;
std::vector<OrderedNode*> Predecessors;
std::vector<OrderedNode*> Successors;
};
class RAValidation;
class IRValidation final : public FEXCore::IR::Pass {
public:
~IRValidation();
bool Run(IREmitter *IREmit) override;
private:
BitSet<uint64_t> NodeIsLive;
OrderedNode *EntryBlock;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, BlockInfo> OffsetToBlockMap;
size_t MaxNodes{};
friend class RAValidation;
};
}
@@ -6,7 +6,12 @@ $end_info$
*/
#include "Interface/IR/PassManager.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <memory>
#include <stdint.h>
namespace FEXCore::IR {
@@ -5,10 +5,16 @@ desc: Sanity checking pass
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/IR/PassManager.h"
#include <memory>
#include <sstream>
#include <string>
namespace FEXCore::IR::Validation {
@@ -0,0 +1,455 @@
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <algorithm>
#include <deque>
#include <unordered_map>
namespace FEXCore::IR::Validation {
// Hold the mapping of physical registers to the SSA id it holds at any given point in the IR
struct RegState {
static constexpr uint32_t UninitializedValue = 0;
static constexpr uint32_t InvalidReg = 0xffff'ffff;
static constexpr uint32_t CorruptedPair = 0xffff'fffe;
static constexpr uint32_t ClobberedValue = 0xffff'fffd;
static constexpr uint32_t StaticAssigned = 0xffff'ff00;
// This class makes some assumptions about how the host registers are arranged and mapped to virtual registers:
// 1. There will be less than 32 GPRs and 32 FPRs
// 2. If the GPRFixed class is used, there will be 16 GPRs and 16 FixedGPRs max
// 3. Same with FPRFixed
// 4. If the GPRPairClass is used, it is assumed each GPRPair N will map onto GPRs N*2 and N*2 + 1
// These assumptions were all true for the state of the arm64 and x86 jits at the time this was written
// Mark a physical register as containing a SSA id
bool Set(PhysicalRegister Reg, uint32_t ssa) {
LOGMAN_THROW_A(ssa != 0, "RegState assumes ssa0 will be the block header and never assigned to a register");
// PhyscialRegisters aren't fully mapped until assembly emission
// We need to apply a generic mapping here to catch any aliasing
switch (Reg.Class) {
case GPRClass:
GPRs[Reg.Reg] = ssa;
return true;
case GPRFixedClass:
// On arm64, there are 16 Fixed and 9 normal
GPRs[Reg.Reg + 16] = ssa;
return true;
case FPRClass:
FPRs[Reg.Reg] = ssa;
return true;
case FPRFixedClass:
// On arm64, there are 16 Fixed and 12 normal
FPRs[Reg.Reg + 16] = ssa;
return true;
case GPRPairClass:
if (Reg.Reg <= 16) {
// Alias paired registers onto both
GPRs[Reg.Reg*2] = ssa;
GPRs[Reg.Reg*2 + 1] = ssa;
return true;
}
break;
}
return false;
}
// Get the current SSA id
// Or an error value there isn't a (sane) SSA id
uint32_t Get(PhysicalRegister Reg) {
switch (Reg.Class) {
case GPRClass:
return GPRs[Reg.Reg];
case GPRFixedClass:
if (GPRs[Reg.Reg + 16] == UninitializedValue) {
return StaticAssigned;
}
return GPRs[Reg.Reg + 16];
case FPRClass:
return FPRs[Reg.Reg];
case FPRFixedClass:
if (FPRs[Reg.Reg + 16] == UninitializedValue) {
return StaticAssigned;
}
return FPRs[Reg.Reg + 16];
case GPRPairClass:
if (Reg.Reg > 16)
break;
// Make sure both halves of the Pair contain the same SSA
if (GPRs[Reg.Reg*2] == GPRs[Reg.Reg*2 + 1]) {
return GPRs[Reg.Reg*2];
}
return CorruptedPair;
}
return InvalidReg;
}
// Mark a spill slot as containing a SSA id
void Spill(uint32_t SpillSlot, uint32_t ssa) {
Spills[SpillSlot] = ssa;
}
// Consume (and return) the SSA id currently in a spill slot
uint32_t Unspill(uint32_t SpillSlot) {
if (Spills.contains(SpillSlot)) {
uint32_t Value = Spills[SpillSlot];
Spills.erase(SpillSlot);
return Value;
}
return UninitializedValue;
}
// Intersect another regstate with this one
// Any registers/slots which contain the same SSA id will be persevered
// Anything else will be marked as Clobbered
//
// Useful for merging two branches of control flow.
// Any register that differs depending on control flow shouldn't be consumed by
// code that follows
void Intersect(RegState& other) {
for (size_t i = 0; i < GPRs.size(); i++) {
if (GPRs[i] != other.GPRs[i]) {
GPRs[i] = ClobberedValue;
}
}
for (size_t i = 0; i < FPRs.size(); i++) {
if (FPRs[i] != other.FPRs[i]) {
FPRs[i] = ClobberedValue;
}
}
for (auto it = Spills.begin(); it != Spills.end(); it++) {
auto& [SlotID, Value] = *it;
if (!other.Spills.contains(SlotID)) {
Spills.erase(it);
} else if (Value != other.Spills[SlotID]) {
Value = ClobberedValue;
}
}
}
// Filter out all registers/slots containing an SSA id larger than MaxSSA
// Mark them as Clobbered.
// Useful for backwards edges, where using an SSA from before the
void Filter(uint32_t MaxSSA) {
for (auto &gpr : GPRs) {
if (gpr > MaxSSA) {
gpr = ClobberedValue;
}
}
for (auto &fpr : FPRs) {
if (fpr > MaxSSA) {
fpr = ClobberedValue;
}
}
for (auto it = Spills.begin(); it != Spills.end(); it++) {
auto& [SlotID, Value] = *it;
if (Value > MaxSSA) {
Spills.erase(it);
}
}
}
private:
std::array<uint32_t, 32> GPRs = {};
std::array<uint32_t, 32> FPRs = {};
std::unordered_map<uint32_t, uint32_t> Spills;
public:
uint32_t Version{}; // Used to force regeneration of RegStates after following backward edges
};
class RAValidation final : public FEXCore::IR::Pass {
public:
~RAValidation() {}
bool Run(IREmitter *IREmit) override;
private:
// Holds the calculated RegState at the exit of each block
std::unordered_map<uint32_t, RegState> BlockExitState;
// A queue of blocks we need to visit (or revisit)
std::deque<OrderedNode*> BlocksToVisit;
};
bool RAValidation::Run(IREmitter *IREmit) {
if (!Manager->HasPass("RA")) return false;
IR::RegisterAllocationData* RAData = Manager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData();
BlockExitState.clear();
// BlocksToVisit will already be empty
// Get the control flow graph from the validation pass
auto ValidationPass = Manager->GetPass<IRValidation>("IRValidation");
LOGMAN_THROW_A(ValidationPass != nullptr, "Couldn't find IRValidation pass");
auto& OffsetToBlockMap = ValidationPass->OffsetToBlockMap;
LOGMAN_THROW_A(ValidationPass->EntryBlock != nullptr, "No entry point");
BlocksToVisit.push_front(ValidationPass->EntryBlock); // Currently only a single entry point
bool HadError = false;
std::ostringstream Errors;
auto CurrentIR = IREmit->ViewIR();
uint32_t CurrentVersion = 1; // Incremented every backwards edge
while (!BlocksToVisit.empty())
{
auto BlockNode = BlocksToVisit.front();
uint32_t BlockID = CurrentIR.GetID(BlockNode);
auto& BlockInfo = OffsetToBlockMap[BlockID];
auto IsFowardsEdge = [&] (uint32_t PredecessorID) {
// Blocks are sorted in FEXes IR, so backwards edges always go to a lower (or equal) Block ID
return PredecessorID < BlockID;
};
// First, make sure we have the exit state for all Predecessors that
// get here via a forwards branch.
bool MissingPredecessor = false;
for (auto Predecessor : BlockInfo.Predecessors) {
auto PredecessorID = CurrentIR.GetID(Predecessor);
bool HaveState = BlockExitState.contains(PredecessorID) && BlockExitState[PredecessorID].Version == CurrentVersion;
if (IsFowardsEdge(PredecessorID) && !HaveState) {
// We are probably about to visit this node anyway, remove it
std::remove(BlocksToVisit.begin(), BlocksToVisit.end(), Predecessor);
// Add the missing predecessor to start of queue
BlocksToVisit.push_front(Predecessor);
MissingPredecessor = true;
}
}
if (MissingPredecessor) {
// We'll have to come back to this block later
continue;
}
// We have committed to processing this block
// Remove from queue
BlocksToVisit.pop_front();
bool FirstVisit = !BlockExitState.contains(BlockID);
// Second, we need to determine the register status as of Block entry
auto BlockOp = CurrentIR.GetOp<IROp_CodeBlock>(BlockNode);
uint32_t FirstSSA = BlockOp->Begin.ID();
auto& BlockRegState = BlockExitState.try_emplace(BlockID).first->second;
bool EmptyRegState = true;
auto Intersect = [&] (RegState& Other) {
if (EmptyRegState) {
BlockRegState = Other;
EmptyRegState = false;
} else {
BlockRegState.Intersect(Other);
}
};
for (auto Predecessor : BlockInfo.Predecessors) {
auto PredecessorID = CurrentIR.GetID(Predecessor);
if (BlockExitState.contains(PredecessorID)) {
if (IsFowardsEdge(PredecessorID)) {
Intersect(BlockExitState[PredecessorID]);
} else {
RegState Filtered = BlockExitState[PredecessorID];
Filtered.Filter(FirstSSA);
Intersect(Filtered);
}
}
}
// Thrid, we need to iterate over all IR ops in the block
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
uint32_t ID = CurrentIR.GetID(CodeNode);
auto CheckArg = [&] (uint32_t i, OrderedNodeWrapper Arg) {
const auto PhyReg = RAData->GetNodeRegister(Arg.ID());
if (PhyReg.IsInvalid())
return;
auto CurrentSSAAtReg = BlockRegState.Get(PhyReg);
if (CurrentSSAAtReg == RegState::InvalidReg) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] unknown Reg: {}, class: {}\n", ID, i, PhyReg.Reg, PhyReg.Class);
} else if (CurrentSSAAtReg == RegState::CorruptedPair) {
HadError |= true;
auto Lower = BlockRegState.Get(PhysicalRegister(GPRClass, uint8_t(PhyReg.Reg*2) + 1));
auto Upper = BlockRegState.Get(PhysicalRegister(GPRClass, PhyReg.Reg*2 + 1));
Errors << fmt::format("%ssa{}: Arg[{}] expects paired reg{} to contain %ssa{}, but it actually contains {{%ssa{}, %ssa{}}}\n",
ID, i, PhyReg.Reg, Arg.ID(), Lower, Upper);
} else if (CurrentSSAAtReg == RegState::UninitializedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but it is uninitialized\n",
ID, i, PhyReg.Reg, Arg.ID());
} else if (CurrentSSAAtReg == RegState::ClobberedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but contents vary depending on control flow\n",
ID, i, PhyReg.Reg, Arg.ID());
} else if (CurrentSSAAtReg != Arg.ID()) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but it actually contains %ssa{}\n",
ID, i, PhyReg.Reg, Arg.ID(), CurrentSSAAtReg);
}
};
switch (IROp->Op)
{
case OP_SPILLREGISTER: {
auto SpillRegister = IROp->C<IROp_SpillRegister>();
CheckArg(0, SpillRegister->Value);
BlockRegState.Spill(SpillRegister->Slot, SpillRegister->Value.ID());
break;
}
case OP_FILLREGISTER: {
auto FillRegister = IROp->C<IROp_FillRegister>();
uint32_t ExpectedValue = FillRegister->OriginalValue.ID();
uint32_t Value = BlockRegState.Unspill(FillRegister->Slot);
// TODO: This only proves that the Spill has a consistent SSA value
// In the future we need to prove it contains the correct SSA value
if (Value == RegState::UninitializedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: FillRegister expected %ssa{} in Slot {}, but was undefined in at least one control flow path\n",
ID, ExpectedValue, FillRegister->Slot);
} else if (Value == RegState::ClobberedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: FillRegister expected %ssa{} in Slot {}, but contents vary depending on control flow\n",
ID, ExpectedValue, FillRegister->Slot);
} else if (Value != ExpectedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: FillRegister expected %ssa{} in Slot {}, but it actually contains %ssa{}\n",
ID, ExpectedValue, FillRegister->Slot, Value);
}
break;
}
default: {
// And check that all args point at the correct SSA
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint32_t i = 0; i < NumArgs; ++i) {
CheckArg(i, IROp->Args[i]);
}
break;
}
}
// Update BlockState map
BlockRegState.Set(RAData->GetNodeRegister(ID), ID);
}
// Forth, Add successors to the queue of blocks to validate
for (auto Successor : BlockInfo.Successors) {
auto SuccessorID = CurrentIR.GetID(Successor);
// Blocks are sorted in FEXes IR, so backwards edges always go to a lower (or equal) Block ID
bool FowardsEdge = SuccessorID > BlockID;
if (FowardsEdge) {
// Always follow forwards edges, assuming it's not already on the queue
if (std::find(BlocksToVisit.begin(), BlocksToVisit.end(), Successor) == std::end(BlocksToVisit)) {
// Push to the back of queue so there is a higher chance all predecessors for this block are done first
BlocksToVisit.push_back(Successor);
}
} else if (FirstVisit) {
// Now that we have the block data for the backwards edge, we can visit it again and make
// sure it (and all it's successors) are still valid.
// But only do this the first time we encounter each backwards edge.
// Push to the front of queue, so we get this re-checking done before examining future nodes.
BlocksToVisit.push_front(Successor);
// Make sure states are reprocessed
CurrentVersion++;
}
}
BlockRegState.Version = CurrentVersion;
if (CurrentVersion > 10000) {
Errors << "Infinite Loop\n";
HadError |= true;
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
uint32_t BlockID = CurrentIR.GetID(BlockNode);
auto& BlockInfo = OffsetToBlockMap[BlockID];
Errors << fmt::format("Block {}\n\tPredecessors: ", BlockID);
for (auto Predecessor : BlockInfo.Predecessors) {
auto PredecessorID = CurrentIR.GetID(Predecessor);
bool FowardsEdge = PredecessorID < BlockID;
if (!FowardsEdge) {
Errors << "(Backwards): ";
}
Errors << fmt::format("Block {} ", PredecessorID);
}
Errors << "\n\tSuccessors: ";
for (auto Successor : BlockInfo.Successors) {
auto SuccessorID = CurrentIR.GetID(Successor);
bool FowardsEdge = SuccessorID > BlockID;
if (!FowardsEdge) {
Errors << "(Backwards): ";
}
Errors << fmt::format("Block {} ", SuccessorID);
}
Errors << "\n\n";
}
break;
}
}
if (HadError) {
std::stringstream IrDump;
FEXCore::IR::Dump(&IrDump, &CurrentIR, RAData);
LogMan::Msg::EFmt("RA Validation Error\n{}\nErrors:\n{}\n", IrDump.str(), Errors.str());
LOGMAN_MSG_A("Encountered RA validation Error");
Errors.clear();
}
return false;
}
std::unique_ptr<FEXCore::IR::Pass> CreateRAValidation() {
return std::make_unique<RAValidation>();
}
}
@@ -5,8 +5,13 @@ desc: This is not used right now, possibly broken
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <array>
#include <memory>
namespace FEXCore::IR {
@@ -4,15 +4,28 @@ tags: ir|opts
$end_info$
*/
#include "Common/BitSet.h"
#include <FEXCore/Utils/BucketList.h>
#include "Common/MathUtils.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/LogManager.h>
#include <iterator>
#include <algorithm>
#include <cstdint>
#include <set>
#include <stddef.h>
#include <string.h>
#include <strings.h>
#include <unordered_map>
#include <unordered_set>
#include <sys/mman.h>
#include <utility>
#include <vector>
#define SRA_DEBUG(...) // printf(__VA_ARGS__)
@@ -26,137 +39,6 @@ namespace {
constexpr uint32_t DEFAULT_INTERFERENCE_SPAN_COUNT = 30;
constexpr uint32_t DEFAULT_NODE_COUNT = 8192;
const PhysicalRegister INVALID_REGCLASS = PhysicalRegister::Invalid();
// BucketList is an optimized container, it includes an inline array of Size
// and can overflow to a linked list of further buckets
//
// To optimize for best performance, Size should be big enough to allocate one or two
// buckets for the typical case
// Picking a Size so sizeof(Bucket<...>) is a power of two is also a small win
template<unsigned _Size, typename T = uint32_t>
struct BucketList {
static constexpr const unsigned Size = _Size;
T Items[Size];
std::unique_ptr<BucketList<Size>> Next;
void Clear() {
Items[0] = 0;
#ifndef NDEBUG
for (int i = 1; i < Size; i++)
Items[i] = 0xDEADBEEF;
#endif
Next.reset();
}
BucketList() {
Clear();
}
template<typename EnumeratorFn>
inline void Iterate(EnumeratorFn Enumerator) const {
int i = 0;
auto Bucket = this;
for(;;) {
auto Item = Bucket->Items[i];
if (Item == 0)
break;
Enumerator(Item);
if (++i == Bucket->Size) {
LOGMAN_THROW_A(Bucket->Next != nullptr, "Interference bug");
Bucket = Bucket->Next.get();
i = 0;
}
}
}
template<typename EnumeratorFn>
inline bool Find(EnumeratorFn Enumerator) const {
int i = 0;
auto Bucket = this;
for(;;) {
auto Item = Bucket->Items[i];
if (Item == 0)
break;
if (Enumerator(Item))
return true;
if (++i == Bucket->Size) {
LOGMAN_THROW_A(Bucket->Next != nullptr, "Bucket in bad state");
Bucket = Bucket->Next.get();
i = 0;
}
}
return false;
}
void Append(uint32_t Val) {
auto that = this;
while (that->Next) {
that = that->Next.get();
}
int i;
for (i = 0; i < Size; i++) {
if (that->Items[i] == 0) {
that->Items[i] = Val;
break;
}
}
if (i < (Size-1)) {
that->Items[i+1] = 0;
} else {
that->Next = std::make_unique<BucketList<Size, T>>();
}
}
void Erase(uint32_t Val) {
int i = 0;
auto that = this;
auto foundThat = this;
auto foundI = 0;
for (;;) {
if (that->Items[i] == Val) {
foundThat = that;
foundI = i;
break;
}
else if (++i == Size) {
i = 0;
LOGMAN_THROW_A(that->Next != nullptr, "Bucket::Erase but element not contained");
that = that->Next.get();
}
}
for (;;) {
if (that->Items[i] == 0) {
foundThat->Items[foundI] = that->Items[i-1];
that->Items[i-1] = 0;
break;
}
else if (++i == Size) {
if (that->Next->Items[0] == 0) {
that->Next.reset();
foundThat->Items[foundI] = that->Items[Size-1];
that->Items[Size-1] = 0;
break;
}
i = 0;
that = that->Next.get();
}
}
}
};
struct Register {
bool Virtual;
uint64_t Index;
@@ -174,7 +56,7 @@ namespace {
RegisterNode *PhiPartner;
} Head { ~0U, ~0U, nullptr };
BucketList<DEFAULT_INTERFERENCE_LIST_COUNT, uint32_t> Interferences;
FEXCore::BucketList<DEFAULT_INTERFERENCE_LIST_COUNT, uint32_t> Interferences;
};
static_assert(sizeof(RegisterNode) == 128 * 4);
@@ -191,7 +73,7 @@ namespace {
uint32_t End{~0U};
uint32_t RematCost{0};
uint32_t PreWritten{0};
PhysicalRegister PrefferedRegister{INVALID_REGCLASS};
PhysicalRegister PrefferedRegister{PhysicalRegister::Invalid()};
bool Written{false};
bool Global{false};
};
@@ -269,7 +151,7 @@ namespace {
Graph->VisitedNodePredecessors.clear();
Graph->AllocData.reset((FEXCore::IR::RegisterAllocationData*)FEXCore::Allocator::malloc(FEXCore::IR::RegisterAllocationData::Size(NodeCount)));
memset(&Graph->AllocData->Map[0], INVALID_REGCLASS.Raw, NodeCount);
memset(&Graph->AllocData->Map[0], PhysicalRegister::Invalid().Raw, NodeCount);
Graph->AllocData->MapCount = NodeCount;
Graph->AllocData->IsShared = false; // not shared by default
Graph->NodeCount = NodeCount;
@@ -545,6 +427,10 @@ namespace FEXCore::IR {
// Set this node's block ID
Graph->Nodes[Node].Head.BlockID = BlockNodeID;
// FillRegister's SSA arg is only there for verification, and we don't want it
// to impact the live range.
if (IROp->Op == OP_FILLREGISTER) continue;
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
if (IROp->Args[i].IsInvalid()) continue;
@@ -639,7 +525,7 @@ namespace FEXCore::IR {
return PhysicalRegister(FPRFixedClass, reg);
} else {
LOGMAN_THROW_A(false, "Unexpected Offset %d", Offset);
return INVALID_REGCLASS;
return PhysicalRegister::Invalid();
}
};
@@ -722,7 +608,7 @@ namespace FEXCore::IR {
// ACCESSED after write, let's not SRA this one
if (LiveRanges[ArgNode].Written) {
SRA_DEBUG("Demoting ssa%d because accessed after write in ssa%d\n", ArgNode, Node);
LiveRanges[ArgNode].PrefferedRegister = INVALID_REGCLASS;
LiveRanges[ArgNode].PrefferedRegister = PhysicalRegister::Invalid();
auto ArgNodeNode = IR->GetNode(IROp->Args[i]);
SetNodeClass(Graph, ArgNode, GetRegClassFromNode(IR, ArgNodeNode->Op(IR->GetData())));
}
@@ -756,7 +642,7 @@ namespace FEXCore::IR {
uint32_t ID = (*StaticMap) - &LiveRanges[0];
SRA_DEBUG("ssa%d cannot be a pre-write because ssa%d reads from sra%d before storereg", ID, Node, -1 /*vreg*/);
(*StaticMap)->PrefferedRegister = INVALID_REGCLASS;
(*StaticMap)->PrefferedRegister = PhysicalRegister::Invalid();
(*StaticMap)->PreWritten = 0;
SetNodeClass(Graph, ID, Op->Class);
}
@@ -912,6 +798,9 @@ namespace FEXCore::IR {
return (uint32_t)PhyReg.Class;
};
// SpanStart/SpanEnd assume SSA id will fit in 24bits
LOGMAN_THROW_A(NodeCount <= 0xff'ffff, "Block too large for Spans");
SpanStart.resize(NodeCount);
SpanEnd.resize(NodeCount);
for (uint32_t i = 0; i < NodeCount; ++i) {
@@ -953,13 +842,13 @@ namespace FEXCore::IR {
for (uint32_t i = 0; i < Graph->NodeCount; ++i) {
RegisterNode *CurrentNode = &Graph->Nodes[i];
auto &CurrentRegAndClass = Graph->AllocData->Map[i];
if (CurrentRegAndClass == INVALID_REGCLASS)
if (CurrentRegAndClass == PhysicalRegister::Invalid())
continue;
auto LiveRange = &LiveRanges[i];
FEXCore::IR::RegisterClassType RegClass = FEXCore::IR::RegisterClassType{CurrentRegAndClass.Class};
auto RegAndClass = INVALID_REGCLASS;
auto RegAndClass = PhysicalRegister::Invalid();
RegisterClass *RAClass = &Graph->Set.Classes[RegClass];
if (CurrentNode->Head.PhiPartner) {
@@ -1451,10 +1340,10 @@ namespace FEXCore::IR {
IREmit->SetWriteCursor(FirstUseOrderedNode);
auto FilledInterference = IREmit->_FillRegister(SpillSlot, InterferenceRegClass);
auto FilledInterference = IREmit->_FillRegister(InterferenceOrderedNode, SpillSlot, InterferenceRegClass);
FilledInterference.first->Header.Size = InterferenceIROp->Size;
FilledInterference.first->Header.ElementSize = InterferenceIROp->ElementSize;
IREmit->ReplaceUsesWithAfter(InterferenceOrderedNode, FilledInterference, FirstUseLocation);
IREmit->ReplaceUsesWithAfter(InterferenceOrderedNode, FilledInterference, FilledInterference);
Spilled = true;
}
}
@@ -6,11 +6,14 @@ $end_info$
#pragma once
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/RegisterAllocationData.h>
#include <vector>
#include <memory>
#include <stdint.h>
namespace FEXCore::IR {
class IRListView;
class RegisterAllocationData;
struct RegisterAllocationDataDeleter;
struct RegisterClassType;
class RegisterAllocationPass : public FEXCore::IR::Pass {
public:
@@ -6,7 +6,15 @@ $end_info$
*/
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
namespace FEXCore::IR {
@@ -5,12 +5,15 @@ desc: Removes unused arguments if known syscall number
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/LogManager.h>
#include <memory>
#include <stdint.h>
namespace FEXCore::IR {
Loaded 100 of 317 files, more files were not shown because too many files have changed in this diff. Show more