Compare commits

...
724 Commits
Author SHA1 Message Date
Ryan Houdek 448c4ec797 Docs: Update for release FEX-2106 2021-06-10 12:59:56 -07:00
Ryan Houdek 08fad8fb0d Merge pull request #1067 from Sonicadvance1/fix_32bit_crash
Fixes 32-bit applications crashing
2021-06-10 12:57:21 -07:00
Ryan Houdek 29083a0b41 Makes sure the TestHarnessRunner also hits the shutdown path 2021-06-10 12:47:05 -07:00
Ryan Houdek a1f4ca873c Have FEXLoader use the new freeing functions
It now calls the new functions for handling a cleaner shutdown
2021-06-10 12:47:05 -07:00
Ryan Houdek c3b0e4820b Allocates the signal handlers altstack using FEX's allocator
Using malloc and free was causing a bad crash
2021-06-10 12:47:05 -07:00
Ryan Houdek bb96955e05 Switches from FILE to raw int for Logging message
Due to our current allocation strategy. This variable was ending up in a weird state
where jemalloc allocated it using the glibc allocator.
On shutdown this was causing it to try and deallocate through FEX's allocator...Which
is very broken.

This lets linux clean up in this case, at least until the allocator lines are more
strongly written
2021-06-10 12:47:05 -07:00
Ryan Houdek 16bb64aaab Merge pull request #1063 from Sonicadvance1/fix_crash
Fixes crash in 32-bit pselect6 and pselect6_time64
2021-06-10 12:41:08 -07:00
Ryan Houdek b9e53b6c46 Merge pull request #1069 from lioncash/unex
Syscalls/Thread: Replace std::unexpected with std::terminate
2021-06-09 20:11:58 -07:00
Lioncash 616aa46c6b Syscalls/Thread: Replace std::unexpected with std::terminate
std::unexpected was deprecated in C++11 and removed from the standard in
C++17. The default unexpected_handler calls std::terminate, so this is
identical behavior.
2021-06-09 22:59:21 -04:00
Ryan Houdek c98fb7ae45 Merge pull request #1068 from lioncash/unique
Core: return unique_ptr by default for cores
2021-06-09 19:19:22 -07:00
Lioncash 6c5abc8819 Core: Make CustomCPUFactory instances return unique_ptr
Makes the ownership intentions of the function explicit.
2021-06-09 21:53:07 -04:00
Lioncash f36726ecf6 Core: return unique_ptr by default for cores
Communicates ownership semantics in the API.
2021-06-09 21:41:43 -04:00
Ryan Houdek 6d20ae5dc6 Merge pull request #1062 from lioncash/parser
IRParser: Minor cleanup
2021-06-09 18:34:27 -07:00
Ryan Houdek e469031c40 Merge pull request #1064 from lioncash/syscall
x32/FD: Construct vectors in place
2021-06-09 18:26:08 -07:00
Ryan Houdek b9a8c383fd Merge pull request #1065 from lioncash/poll
{x32, x64}/EPoll: Prevent edge-case out-of-bounds access scenarios
2021-06-09 18:19:25 -07:00
Ryan Houdek 57df4c6d61 Merge pull request #1066 from lioncash/signal
SignalDelegator: Minor changes
2021-06-09 18:17:55 -07:00
Ryan Houdek c5cfb45f48 In ELFCodeLoader2, changes Sections variable over to a unique_ptr
This class is partially allocated on either side of the allocator fence.
Moves `Sections` over to a unique_ptr so we can clear this variable prior to shutdown.
The rest of the class is allocated prior to switching to FEX allocator
2021-06-09 18:05:57 -07:00
Ryan Houdek b6417153d7 Hook more glibc hooks and allow clearing hooks in Allocator
Due to a disjoint mechanism inside of glibc we need to override both the glibc
publicly visible allocation functions AND the hooks.

We had broken the overriding when we changed the default visibility. So first fix that.
Then only override to FEX's allocators once we are ready for it.
Then also replace the hooks.

Additionally have a way to clear the hooks back to default.
2021-06-09 18:03:27 -07:00
Ryan Houdek 599b579192 Adds a ShutdownStaticTables context function
Due to a mixture of allocators this needs to be shutdown in the correct location.
A little bit dirty but is necessary without a full refactor
2021-06-09 18:00:54 -07:00
Lioncash 9ed0572fa4 IRParser: Make use of unique_ptr for Parse()
Communicates ownership semantics to the user of the API a little better,
and makes it harder to accidentally leak memory.
2021-06-09 21:00:30 -04:00
Lioncash 00b0222129 IRParser: Make use of find character overloads
Results in a tiny bit better codegen. May as well, since the changes are
essentially 'free'.
2021-06-09 21:00:30 -04:00
Lioncash 32cc98856f IRParser: Add missing default return in DecodeErrorToString()
Prevents undefined behavior in the event anything actually reaches this.
2021-06-09 21:00:30 -04:00
Lioncash dee7441795 IRParser: Make input string to DecodeValue a const reference
No specializations modify this, and even if they did, it would be a
little confusing to statefully modify the string this way.

Instead, we can make it read-only.
2021-06-09 21:00:29 -04:00
Lioncash f6a6f2eac8 x32/FD: Construct vectors in place
The syscall is defined as taking in an array, so we can construct the
std::vector in place and surround it.

This also fixes an edge case where an out of bounds access on the
Host_iovec vectors could occur if these syscalls were called with an
iovcnt of zero (.at would cause an exception to be thrown).

Looking into the syscalls for readv and writev, these just return early
if a size of zero is passed in.
2021-06-09 21:00:01 -04:00
Lioncash c10e139964 x64/EPoll: Prevent edge-case out-of-bounds access scenarios
Similarly to #1064, an application could potentially pass in bad values
that could result in assertions/exceptions being thrown, so we can
guard against that.

Despite the manual saying that applications *must* not pass
maxevents less than or equal to zero, against all odds, this is
still a handled case in the kernel, so some application out there
may rely on this behavior.
2021-06-09 20:59:44 -04:00
Lioncash 379b200d05 x32/EPoll: Prevent edge-case out-of-bounds access scenarios
Similarly to #1064, an application could potentially pass in bad values
that could result in assertions/exceptions being thrown, so we can guard
against that.

Despite the manual saying that applications *must* not pass maxevents
less than or equal to zero, against all odds, this is still a handled
case in the kernel, so some application out there may rely on this behavior.
2021-06-09 20:59:44 -04:00
Lioncash b4cde28137 SignalDelegator: Move virtual destructor to base class
Ensures that destruction behavior will always handle the polymorphic
case, no matter where it occurs in the hierarchy.
2021-06-09 20:59:21 -04:00
Lioncash 3f88150581 SignalDelegator: Make std::vector into a constexpr array
Minor change, but gets rid of some heap usage.
2021-06-09 20:59:20 -04:00
Lioncash e95076f321 SignalDelegator: Migrate from std::unexpected() to std::terminate()
std::unexpected() was deprecated in C++11 and removed from the standard
in C++17. The default unexpected_handler would call std::terminate
anyway, and given we don't explicitly set a termination handler (as far
as I can tell), this retains identical behavior.
2021-06-09 20:59:20 -04:00
Lioncash e2a16c53b6 SignalDelegator: std::move std::function instances
std::function is allowed to allocate if the size of captures exceeds its
internal storage buffer. It's unlikely this is regularly going to be the
case, but we can allow for avoiding it where necessary.
2021-06-09 20:59:20 -04:00
Lioncash ee89b7a794 SignalDelegator: Leverage std::array
Allows for better detection of out of bounds accesses, as library
implementations generally allow conditional enabling of bounds checks
through preprocessor defines.
2021-06-09 20:59:20 -04:00
Ryan Houdek 5ac9b45528 Merge pull request #1059 from Sonicadvance1/fix_rounding
Fixes floating point conversion rounding bugs
2021-06-09 17:58:35 -07:00
Ryan Houdek 6a94011b00 Fixes crash in 32-bit pselect6 and pselect6_time64
The kernel can be provided a sigmaskpack pointer without a sigset inside of it.

This was causing a crash in some random applications
2021-06-08 21:20:16 -07:00
Ryan Houdek 104fabc11f Merge pull request #1061 from lioncash/dump
IRParser: Make use of fmt where applicable
2021-06-08 07:10:10 -07:00
Ryan Houdek 15b2eb48e5 Disable rounding tests that only fail on Nvidia Xavier
Spooky rounding behaviour changes.
2021-06-08 06:54:11 -07:00
Ryan Houdek 62c446a5f9 Adds more unit tests for conversion operations 2021-06-08 06:54:11 -07:00
Ryan Houdek d57a034220 Fixes floating point conversion rounding bugs
When floats and doubles were converting to integers we weren't doing the correct transformation.
AArch64 provides direct ops for all four of the op types. So lets use them
  - f32 -> int64
  - f32 -> int32
  - f64 -> int64
  - f64 -> int32

Doesn't fully fix the case of overflow for AArch64 since overflow behaviour is different.
x86 returns 0x8000'0000 or 0x8000'0000'0000'0000 while AArch64 saturates to the maximum signed
integers. Can't work around that without checking overflow flags.

This does solve the typical case though.
Fixes audio problems in all FMod games.
2021-06-08 06:54:10 -07:00
Ryan Houdek 18fe049976 Extends a couple tests to ensure correct zext 2021-06-08 06:54:08 -07:00
Stefanos Kornilios Mitsis Poiitidis c6e09ddec5 Merge pull request #1046 from Sonicadvance1/more_sse41
Implements SSE4.1
2021-06-08 16:52:12 +03:00
Lioncash 05ee4c363b IRParser: Make use of fmt where applicable 2021-06-08 09:47:13 -04:00
Lioncash 6f2f4a9bc0 IRDumper: Make use of std::string_view over std::string where applicable
Same behavior, but with a smaller footprint (and the ability to be
constant data).
2021-06-08 09:36:59 -04:00
Ryan Houdek f699726409 Merge pull request #1060 from lioncash/byte
Utils: Add header for bit-related utilities
2021-06-08 05:40:05 -07:00
Lioncash a2b8fedba5 Utils: Add header for bit-related utilities
Places a layer of separation around the remaining compiler builtins.
2021-06-08 08:25:32 -04:00
Ryan Houdek ed2336c4f2 Disables some new sse4.1 gcc tests
The SSE 4.1 path of these tests seemingly work, it's the setup code before the
test that is broken.

We will need to find out why these are failing later
2021-06-08 02:44:40 -07:00
Ryan Houdek 748c3182a7 Fixes a typo in interpreter LSHL
This was causing uint32_t pointers to sign extend on 32-bit
New tests being run were hitting this when we were expecting them to zero extend
2021-06-08 02:44:40 -07:00
Ryan Houdek dfc454873c Moves AOTIR cache queue object to be inside the context
This fixes a crash that occurs due to mixing memory allocators.
This PR has tickled the allocation just enough that it broke
2021-06-08 02:44:40 -07:00
Ryan Houdek 1b06a06a4e Makes RoundType a unique type
This allows us to have the IRDumper have unique output for the type
2021-06-08 02:44:40 -07:00
Ryan Houdek b2ad8d73ac Adds unit test for FEX frontend decoder
This could potentially show up as being ch instead of edi.
Ensure we don't regress behaviour in the future
2021-06-08 02:44:39 -07:00
Ryan Houdek fec078f924 Fixes bug in interpreter VInsGPR
If the source GPR had data that was larger than the element it would overwrite other elements
Mask it correctly. This then matches behaviour with the other CPU backends
2021-06-08 02:44:39 -07:00
Ryan Houdek 38f1cded2c Implements unit tests for all the SSE4.1 instructions 2021-06-08 02:44:39 -07:00
Ryan Houdek a313253510 Enables SSE 4.1 in CPUID 2021-06-08 02:44:39 -07:00
Ryan Houdek 7b2b2a3d07 Implements all the remaining SSE4.1 instructions
Theres a fair number of these so I won't describe them all.

A couple highlights are MPSADBW and PHMINPOSUW.
These don't really match with AArch64 very well so their IR is a bit ugly.
2021-06-08 02:44:39 -07:00
Ryan Houdek 910a25b0ed Cleans up some vector ops that don't need dedicated functions
All of these ops can use the VectorALUOp and VectorUnaryOp templated functions
2021-06-08 02:44:39 -07:00
Ryan Houdek 75bae81331 Describe the remaining SSE4.1 ops in the tables 2021-06-08 02:44:39 -07:00
Ryan Houdek 02abbd216a Fixes an issue with debug printing vectors
We need to save the vector registers otherwise we will corrupt them.
Also in the case of printing a vector register, fall down the specialized path

Only useful when debugging
2021-06-08 02:44:39 -07:00
Ryan Houdek 8b7a7e91c9 Implements new IR ops
Adds VBic, VUMinV, VPopcount, VUnZip, VUnZip2, VDupElement, Vector_FToI, and VUABDL

We will need these for the SSE4.1 ops
2021-06-08 02:44:39 -07:00
Ryan Houdek 130f82a310 Merge pull request #1058 from lioncash/builtins
General: Place more compiler specifics into CompilerDefs.h
2021-06-07 05:31:07 -07:00
Lioncash 18e0f2636f General: Make use of the <bit> header where applicable
Since C++20, a bunch of bit manipulation functions finally have a common
interface, so lets make use of those
2021-06-07 08:16:51 -04:00
Lioncash 2b9029623e General: Abstract trapping behind a define
Provides a layer of separation from direct usages of compiler builtins.
2021-06-07 06:18:44 -04:00
Lioncash 6798afe91f General: Abstract unreachable behind a define
Provides a layer of separation from direct use of compiler built-ins.
2021-06-07 06:13:26 -04:00
Ryan Houdek 4fa8522d9e Merge pull request #1057 from lioncash/cprop
ConstProp: Separate out constituent chunks of const prop
2021-06-05 14:08:40 -07:00
Lioncash 58272ffc47 ConstProp: Separate out constituent chunks of const prop
Makes it nicer to find where each part of the pass is, and also see how
the pass is operating at a high level.
2021-06-05 16:52:03 -04:00
Ryan Houdek 4c275ecc08 Merge pull request #1056 from lioncash/compiler-defs
Utils: Add CompilerDefs.h for compiler-specifics
2021-06-05 12:57:51 -07:00
Ryan Houdek 294ec80aaf Merge pull request #1055 from lioncash/enum-cls
InternalThreadState: Make SignalEvent an enum class
2021-06-05 12:55:10 -07:00
Ryan Houdek e78d3d1a80 Merge pull request #1054 from lioncash/fmt-1
GdbServer: Migrate logging/string handling to fmt where applicable
2021-06-05 12:54:11 -07:00
Lioncash 262389f4ea ConstProp: Mark internal functions as static
Allows them to have internal linkage.
2021-06-05 15:25:35 -04:00
Lioncash 341b811642 Utils: Add CompilerDefs.h for compiler-specifics
Puts a layer of separation around compiler specifics (and also makes
them nicer to write).
2021-06-05 13:55:54 -04:00
Lioncash 32744d074e InternalThreadState: Make SignalEvent an enum class
Makes the enumeration strongly typed, preventing implicit conversions,
minimizing the potential chances of an invalid value being used by
accident.
2021-06-05 10:52:56 -04:00
Lioncash 5f91bbe28d GdbServer: Migrate code to fmt
Simplifies a bunch of string manipulation code.

Much nicer to grok than stream formatting in many cases.
2021-06-05 10:34:59 -04:00
Ryan Houdek 69035348f5 Merge pull request #1053 from lioncash/alias
CodeLoader: Add type aliases for mapper and unmapper functions
2021-06-05 06:16:27 -07:00
Lioncash be46db10e4 CodeLoader: Add virtual destructor
Ensures that no matter the context the hierarchy tree is used
polymorphically, that the deallocation will always be well-defined.

Gets rid of a potential bug vector.
2021-06-05 08:54:09 -04:00
Lioncash aa2c18d8cc CodeLoader: Add type aliases for mapper and unmapper functions
Centralizes the long types in one place for less reading.
2021-06-05 08:51:40 -04:00
Ryan Houdek 0a7f0a3441 Merge pull request #1052 from lioncash/fmtimpl
LogManager: Add fmt-capable logging functions
2021-06-05 05:46:56 -07:00
Lioncash ca2b04b309 LogManager: Add fmt-capable logging functions
Addresses #146 a little more by providing an interface to perform
fmt-compatible logging.

No more, will people on the project be tormented by classic printf
features like:

- Accidentally passing in a non-trivial type
- PRI macros
- Not being able to add support for custom types
- Mixing up signed/unsigned printf formatting specifiers accidentally

fmt-capable versions of the logging functions are named the same as the
existing functions, just with a "Fmt" or _FMT suffix (depending on
whether or not it's a function being used or a macro, respectively).
2021-06-05 08:29:48 -04:00
Ryan Houdek 376892793f Merge pull request #1051 from lioncash/view
IRParser: Convert array of std::string over to array of std::string_view
2021-06-05 04:05:28 -07:00
Stefanos Kornilios Mitsis Poiitidis 61f73cf0bf Merge pull request #1050 from Sonicadvance1/fix_semctl_shmctl
Fixes semctl and msgctl
2021-06-05 13:54:30 +03:00
Stefanos Kornilios Mitsis Poiitidis 9716a4f0cf Merge pull request #1049 from Sonicadvance1/drm_headers
Adds drm headers to an external repository
2021-06-05 13:53:16 +03:00
Lioncash 35f3777797 IRParser: Make use of insert_or_assign where applicable
Avoids some minor potential default constructions that get overwritten
immediately.
2021-06-05 06:43:17 -04:00
Lioncash d73b79470e IRParser: Mark parameter of CheckPrintError as a const reference
Def is only ever accessed to read members, so we can signify to the
reader to not expect it to be modified.
2021-06-05 06:43:16 -04:00
Lioncash 676d9cb665 IRParser: std::move elements where advantageous
e.g. LineDefinitions are moderately beefy, they contain
two strings and a vector of strings among other things,
so we can move instances into their containing vector to avoid some
allocation churn.
2021-06-05 06:43:13 -04:00
Lioncash 993b4513d9 IRParser: Convert array of std::string over to array of std::string_view
Same behavior, but allows the arrays to be constexpr (and use less
space; 16 bytes vs 32 bytes per element).
2021-06-05 06:16:42 -04:00
Ryan Houdek 2958744777 Fixes semctl and msgctl
There were some problems in both of these implementations.

Fixes #744
Fixes #745
2021-06-04 23:40:24 -07:00
Ryan Houdek 2e8aacffe6 Fixes compat_ptr to return reference instead of copy
The expectation was that a reference would be returned rather than a copy.
Oops
2021-06-04 23:39:13 -07:00
Ryan Houdek 2101914c9d Adds drm headers to an external repository
It's highly likely that the host system won't have these headers installed.
Carry them in an external repository to ensure they are available.

Fixes #1047
2021-06-04 21:26:45 -07:00
Ryan Houdek 97c4ba018b Merge pull request #1048 from lioncash/lookup
Config: Avoid a few minor string copies where trivially possible
2021-06-04 17:13:10 -07:00
Lioncash 525d50f07e config: Pass strings by const reference where applicable
In a few cases, the strings aren't ever directly modified, so they can
be passed by reference to eliminate a few trivial copies.
2021-06-04 12:46:34 -04:00
Lioncash 5cec8ddca0 config: Move input strings in Loader constructors
Eliminates a copy, minor, but basically a "free" change.
2021-06-04 12:40:38 -04:00
Lioncash 39b2be15ab config: Make use of heterogenous map lookup
Same behavior, but allows lookups without constructing a std::string
(in most cases find() inputs use const char*, so this gets rid of some
string churn).
2021-06-04 12:36:54 -04:00
Ryan Houdek 83bf79ab42 Merge pull request #1045 from lioncash/default
IR: Make use of defaulted operator==
2021-06-03 22:01:58 -07:00
Lioncash 1f97a2baa9 IR: Make use of defaulted operator==
Same behavior, but allows both operator== and operator!= to be
automatically generated with a single declaration.
2021-06-04 00:52:25 -04:00
Ryan Houdek 0c2344166c Merge pull request #1044 from lioncash/moves
Config: Move strings where applicable
2021-06-03 18:33:00 -07:00
Lioncash 0bf30e2024 Config: Simplify qualifiers
These functions are part of the same class, so we can use the function
names directly without qualifiers.
2021-06-03 21:06:14 -04:00
Lioncash 933cdf76b4 Config: Move strings where applicable
Noticed when adding amending missing const qualifiers on interfaces. We
can make use of std::move here to avoid potential allocation churn a
little.

While we're at it, we can implement EraseSet in terms of, well, Erase()
and Set().
2021-06-03 21:03:19 -04:00
Ryan Houdek 36a77bb396 Merge pull request #1043 from lioncash/const
General: Add missing const specifiers where applicable
2021-06-03 17:40:13 -07:00
Lioncash 1936ebc59c General: Add missing const specifiers where applicable
Minor change that adds missing const specifiers to getters that don't
modify internal class state.
2021-06-03 20:24:27 -04:00
Ryan Houdek 9838309560 Merge pull request #1042 from lioncash/strong
DecodedOperand: Convert operand type into an enum class
2021-06-02 05:01:56 -07:00
Lioncash 37e2210f64 DecodedOperand: Convert operand type into an enum class
Now possible in a less messy way, since the type is now centralized in
one location.

Makes it strongly typed and prevents any potential accidental implicit
assignments to Type instead of something intended for the Data members.
2021-06-02 06:51:15 -04:00
Ryan Houdek 71c32c0135 Merge pull request #1041 from lioncash/info
DecodedOperand: Add helper functions for type testing
2021-06-02 03:10:49 -07:00
Lioncash 1f0bc49e54 DecodedOperand: Add helper functions for type testing
Significantly shortens the amount of code necessary for testing the type
of a decoded operand.

Also relocates the type field out of all the union types to have it in a
central location.
2021-06-02 05:59:12 -04:00
Lioncash 718f7ef8e6 X86InstInfo: Add operator!=
Provides logical symmetry.
2021-06-02 05:44:31 -04:00
Ryan Houdek edd1dfdfe8 Merge pull request #1040 from Sonicadvance1/more_syscalls_5_12
Implements more syscalls for supporting a higher guest kernel version
2021-06-01 00:27:50 -07:00
Ryan Houdek 4243ed19a4 Merge pull request #1037 from Sonicadvance1/fexconfig_downgrade
FEXConfig: Allow lower GL versions
2021-06-01 00:14:04 -07:00
Scott Mansell 16467c0fb7 Merge pull request #1038 from Sonicadvance1/implement_insertps
Implements SSE4.1 insertps
2021-06-01 19:04:53 +12:00
Ryan Houdek 0b7587768d Unifies 32bit and 64bit clone implementation
Syscall entry points still have different argument orders,
Moves the arguments to the clone3 argument structure and passes to generic handler.
Also implements clone3 while doing this
2021-05-27 23:00:00 -07:00
Ryan Houdek d115a57fe6 Implements support for execveat 2021-05-27 22:59:57 -07:00
Ryan Houdek 4c74478610 Implements iouring syscalls
This is a very simple initial implementation.
io_uring allows some things that are hard to capture like setting personalities.

Assume sane usage for now
2021-05-27 22:59:55 -07:00
Ryan Houdek 77a111cf19 Implements process_madvise 64-bit syscall 2021-05-27 22:59:52 -07:00
Ryan Houdek 7e551e5f88 Implements epoll_pwait2 syscall 2021-05-27 22:59:49 -07:00
Ryan Houdek 9d44dfdca6 Actually updated the hardcoded guest kernel locations
Just a couple of files and uname syscall
2021-05-27 22:59:47 -07:00
Ryan Houdek a7858f4d37 Calculates a guest kernel version for FEX instead
Takes the host kernel version and makes sure it fits in our supported kernel range
Minimum kernel version FEX reports to the guest is 5.0
Maximum kernel version FEX reports to the guest is currently 5.12
2021-05-27 22:59:44 -07:00
Ryan Houdek 6b3b10f8ae Implements pidfd_send_signal syscall
Now that we know when to forward a siginfo_t to the guest, we can allow this syscall
2021-05-27 22:59:42 -07:00
Ryan Houdek 9d5fa64f68 Implements pidfd_open syscall 2021-05-27 22:59:39 -07:00
Ryan Houdek 8b3bbd0ea1 Implements openat2 2021-05-27 22:59:37 -07:00
Ryan Houdek cc477b6964 Implements close_range syscall 2021-05-27 22:59:34 -07:00
Ryan Houdek 7a64bba8c4 Removes check for ThreadState being standard layout
Latest clang and libstdc++ makes unique_ptr not be standard layout.
Results in a compile error
2021-05-27 22:59:31 -07:00
Ryan Houdek 2d342d4662 Handle user provided siginfo_t with user signal
If the guest has sent a signal and the si_code is SI_USER then
we need to pass that siginfo_t through without touching it.
User could be sticking whatever they want in to that struct
2021-05-27 22:59:27 -07:00
Stefanos Kornilios Mitsis Poiitidis 7b41808806 Merge pull request #1039 from Sonicadvance1/pidfd_getfd
Implements pidfd_getfd syscall
2021-05-28 07:30:30 +03:00
Ryan Houdek 912d019ad8 Implements pidfd_getfd syscall
Rise of the Tomb Raider launcher uses this without checking host kernel version.
Might be part of their crash handler.
2021-05-27 18:44:12 -07:00
Ryan Houdek 1f9405b880 Implements insertps unit test 2021-05-26 21:41:16 -07:00
Ryan Houdek 211e7bf0f0 Implements SSE4.1 insertps
Cheap Golf is using this instruction unconditionally.
With this implemented the game now runs
2021-05-26 21:40:39 -07:00
Ryan Houdek 822a08f271 FEXConfig: Allow lower GL versions
Automatically fall back through older GL versions

This allows us to freely support GL 3.0, 2.1 and ES 2.0.
Should fix an issue where Pi devices don't support GL 3.0 in all configs

External imgui had to be updated to fix an issue with ES 2.0

Fixes #1036
2021-05-26 18:45:03 -07:00
Ryan Houdek 6cba775d4e Merge pull request #1032 from FEX-Emu/skmp/aotir-gen-mt
Multi threaded AOTGen
2021-05-24 23:06:13 -07:00
Ryan Houdek 8b5873061a Merge pull request #1033 from lioncash/fmtlib
Externals: Add fmtlib as an external
2021-05-20 14:58:33 -07:00
Lioncash 37a33bf127 Externals: Add fmt as an external
Begins the process of addressing issue #146.

This is separated off, so that others can make use of fmt for other
purposes (general localized non-sucky string formatting), while the
logging rework is being tackled.
2021-05-20 13:18:38 -04:00
Stefanos Kornilios Mitsis Poiitidis d70c91ddc6 AOTIR: Multi-threaded AOTGen 2021-05-19 14:19:55 +03:00
Stefanos Kornilios Mitsis Poiitidis 370f36c8f7 Merge pull request #1027 from Sonicadvance1/fix_ppoll
ppoll fixes
2021-05-18 09:20:32 +03:00
Ryan Houdek 42bb27b1fb Merge pull request #1025 from Sonicadvance1/ignore_non_canonical
Ignore non-canonical addresses in FS/GS setting
2021-05-12 22:31:09 -07:00
Ryan Houdek f522303837 gvisor: arch_prctl now passes 2021-05-12 22:17:13 -07:00
Ryan Houdek 3d09a55715 Ignore non-canonical addresses in FS/GS setting 2021-05-12 22:17:13 -07:00
Ryan Houdek c103f54774 Merge pull request #1028 from Sonicadvance1/remove_syscall_forwards
Removes syscall forward errno preprocessor implementation
2021-05-12 22:04:30 -07:00
Ryan Houdek dc7437b8c6 Merge pull request #1018 from FEX-Emu/skmp/streamable-aotir
AOTIR: .aotir files are now streamed out
2021-05-12 22:04:17 -07:00
Ryan Houdek b45503f93c gvisor: chown and sync tests now pass 2021-05-12 21:54:08 -07:00
Ryan Houdek 8cf1b0263d Removes syscall forward errno preprocessor implementation
This was causing issues with syscalls returning errors.
Remove it and move on
2021-05-12 21:53:45 -07:00
Stefanos Kornilios Mitsis Poiitidis 7ecaf24e8b Merge pull request #1030 from Sonicadvance1/disable_userfaultfd
Disable userfaultfd until supported
2021-05-12 09:45:43 +03:00
Stefanos Kornilios Mitsis Poiitidis 52f7ea5433 Merge pull request #1029 from Sonicadvance1/fix_fadvise64
Fix fadvise64
2021-05-12 09:45:21 +03:00
Stefanos Kornilios Mitsis Poiitidis 023a32ea26 Merge pull request #1026 from Sonicadvance1/fix_flag_remapping
Handle FD flag remapping correctly
2021-05-12 09:39:39 +03:00
Stefanos Kornilios Mitsis Poiitidis 36980131e0 Merge pull request #1024 from Sonicadvance1/signal_fixes
Sigaction fixes
2021-05-12 09:34:51 +03:00
Stefanos Kornilios Mitsis Poiitidis b2eb13fa1c Merge pull request #1023 from Sonicadvance1/emulate_map_32bit
Emulate MAP_32BIT in mmap
2021-05-12 09:34:18 +03:00
Stefanos Kornilios Mitsis Poiitidis 27b6497bee Merge pull request #1031 from Sonicadvance1/fix_epoll
Use epoll syscalls directly
2021-05-12 09:33:59 +03:00
Stefanos Kornilios Mitsis Poiitidis 66ee89ec8a Merge pull request #1022 from Sonicadvance1/fix_uname
Fix uname
2021-05-12 09:22:58 +03:00
Stefanos Kornilios Mitsis Poiitidis bee73199e1 Merge pull request #1021 from Sonicadvance1/fix_thread_self
Support redirecting thread-self exe softlink
2021-05-12 09:22:36 +03:00
Stefanos Kornilios Mitsis Poiitidis 0e6db843ef Merge pull request #1020 from Sonicadvance1/invalid_syscall
Return ENOSYS on too large of syscall number
2021-05-12 09:21:17 +03:00
Stefanos Kornilios Mitsis Poiitidis b45538f01e Merge pull request #1019 from Sonicadvance1/fix_emufd
Fix two EmuFD issues
2021-05-12 09:20:53 +03:00
Ryan Houdek f8cf98d378 gvisor: fadvise64 test now passes 2021-05-11 18:55:48 -07:00
Ryan Houdek 796c5ccbc9 gvisor: sigaction_test now passes 2021-05-11 18:54:08 -07:00
Ryan Houdek b157a5a0fb gvisor: bad_test now passes 2021-05-11 18:53:22 -07:00
Ryan Houdek 783ceb2d56 Use epoll syscalls directly 2021-05-11 18:51:06 -07:00
Ryan Houdek d1fc65daaa Disable userfaultfd until supported
Until we properly wrap this we can't support it
2021-05-11 18:44:31 -07:00
Ryan Houdek 94f464e1a4 Fix fadvise64
posix variant still doesn't quite match the actual syscall implementation
2021-05-11 18:39:00 -07:00
Ryan Houdek dfb15aaeb9 ppoll fixes
glibc implementation makes a copy of the timeout and the kernel is expected to update it
Can't use the glibc implementation because of this.
2021-05-11 18:33:24 -07:00
Ryan Houdek 565f0e6af3 Handle FD flag remapping correctly
FD flag remapping was broken. It would remap one flag on to another and then the next check would remap it back.

Instead keep a mask of the flags to be remapped then remap them all at the end.
Also goes through the ops and fixes a few cases where it was remapping wrong flags.
2021-05-11 18:29:21 -07:00
Ryan Houdek b03554bb7d Sigaction fixes
Check for invalid sigsetsize

Since we are using SignalDelegator we don't use errno, so just return Result and check for Result == 0
2021-05-11 18:22:23 -07:00
Ryan Houdek 7b8cb5107a Emulate MAP_32BIT in mmap
If we are on AArch64 then MAP_32BIT doesn't exist.
Emulate it by setting the address as a hint to the kernel so it scans bottom up
2021-05-11 18:18:59 -07:00
Ryan Houdek cbb20ef6c6 Fix uname
There is a domainname variable that was missed
2021-05-11 18:16:35 -07:00
Ryan Houdek 298481f9af Support redirecting thread-self exe softlink 2021-05-11 18:15:14 -07:00
Ryan Houdek 0d108cd8fd Return ENOSYS on too large of syscall number 2021-05-11 18:14:02 -07:00
Ryan Houdek 490d7e57d7 Arguments file needs one additional null argument to finish the arguments. 2021-05-11 18:12:16 -07:00
Ryan Houdek fe77c3ef06 Fix cpuinfo needing tabs on its options and PM option
Depending on option it is split by anywhere from zero to two tabs
2021-05-11 18:12:16 -07:00
Ryan Houdek eb02afe952 Merge pull request #1013 from Sonicadvance1/fix_openat_symlinks
Handles symlinks in rootfs in openat
2021-05-11 17:52:07 -07:00
Stefanos Kornilios Mitsis Poiitidis b8dc63754b AOTIR: Don't keep IR,RA,Code data around when aotirgenerating 2021-05-11 15:15:39 +03:00
Stefanos Kornilios Mitsis Poiitidis df8f1850ae AOTIR: .aotir files are now stream-written 2021-05-11 14:37:12 +03:00
Ryan Houdek e282fd2221 Merge pull request #1016 from Sonicadvance1/ioctl32_micro_optimization
Microoptimization for the DRM ioctls
2021-05-10 21:14:16 -07:00
Stefanos Kornilios Mitsis Poiitidis 3147f0d84f Merge pull request #1017 from Sonicadvance1/fix_x86_jitsymbols
Fixes JITSymbols for x86-64 JIT
2021-05-10 00:36:30 +03:00
Stefanos Kornilios Mitsis Poiitidis 4c02c9f037 Merge pull request #1015 from Sonicadvance1/fix_inotify_thread
FEXConfig: Fixes INotify watcher never coming up
2021-05-10 00:34:27 +03:00
Ryan Houdek 83b4188569 Microoptimization for the DRM ioctls
For DRM applications have a three FD deep MRU cached for faster lookups of FD to DRM handlers.
In a completely DRM ioctl bound situation like es2gears or GL application without threaded context
Then this puts us /nearly/ at the performance of calling the ioctl32 handler directly.
Sadly there is overhead that can't be overcome so this is the best that can be done from userland
2021-05-06 02:55:22 -07:00
Ryan Houdek 82b04d9886 Fixes JITSymbols for x86-64 JIT 2021-05-06 01:12:51 -07:00
Ryan Houdek 36fdd7f6e2 FEXConfig: Fixes INotify watcher never coming up
Inverted this check on accident
2021-05-05 19:44:31 -07:00
Ryan Houdek 6a1ca654db Docs: Update for release FEX-2105 2021-05-05 17:53:55 -07:00
Ryan Houdek 623b29de8c Merge pull request #1014 from Sonicadvance1/fix_host_stacks
Fixes host thread stacks ending up in lower 32-bit VA
2021-05-05 15:51:40 -07:00
Ryan Houdek 9be67fb22a Merge pull request #1009 from Sonicadvance1/handle_self
Handle more cases of an application pinging self
2021-05-05 14:51:11 -07:00
Ryan Houdek cb9286c11e Merge pull request #1012 from Sonicadvance1/fix_ioctl_errors
Fixes ioctl returning error
2021-05-05 14:50:46 -07:00
Ryan Houdek 1e40148f60 Fixes host thread stacks ending up in lower 32-bit VA
Overriding glibc's mmap and munmap functions don't encapsulate that functions they use
for allocating stack data so it was falling down the system mmap/munmap path.
This was causing host side stacks to end up in the 32-bit VA space in 32-bit applications.
2021-05-04 22:15:46 -07:00
Ryan Houdek 19e45544fb Handles symlinks in rootfs in openat
This fixes an issue where rootfs has symlinks to other things in the rootfs so we need to track it through.

Relies on #1009 to be merged first.
Fixes any application that relies on libblas, mpv for example.
2021-05-04 22:13:52 -07:00
Ryan Houdek d9d5303aa0 Fixes ioctl returning error
Fixes evdev device description at the very least
2021-05-04 18:45:25 -07:00
Ryan Houdek 517b575783 Merge pull request #1010 from FEX-Emu/skmp/aotupdate-install-cmake
AOTIR: Add install rule for FEXUpdateAOTIRCache
2021-05-04 18:11:27 -07:00
Ryan Houdek 605994dca2 Merge pull request #1011 from FEX-Emu/skmp/fix-0xcd-breaks
JIT: Add reason 1 (0xCD) to OP_BREAK implementations
2021-05-04 18:11:17 -07:00
Ryan Houdek ac97c05271 Handle more cases of an application pinging self
AppImage uses this to inspect its own executable to ensure it is an AppImage file.
2021-05-04 18:09:15 -07:00
Stefanos Kornilios Mitsis Poiitidis c8bb02f3a6 JIT: Add reason 1 (0xCD) to OP_BREAK implementations 2021-05-04 13:22:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 6b12c1768f AOTIR: Add install rule for FEXUpdateAOTIRCache 2021-05-04 12:33:59 +03:00
Stefanos Kornilios Mitsis Poiitidis 693097c5b1 Merge pull request #1008 from Sonicadvance1/fix_binfmt_misc
Updates binfmt_misc files to support AppImage
2021-05-04 11:14:49 +03:00
Ryan Houdek 47e823b773 Updates binfmt_misc files to support AppImage
AppImage files stick additional data in the ABI Version and PAD areas that it was getting blocked on.
2021-05-03 18:08:00 -07:00
Ryan Houdek fd6d33d197 Merge pull request #996 from FEX-Emu/skmp/aotirgen
AOTIR: Offline Generation
2021-05-03 15:58:57 -07:00
Stefanos Kornilios Mitsis Poiitidis d8b583bc51 Remove assert that was used in debugging 2021-05-04 01:47:39 +03:00
Stefanos Kornilios Mitsis Poiitidis be41f14452 Fix generation script 2021-05-03 18:16:43 +03:00
Stefanos Kornilios Mitsis Poiitidis cf5bd202d4 AOTIR: Update gen script to default to global fex installation 2021-05-03 06:15:25 +03:00
Stefanos Kornilios Mitsis Poiitidis 8893cc6a7d Undo xbyak update 2021-05-03 05:48:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 1d82e9e59e JIT: Stub Break reasons 2, 3 2021-05-03 05:42:46 +03:00
Stefanos Kornilios Mitsis Poiitidis 5b017f9a66 Merge pull request #1007 from Sonicadvance1/more_ioctl32_emulation2
Adds Nouveau and Lima ioctl32 handlers
2021-05-03 05:10:21 +03:00
Ryan Houdek 3399052eaa Make sure to setup ioctl32 handler table 2021-05-02 14:30:23 -07:00
Ryan Houdek 3df951695e Adds Nouveau and Lima ioctl32 handlers 2021-05-02 13:05:46 -07:00
Ryan Houdek 8d1f9a9081 Merge pull request #1006 from Sonicadvance1/implement_clflush
Implements support for CLFLUSH
2021-05-02 12:27:54 -07:00
Ryan Houdek 68adb1ae12 Basic clflush test 2021-05-02 12:18:35 -07:00
Ryan Houdek 4961604807 Implements support for CLFLUSH
JRE requires this to run
2021-05-02 12:18:35 -07:00
Stefanos Kornilios Mitsis Poiitidis d6b7f32d33 AOTGen: Cleanup unwind symbol logic, add to x86-32 binaries 2021-05-02 11:08:12 +03:00
Stefanos Kornilios Mitsis Poiitidis 6dc69063f3 Merge pull request #1005 from Sonicadvance1/remove_numa
Remove reliance on librt and libnuma
2021-05-02 09:50:10 +03:00
Stefanos Kornilios Mitsis Poiitidis 2916c4f34c Merge pull request #1004 from Sonicadvance1/redirect_self
Redirect applications using execve with self
2021-05-02 09:49:55 +03:00
Stefanos Kornilios Mitsis Poiitidis 04137d0da5 Merge pull request #1002 from Sonicadvance1/sigtimedwait
Initial base implementation of sigtimedwait
2021-05-02 09:49:49 +03:00
Stefanos Kornilios Mitsis Poiitidis afa81294c9 Merge pull request #1001 from Sonicadvance1/more_ioctl32_emulation
More ioctl32 emulation
2021-05-02 09:49:44 +03:00
Ryan Houdek 86bdbe65e5 Implements missed x86-32 message queue syscalls
Nothing hit these but good to have them implemented
2021-05-01 17:40:33 -07:00
Ryan Houdek 63dee0e132 Remove reliance on librt and libnuma
These can fall down the raw syscall path
2021-05-01 17:40:33 -07:00
Ryan Houdek 6618bb809b Redirect applications using execve with self
Fixes a step in the Java JRE and shapez.io
2021-05-01 11:09:44 -07:00
Ryan Houdek d2b19c0b1c Disable gvisor sigtimedwait test 2021-05-01 10:19:13 -07:00
Ryan Houdek 90a8c4f114 Shuffle timed sigwait unit tests that have changed behaviour 2021-05-01 10:11:20 -07:00
Ryan Houdek 866baf66db Merge pull request #1003 from Sonicadvance1/cleaned_sendrecvmsg
Workaround fixes in sendmsg/recvmsg
2021-05-01 10:06:41 -07:00
Ryan Houdek 9498a41f9b More ioctl32 emulation
More DRM and i915 things, then wireless for some reason renderdoc relies on it
2021-05-01 09:40:02 -07:00
Ryan Houdek 1340ffabe4 Workaround fixes in sendmsg/recvmsg
Rebased and cleaned up #933
Needs to be larger allocations than what was originally proposed by #933
2021-05-01 09:37:01 -07:00
Ryan Houdek b5115e096f Initial base implementation of sigtimedwait 2021-05-01 09:35:29 -07:00
Ryan Houdek 235e05e8b9 Merge pull request #978 from Sonicadvance1/wip_allocator
64-bit allocator and ioctl emulation
2021-05-01 09:29:46 -07:00
Ryan Houdek 26b4bd80bf Force 32-bit allocator for CI 2021-05-01 09:19:55 -07:00
Ryan Houdek a8881e8835 Adds option to force 32-bit allocator 2021-05-01 09:19:55 -07:00
Ryan Houdek d1da1b0be5 Disables 20080723-1.c.gcc-target-test-32 CI test
This test consistently fails on the solid run board.
This requires some deep dive investigation
2021-05-01 09:06:39 -07:00
Ryan Houdek a19c0184a6 Keep using the 32-bit allocator if the kernel is old
Works around the Xavier boards running kernel 4.9
2021-05-01 09:06:39 -07:00
Ryan Houdek 2d0dced9e7 Work around old asound headers 2021-05-01 09:06:39 -07:00
Ryan Houdek 52cb96d630 Work around old DRM headers for i915 2021-05-01 09:06:39 -07:00
Ryan Houdek 94017d2e65 On old kernels, steal the 32-bit address space before the 64-bit
Increases speed of startup
2021-05-01 09:06:39 -07:00
Ryan Houdek 13b93ef6b5 Allow memory allocation to go through if the base was above the lower bound 2021-05-01 09:06:39 -07:00
Ryan Houdek c83a5d1bc3 Allow some more clone flags through with info log 2021-05-01 09:06:39 -07:00
Ryan Houdek c20777d558 Change these log messages so we know which one failed 2021-05-01 09:06:38 -07:00
Ryan Houdek 9a0fc4e7e8 Workaround quick with memory allocator. Aligned 4GB regions are hard to allocate 2021-05-01 07:39:14 -07:00
Ryan Houdek 417e8836a4 Arm64: Fixes some wrong sign extensions 2021-05-01 07:39:14 -07:00
Ryan Houdek 237865e064 Adds ioctl verification to struct verifier 2021-05-01 07:39:14 -07:00
Ryan Houdek 45223cc7e7 Adds initial ioctl emulation code 2021-05-01 07:39:14 -07:00
Ryan Houdek 8d34222556 Add vardecl type matching for ioctl verification 2021-05-01 07:39:14 -07:00
Ryan Houdek b6344c4753 Removes old 32-bit allocator usage
Replaces with using host side allocations since we always have a 64-bit allocator in this case
2021-05-01 07:39:14 -07:00
Ryan Houdek c5bda47b2a Adds more allocator hooking 2021-05-01 07:39:14 -07:00
Ryan Houdek 5067737052 Adds 64-bit allocator 2021-05-01 07:39:14 -07:00
Ryan Houdek 5143100a73 Switches more usages of memory allocation routines over to FEX internal 2021-05-01 07:39:14 -07:00
Stefanos Kornilios Mitsis Poiitidis 8b6e767742 Frontend: Slightly more resilient max instruction size handling 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 331549c64a RLSE: Use Constant, not inline constant, let ConstProp generate all inline consts 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis be0b29bbc6 OpDisp: Convert more asserts to soft-errors 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis ee8f8ed176 Improve error handling 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 5b8cbae1e1 Core/Frontend: Improve error handling 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis b98ef9e48b Core/Frontend: Improve error handling 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis fd724448ab Front End: Fix an error check 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis f4083883ca Frontend: Fix ExternalBranches initialization 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 9a4fc758c1 AOTIR: FEXUpdateAOTIRCache.sh now passes compilation modifiers 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 40d2cd8992 Fix asserts 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis ab44f800f7 AOTgen: Cleaner interfaces, use LogMan 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 1a5bc39d09 Rebase fixes 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 4bd997adaa AOT: Postfix ir files with .aotir 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 244ef94e9b AOTIR: Generate list of files with code, add Scripts/FEXUpdateAOTIRCache.sh to generate all files 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis a85ef061d1 AOITR: Add AOTIRGenerate fo FEXLoader option that does generation 2021-05-01 09:40:50 +03:00
Stefanos Kornilios Mitsis Poiitidis 4e847d03b7 Remove more assertions for invalid code 2021-05-01 09:40:49 +03:00
Stefanos Kornilios Mitsis Poiitidis ce9324f50f Kinda-parse unwind table 2021-05-01 09:40:49 +03:00
Stefanos Kornilios Mitsis Poiitidis 0897fc4a0a pretrans wip 2021-05-01 09:40:49 +03:00
Ryan Houdek d256bc59fd Override glibc allocator functions 2021-04-30 18:28:45 -07:00
Ryan Houdek e0d6f4bc34 Make sure link jemalloc in FEXCore 2021-04-30 18:28:45 -07:00
Ryan Houdek 09284331a7 Adds pregen include directory for jemalloc 2021-04-30 18:28:45 -07:00
Ryan Houdek 73e61b0c8e Switches default logging output to stderr
stdout is too likely to break something
2021-04-30 18:28:45 -07:00
Ryan Houdek a30866a886 Merge pull request #976 from FEX-Emu/skmp/aotir-mmap
AOTIR: Switch over to mmap-based loading
2021-04-30 18:24:54 -07:00
Stefanos Kornilios Mitsis Poiitidis 3e1da51cae Merge pull request #998 from Sonicadvance1/remove_xxhash_check
Removes xxhash version check
2021-04-30 12:05:18 +03:00
Stefanos Kornilios Mitsis Poiitidis 4efe973b65 Merge pull request #999 from Sonicadvance1/fix_sdl2_interface_Again
Fixes SDL2 interface check...again.
2021-04-30 12:04:50 +03:00
Ryan Houdek f6cc75de87 Fixes SDL2 interface check...again. 2021-04-30 01:55:08 -07:00
Ryan Houdek 6b082f4ef7 Removes xxhash version check
xxhash doesn't yet have a version of the library with pkg-config giving
a version.
Debian has backported a change for this but Arch hasn't.
https://github.com/Cyan4973/xxHash/issues/524
2021-04-30 01:35:03 -07:00
Stefanos Kornilios Mitsis Poiitidis 8a73783a07 AOTIR: Use FEXCore::Allocator 2021-04-29 08:16:21 +03:00
Stefanos Kornilios Mitsis Poiitidis 32b4b02302 AOTIR: mmap based loading 2021-04-28 13:08:43 +03:00
Stefanos Kornilios Mitsis Poiitidis 528b01ad7a Merge pull request #988 from Sonicadvance1/memory_layout_optimizations
Memory layout optimizations
2021-04-28 12:31:53 +03:00
Stefanos Kornilios Mitsis Poiitidis cd0492a143 Merge pull request #994 from Sonicadvance1/enable_armv8.4
Enables ARMv8.4 CI runner
2021-04-28 09:53:47 +03:00
Ryan Houdek 26bf903b93 Enables ARMv8.4 CI runner
Requires #990 and #993 to be merged first
2021-04-27 23:16:39 -07:00
Stefanos Kornilios Mitsis Poiitidis aae4de106a Merge pull request #992 from Sonicadvance1/optimize_cpuid
Optimizes CPUID generation a bit
2021-04-28 09:12:09 +03:00
Stefanos Kornilios Mitsis Poiitidis c495d8120d Merge pull request #993 from Sonicadvance1/fix_sdl_missed_line
Oops: Missed line on FEXConfig SDL2
2021-04-28 09:10:52 +03:00
Stefanos Kornilios Mitsis Poiitidis a6d9425091 Merge pull request #990 from Sonicadvance1/fix_native_m1
Work around M1 Parallels hypervisor not showing CPU type
2021-04-28 09:09:08 +03:00
Ryan Houdek b06ebf486f Work around M1 Parallels hypervisor not showing CPU type
Parallels claims that the CPU is of implementor 0x41 and part number 0.
Easy enough to work around before real numbers get exposed
2021-04-27 22:26:35 -07:00
Ryan Houdek b17464d81a Oops: Missed line on FEXConfig SDL2 2021-04-27 16:51:12 -07:00
Ryan Houdek 2ab8a97055 Merge pull request #989 from Sonicadvance1/sdl2_fix
Only use SDL2 import target if it exists
2021-04-27 01:11:28 -07:00
Ryan Houdek 34a5eb6d6f Optimizes CPUID generation a bit
The flags calculation is fairly intensive and we were doing it for every core.
Do it once then duplicate it per core instead

Improves generation time quite significantly. A couple of milliseconds converted to half a millisecond
Slight improvement to startup time
2021-04-27 01:09:28 -07:00
Ryan Houdek 2a1a863c58 Only use SDL2 import target if it exists
Should fix Arch and Ubuntu 20.04 together
2021-04-27 00:45:20 -07:00
Ryan Houdek bab5927931 Merge pull request #981 from Sonicadvance1/optimize_long_divide
Adds Long Divide removal pass
2021-04-26 23:48:50 -07:00
Ryan Houdek 6255eba524 Memory layout optimizations
Some of these memory layouts weren't optimal for ARM loading.
This is now significantly improved
2021-04-26 23:48:22 -07:00
Ryan Houdek 37108058d1 Removes some unnecessary branching checks
It doesn't actually matter if the invalid node has has its linked list nodes changed
Faster to just set than to check while in a tight loop
2021-04-26 23:48:22 -07:00
Stefanos Kornilios Mitsis Poiitidis c3ad571062 Merge pull request #987 from Sonicadvance1/fix_assert_dumb_code
Changes LogMan assert to defines
2021-04-27 09:41:12 +03:00
Ryan Houdek 04e0baadc6 Changes LogMan assert to defines
C++ no-op functions can't optimize out the predicate arguments in all cases.
This was causing a problem where zero cost assertions weren't actually zero cost.

The only way to resolve this is to actually use macros sadly enough.

This will give a fairly hefty performance uplift with anything operating on IR.
2021-04-26 13:49:35 -07:00
Ryan Houdek 5ffbd97f01 IROp types can be marked const
Allows the compiler to know it can be optimized out without LTO
2021-04-26 13:45:37 -07:00
Ryan Houdek 7d4380fe6d Merge pull request #984 from FEX-Emu/skmp/relocation-fixes
IR: Fix RIP relocation edge-cases
2021-04-26 10:11:24 -07:00
Ryan Houdek 7d5157b602 Merge pull request #983 from FEX-Emu/skmp/faster-l1
LookupCache: Add new entries to L1C
2021-04-26 10:10:51 -07:00
Ryan Houdek e9a32da997 Merge pull request #986 from FEX-Emu/skmp/fix-20.04-build
Deps: Require xxhash 0.7.3, ubuntu 20.04 ships with that version
2021-04-26 10:05:35 -07:00
Stefanos Kornilios Mitsis Poiitidis b3564a4a48 IR: Fix RIP relocation edge-cases 2021-04-26 19:20:01 +03:00
Stefanos Kornilios Mitsis Poiitidis 2294419353 Deps: Require xxhash 0.7.3, ubuntu 20.04 ships with that version 2021-04-26 19:07:23 +03:00
Stefanos Kornilios Mitsis Poiitidis df8b78b327 Dispatcher: Use Also lookup L1C when dispatching interpreter 2021-04-26 19:06:11 +03:00
Stefanos Kornilios Mitsis Poiitidis 2ea6a2a141 LookupCache: Fix L1C to actually cache entries 2021-04-26 18:38:03 +03:00
Ryan Houdek db776fae4e Merge pull request #985 from FEX-Emu/skmp/rip-relative-compilation
IR: Remove Entry from OP_HEADER, pass as parameter to CompileCode
2021-04-25 18:27:37 -07:00
Stefanos Kornilios Mitsis Poiitidis b45b7c3441 IR: Remove Entry from OP_HEADER, pass as parameter to CompileCode 2021-04-25 13:15:23 +03:00
Stefanos Kornilios Mitsis Poiitidis 7805552edb LookupCache: Add new entries to L1C 2021-04-25 13:03:43 +03:00
Ryan Houdek 520dfd7edf Adds Long Divide removal pass
If the long divides don't have anything in the upper bits then they can be optimized away.
Teeworlds, SuperTuxKart, and RRootage were between 50%-55% getting removed
2021-04-24 23:10:59 -07:00
Stefanos Kornilios Mitsis Poiitidis 60671ee6cb Merge pull request #975 from Sonicadvance1/use_xxhash
Switch from fasthash64 to xxhash's XXH3
2021-04-23 09:49:06 +03:00
Stefanos Kornilios Mitsis Poiitidis 6635765ea9 Merge pull request #972 from Sonicadvance1/naked_mman
Naked mman usage removal
2021-04-22 09:47:41 +03:00
Stefanos Kornilios Mitsis Poiitidis 3e4ff23ab6 Merge pull request #977 from Sonicadvance1/fix_relative_elf
Fixes ELFCodeLoader2 loading relative ELF files
2021-04-22 09:40:08 +03:00
Ryan Houdek 93ed411899 Fixes ELFCodeLoader2 loading relative ELF files 2021-04-20 01:58:37 -07:00
Ryan Houdek 0b0db08ebf Switch from fasthash64 to xxhash's XXH3
Fixes #797
2021-04-18 18:13:27 -07:00
Stefanos Kornilios Mitsis Poiitidis 3419f40c00 Merge pull request #971 from Sonicadvance1/template_specialization
Declare missing template specialization for GetListIfExists
2021-04-19 00:03:25 +03:00
Stefanos Kornilios Mitsis Poiitidis 8260ddda13 Merge pull request #970 from Sonicadvance1/sdl2_target_imports
Use cmake SDL2 target properties
2021-04-19 00:02:59 +03:00
Stefanos Kornilios Mitsis Poiitidis 872e49f02a Merge pull request #969 from Sonicadvance1/mcpu_native
Use mcpu=native with Clang 12
2021-04-19 00:02:22 +03:00
Stefanos Kornilios Mitsis Poiitidis f0f6bc7a73 Merge pull request #973 from Sonicadvance1/fix_openat
Fixes OpenAt breaking with anonymous objects
2021-04-17 15:51:24 +03:00
Ryan Houdek 5820c251e8 Fixes OpenAt breaking with anonymous objects
Fixes #753
2021-04-16 21:21:08 -07:00
Ryan Houdek eff97509c0 Replaces naked mman usage in TestHarnessRunner 2021-04-16 21:01:59 -07:00
Ryan Houdek 3128d0c148 Replaces naked mman usage in SyscallHandler 2021-04-16 21:01:39 -07:00
Ryan Houdek 0216d7b552 Replaces naked mman usage in x86-64 syscalls 2021-04-16 21:01:07 -07:00
Ryan Houdek 87b473c7ee Replaces naked mman usage in IRLoader 2021-04-16 21:00:51 -07:00
Ryan Houdek f9d647c852 Replaces naked mman usage in HarnessHelpers 2021-04-16 21:00:19 -07:00
Ryan Houdek 2c175f1e0b Replaces naked mman usage in FEXLoader 2021-04-16 21:00:01 -07:00
Ryan Houdek 96813b0d6c Replaces naked mman usage in ELFCodeLoader 2021-04-16 20:59:27 -07:00
Ryan Houdek 967ac863be Replaces naked mman usage in X86Dispatcher 2021-04-16 20:50:14 -07:00
Ryan Houdek fa546fb492 Replaces naked mman usage in ELFSymbolDatabase 2021-04-16 20:50:09 -07:00
Ryan Houdek 99ab9864aa Replaces naked mman usage in X86HelperGen 2021-04-16 20:50:03 -07:00
Ryan Houdek 11246355e2 Replaces naked mman usage in LookupCache 2021-04-16 20:49:58 -07:00
Ryan Houdek 0d236940e0 Replaces naked mman usage in x86-64 JIT 2021-04-16 20:49:52 -07:00
Ryan Houdek a0afb4be5a Replaces naked mman usage in Arm64 JIT 2021-04-16 20:49:47 -07:00
Ryan Houdek ca45485d48 Adds new Allocator namespace for replace mmap and munmap
This is necessary for overriding these functions
2021-04-16 20:49:35 -07:00
Ryan Houdek c8bb0e2d51 Declare missing template specialization for GetListIfExists
Clang 12 picks up that this was missing
2021-04-16 18:54:16 -07:00
Ryan Houdek e9aeebb3a1 Use cmake SDL2 target properties
Some distros include generated cmake files instead of the two cmake files from the source.
Distros that use the generated cmake files only have SDL import properties, so we are required
to use those.
2021-04-16 18:45:06 -07:00
Ryan Houdek 1ad46c5af1 Use mcpu=native with Clang 12
Clang 12 fixes the bug with big.little configurations so we
can keep using mcpu=native past this point
2021-04-16 18:44:34 -07:00
Ryan Houdek 08bc7a8364 Merge pull request #967 from FEX-Emu/skmp/new-elf-loading
ELF: Simpler loader, uses mmap
2021-04-16 18:35:56 -07:00
Ryan Houdek e7179ff84b Merge pull request #968 from FEX-Emu/skmp/fix-self-exe-after-fork
Emulated Files: Don't precalculate /proc/pid path, it changes after fork
2021-04-15 13:01:41 -07:00
Stefanos Kornilios Mitsis Poiitidis 13bfdcad06 Emulated Files: Don't precalculate /proc/pid path, it changes after fork 2021-04-15 15:33:10 +03:00
Stefanos Kornilios Mitsis Poiitidis 38a707b38e Add missing functions to elfcodeloader2 2021-04-15 13:51:22 +03:00
Stefanos Kornilios Mitsis Poiitidis ada7915b01 Convert an assert to log 2021-04-15 13:40:28 +03:00
Stefanos Kornilios Mitsis Poiitidis 9775a2148e Fix rootfs interpreter loading 2021-04-15 12:16:33 +03:00
Stefanos Kornilios Mitsis Poiitidis bc8696828b More logging 2021-04-15 12:14:55 +03:00
Stefanos Kornilios Mitsis Poiitidis 7d792e65f0 Rootfs, assert fixes 2021-04-15 12:04:42 +03:00
Stefanos Kornilios Mitsis Poiitidis 14b5b881c5 Add more logging 2021-04-15 11:47:42 +03:00
Stefanos Kornilios Mitsis Poiitidis 830e41459c Fix build for test harness, irloader 2021-04-15 11:05:07 +03:00
Stefanos Kornilios Mitsis Poiitidis 0db99ea4f6 Use allocator for x32 2021-04-15 10:51:16 +03:00
Stefanos Kornilios Mitsis Poiitidis 22f1cf0e2a More wip 2021-04-15 10:05:59 +03:00
Stefanos Kornilios Mitsis Poiitidis 80493487db ELF Loading: Cleanups all around 2021-04-14 10:59:23 +03:00
Stefanos Kornilios Mitsis Poiitidis 59ce750af6 Elf: Add ELFCodeLoader2, with simplified loading logic 2021-04-13 15:59:36 +03:00
Stefanos Kornilios Mitsis Poiitidis 4ac81e22de Move some more files around 2021-04-13 15:59:36 +03:00
Stefanos Kornilios Mitsis Poiitidis a0a382cc5c Clear up file naming a big 2021-04-13 15:59:36 +03:00
Ryan Houdek 6cc7d065aa Merge pull request #964 from Sonicadvance1/paranoid_tso
First step in paranoid TSO mode
2021-04-12 23:34:54 -07:00
Ryan Houdek 93b9ca4bdb Merge pull request #965 from Sonicadvance1/add_jemalloc
Switches FEX over to using jemalloc
2021-04-08 14:04:12 -07:00
Ryan Houdek cc0c45cf47 Switches FEX over to using jemalloc
This is the first step that we need to do for correct 32-bit memory allocations.
2021-04-08 12:42:37 -07:00
Ryan Houdek 4a504bdd00 Merge pull request #939 from Sonicadvance1/hidden_visibility
Default hidden visibility and strip symbols
2021-04-06 18:56:28 -07:00
Ryan Houdek 66c8dd2013 First step in paranoid TSO mode
This moves vector ops down the GPR atomic backpatch path.
Next step after this is to add SIGBUS based handlers for these rather than backpatching.
Needs more work to treat aligned versus unaligned vector loadstores differently
2021-04-06 18:48:36 -07:00
Ryan Houdek 91ac4ccdd8 Adds LDAXP and STLXP backpatch handlers
This will be used in the next commit
2021-04-06 18:47:50 -07:00
Ryan Houdek 849b83bd02 Adds new ParanoidTSO option 2021-04-06 18:46:52 -07:00
Ryan Houdek 286f103dab Adds visibility attributes to everything we need to expose 2021-04-06 01:01:31 -07:00
Ryan Houdek 2a49037789 update external vixl 2021-04-06 00:39:03 -07:00
Ryan Houdek 235f6fbe39 Strip symbols and dead sections
These are unnecessary and just add some file bloat. Strip them out
2021-04-06 00:39:03 -07:00
Ryan Houdek b43db294ec Change our compilation visibility to hidden
This bloats our compilation with symbols that don't need to be visible
2021-04-06 00:39:03 -07:00
Stefanos Kornilios Mitsis Poiitidis 5e6c6adf8b Merge pull request #956 from Sonicadvance1/enable_gcc_32_tests
Enables GCC 32-bit unit tests
2021-04-06 09:48:38 +03:00
Stefanos Kornilios Mitsis Poiitidis 50c3ee047c Merge pull request #961 from Sonicadvance1/fix_segment_override
Fixes segment register override on string instructions.
2021-04-06 09:24:23 +03:00
Stefanos Kornilios Mitsis Poiitidis c0ef6da6a8 Merge pull request #960 from Sonicadvance1/more_cpuid
Adds some more CPUID functions
2021-04-06 09:23:13 +03:00
Stefanos Kornilios Mitsis Poiitidis 2063411443 Merge pull request #959 from Sonicadvance1/fix_tag_order
Fixes ordering problem of tag generation in release script
2021-04-06 09:22:12 +03:00
Stefanos Kornilios Mitsis Poiitidis 4a5eefec25 Merge pull request #958 from Sonicadvance1/fix_atomic_zext
Fixes Zext semantics on atomic ops with x64 JIT
2021-04-06 09:21:44 +03:00
Stefanos Kornilios Mitsis Poiitidis eaad9e99a9 Merge pull request #954 from Sonicadvance1/posix_tests_docs
Updates a couple POSIX test case failures
2021-04-06 09:18:18 +03:00
Stefanos Kornilios Mitsis Poiitidis 62406a45be Merge pull request #953 from Sonicadvance1/env_var_override
Allows overriding FEX folder locations with environment variables
2021-04-06 09:16:12 +03:00
Stefanos Kornilios Mitsis Poiitidis 03537d351d Merge pull request #952 from Sonicadvance1/add_asan_use_after_scope
Adds use-after-scope to the ASAN options
2021-04-06 09:09:09 +03:00
Ryan Houdek 0cb53fff89 Change kernel version error to not be fatal 2021-04-05 23:09:06 -07:00
Ryan Houdek 3e250fa071 Enables GCC 32-bit unit tests
Relies on #950 getting merged first
2021-04-05 23:09:06 -07:00
Stefanos Kornilios Mitsis Poiitidis 1d4b81f556 Merge pull request #951 from Sonicadvance1/fix_getenv_aotir
Removes getenv usage for aotir
2021-04-06 09:08:46 +03:00
Stefanos Kornilios Mitsis Poiitidis c26f19ed4d Merge pull request #950 from Sonicadvance1/fix_gcc_tests
Fixes GCC tests
2021-04-06 09:05:45 +03:00
Stefanos Kornilios Mitsis Poiitidis 9619cea097 Merge pull request #949 from Sonicadvance1/arm64_print
Adds support for Arm64 JIT to Print
2021-04-06 08:58:13 +03:00
Stefanos Kornilios Mitsis Poiitidis 44dbecc2ac Merge pull request #948 from Sonicadvance1/fix_leak
Fixes memory leak in CPU backends
2021-04-06 08:56:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 415632209f Merge pull request #947 from Sonicadvance1/minimal_siginfo
Implements a minimal siginfo_t in the dispatcher's signal handler
2021-04-06 08:54:01 +03:00
Ryan Houdek b7e81e867d Fixes segment register override on string instructions.
The segment prefix only overrides one side of the string instruction.
Passing in the flags was causing the instruction to have the segment override on both sides of the copy.
So something like `movsb es:[rdi], fs:[rsi]` was being interpreted as `movsb fs:[rdi], fs:[rsi]`
Same for all the string ops which is now fixed.

We can't unit test these currently since we can't set host segment register for confirmed host behaviour.
2021-04-05 21:28:31 -07:00
Ryan Houdek 4f276aae82 Adds comment for LEA and segment registers
LEA doesn't give you the result with segment register applied to it.
Adds a comment to OpDispatcher to make it more clear
2021-04-05 21:28:25 -07:00
Ryan Houdek c80254f6f4 Adds some more CPUID functions
Noticed a game poking at these
2021-04-05 18:37:05 -07:00
Ryan Houdek 8adb1ca3a2 Fixes ordering problem of tag generation in release script
Wasn't actually creating an annotated tag with this order
2021-04-05 18:35:09 -07:00
Ryan Houdek 0dd1959d65 Fixes Zext semantics on atomic ops with x64 JIT
Fixes Factorio 1.1.30 on x86-64 hosts

Fixes #957
2021-04-05 18:19:15 -07:00
Ryan Houdek f4ab7cceab Updates a couple POSIX test case failures
The behaviour of these is broken but now they are documented in how they are broken
2021-04-05 04:27:52 -07:00
Ryan Houdek 8cc799ab65 Allows overriding FEX folder locations with environment variables
Introduces three new environment variables that need to live outside of the scope
of the regular argument loader path.
This will allow external applications to adjust FEX parameters in interesting ways.

FEX_APP_CONFIG: Allows you to override where Config.json lives
FEX_APP_CONFIG_LOCATION: Allows you to provide an entire folder of app profiles
FEX_APP_DATA_LOCATION: Allows you to override where the data gets stored and loaded from

Fixes #572
2021-04-05 04:09:41 -07:00
Ryan Houdek acb329b8e7 Adds use-after-scope to the ASAN options
Useful for seeing stack scoping problems that can crop up
2021-04-04 22:07:40 -07:00
Ryan Houdek 3020e39428 Remove DestSize from Print IR Op 2021-04-04 22:07:03 -07:00
Ryan Houdek 4a094c5dee Reenables new gcc unit tests that are no longer failing
Documents the few remaining tests with details as to why the test fails.
2021-04-04 22:05:55 -07:00
Ryan Houdek f8eabb4f33 Implements 64bit scalar VUMin on x86 JIT 2021-04-04 22:05:55 -07:00
Ryan Houdek 4da492d361 Fixes vector integer min and max on AArch64
These instructions don't have a 64bit element variant.
Switch the 64bit element variant over to a coded variant similar to float min/max
2021-04-04 22:05:55 -07:00
Ryan Houdek a200d2c91c Fixes packed shift ops
In the case of a shift value larger than the element size then the shift will zero the element.
Incoming shift amount is always a 64bit value which means that the value can become quite large on shifts.
Checks the incoming source and sets the maximum to be the element size, which works for AArch64's vector shifts
Which only operate on 8bit incoming shift amounts
2021-04-04 22:05:55 -07:00
Ryan Houdek f1663abe81 Removes getenv usage for aotir
We can't rely on `getenv("HOME")` always working.
Have the AOTIR code use the helper method for getting the data directory instead.
2021-04-04 21:43:14 -07:00
Ryan Houdek af205e9f5f Fixes PEXTRW
In the case of 16bit PEXTRW it will zext to 32bits which wasn't being done
2021-04-04 21:32:12 -07:00
Ryan Houdek 032d1fbba2 Fixes VFMIN/VFMax NaN behaviour
In a world with NaNs we were hitting some undefined behaviour with fmin/fmax
Needed to change these over to something more complex to handle the NaN case.

ARMv8.7 alternative FP mode would solve this for the scalar case, but it still doesn't work for the vecto case sadly.
2021-04-04 21:29:46 -07:00
Ryan Houdek 313c4bb784 Fixes bug in VectorImm
If passing in a 64bit immediate then movi doesn't behaviour how you would expect.

Instead to a tmp move + dup to implement it the way we want
2021-04-04 20:24:23 -07:00
Ryan Houdek ceb484fd3a Adds support for Arm64 JIT to Print
Useful for debugging, supports both vector and GPR
2021-04-04 20:22:44 -07:00
Ryan Houdek b98c7bd853 Moves AArch64's GetPhys to be a class function
No need for it to be static
2021-04-04 20:21:06 -07:00
Ryan Houdek 4c045148b6 Fixes memory leak in CPU backends
Dispatcher wasn't ever getting cleaned up.
Also ensures sharing it with the CompileService otherwise it crashes on use.
2021-04-04 20:19:35 -07:00
Ryan Houdek 9f29bba4aa Fixes missing virtual destructor on Dispatcher class 2021-04-04 20:10:58 -07:00
Ryan Houdek 2c592bd8a8 Implements a minimal siginfo_t in the dispatcher's signal handler
This lets Carrion get farther in game.
Seems like one of its threads hits sigsegv and it just...lets it die after checking siginfo?
2021-04-03 17:03:19 -07:00
Ryan Houdek 8319a85524 Merge pull request #946 from Sonicadvance1/fix_eventfd
Fixes eventfd and eventfd2
2021-04-03 06:52:32 -07:00
Ryan Houdek 8ad45071e4 Merge pull request #945 from Sonicadvance1/fix_truncate_creat
Fixes truncate and creat
2021-04-03 06:47:31 -07:00
Ryan Houdek 3379f3a8f9 Fixes eventfd and eventfd2
glibc function signature doesn't match kernel
2021-04-03 06:46:28 -07:00
Stefanos Kornilios Mitsis Poiitidis c76205040e Merge pull request #944 from Sonicadvance1/fix_clock_nanosleep
Fixes clock_nanosleep
2021-04-03 11:09:46 +03:00
Stefanos Kornilios Mitsis Poiitidis 648ef617cf Merge pull request #943 from Sonicadvance1/fix_fexconfig_unnamed_options
Fixes FEXConfig filling in unnamed options
2021-04-03 11:09:21 +03:00
Stefanos Kornilios Mitsis Poiitidis 5fac51f659 Merge pull request #942 from Sonicadvance1/implement_cmdline
Fixes wrapping cmdline arguments
2021-04-03 11:08:50 +03:00
Ryan Houdek 5d63dc179c Fixes truncate and creat
These were using the passthrough helper which doesn't was breaking the result value
2021-04-02 20:45:33 -07:00
Ryan Houdek 858566c356 Fixes clock_nanosleep
glibc version behaves slightly differently than kernel
2021-04-02 20:41:04 -07:00
Ryan Houdek fb94e62df1 Fixes FEXConfig filling in unnamed options
These aren't meant to be set by the user
2021-04-02 20:28:12 -07:00
Ryan Houdek 51e4fe7f0a Fixes wrapping cmdline arguments
In case an application checks this
2021-04-02 20:26:28 -07:00
Stefanos Kornilios Mitsis Poiitidis be03767c0b Merge pull request #941 from Sonicadvance1/pselect6_fix
Fixes pselect6 signal arg pack argument
2021-04-02 23:42:02 +03:00
Stefanos Kornilios Mitsis Poiitidis adeea56ea8 Merge pull request #940 from Sonicadvance1/update_release_script
Updates release script to match what I want it to be
2021-04-02 23:41:17 +03:00
Ryan Houdek 6923c47a70 Fixes pselect6 signal arg pack argument
This doesn't exactly match the glibc call so do the syscall directly
2021-04-02 13:23:07 -07:00
Ryan Houdek 26d9ddef9f Updates release script to match what I want it to be
Generates the current and previous tag automatically.
Checks to ensure the previous tag exists and that the current tag does not
2021-04-02 11:49:19 -07:00
Ryan Houdek c6c2555e1b Merge pull request #938 from Sonicadvance1/FEXConfig_improvements
FEXConfig improvements
2021-04-01 23:55:16 -07:00
Stefanos Kornilios Mitsis Poiitidis 283a5c1293 Merge pull request #937 from Sonicadvance1/fix_inverted_bool
Fixes inverted boolean FEXLoader arguments
2021-04-02 09:38:21 +03:00
Stefanos Kornilios Mitsis Poiitidis b31976d0e2 Merge pull request #936 from Sonicadvance1/fix_cp_config_error
Fixes Thunk Guest libs short arg. C&P error made it not be j
2021-04-02 09:36:53 +03:00
Ryan Houdek f922e6420b Allow opening the first configuration file passed in
Lets us easily open any configuration file
2021-04-01 19:25:17 -07:00
Ryan Houdek 7ff06c828f Loads default config if doesn't exist
Allows us to save a folder with the default if we want it
2021-04-01 19:22:53 -07:00
Ryan Houdek be16cc1456 Fixes window ordering so popup message is visible without config open
This was a child of the config popup before.
Moves it to a child of the workspace instead so that if we had a load error the message is still visible
2021-04-01 19:16:03 -07:00
Ryan Houdek 59df3d13cc Don't set up RootFS folder or inotify thread on config load failure
Can happen in the case that the default file tries to get loaded and doens't exist
2021-04-01 19:15:08 -07:00
Ryan Houdek 1adf518e6f Don't try creating inotify thread twice
If the inotify FD is already opened then just early exit
Can happen in the case that a config tries to get opened and didn't exist
2021-04-01 19:13:43 -07:00
Ryan Houdek 68171ad9b5 Creates RootFS data folder as a user convenience
To make sure users don't mess this up, create the folder if it isn't found
2021-04-01 19:12:56 -07:00
Ryan Houdek 78a923b5b0 Fixes inverted boolean FEXLoader arguments
The default_value doesn't have anything to do with the boolean argument being set.
Setting a boolean argument will always set true, inverted always sets false.
We don't currently have a need to support the inverted case where passing in an argument sets it to false
2021-04-01 19:10:29 -07:00
Ryan Houdek 0c6741438b Update config generator python to check for duplicate options
Checks for duplicate short and long option arguments
2021-04-01 19:08:19 -07:00
Ryan Houdek 5d0c080a9d Fixes Thunk Guest libs short arg. C&P error made it not be j 2021-04-01 19:08:00 -07:00
Ryan Houdek 5f49b57948 Merge pull request #928 from Sonicadvance1/pthreads_threading
Switches FEXCore over to pthreads implementation
2021-04-01 01:15:43 -07:00
Ryan Houdek 38778953a1 Merge pull request #924 from Sonicadvance1/select_named_rootfs
Adds support for Named RootFS folders in FEXConfig
2021-04-01 00:32:15 -07:00
Ryan Houdek 21a364c2ad Switches FEXCore over to pthreads implementation
This defaults to a pthread implementation but it can be switched over to a custom thread
handler if the frontend desires.
2021-04-01 00:31:45 -07:00
Ryan Houdek 3294cc209f Adds support for Named RootFS folders in FEXConfig
Provides a drop down dialog of RootFS folders to select.
Will be empty if the user doesn't have any folders in place.
Just allowing directory paths to be inserted

Fixes #689
2021-04-01 00:15:29 -07:00
Stefanos Kornilios Mitsis Poiitidis ae5d41e46b Merge pull request #926 from Sonicadvance1/disable_seccomp
Disables seccomp
2021-04-01 10:03:36 +03:00
Ryan Houdek 43f83fd6d2 Merge pull request #923 from Sonicadvance1/named_rootfs
Adds support for named rootfs configurations
2021-04-01 00:02:11 -07:00
Ryan Houdek 8cf47f8221 Adds support for named rootfs configurations
If the relative or absolute folder doesn't exist then FEX will search in the data folder for a rootFS named the same thing
If the path exists then it is used.

For example if my data folder contains `$HOME/.fex-emu/RootFS/Ubuntu_main/` and I set the RootFS option to `Ubuntu_main`
Then this rootfs will be used. Allows you to easily select a rootfs in that folder without having the full file path.
2021-03-31 23:46:07 -07:00
Ryan Houdek 43b6cc1cb9 Adds GetDataDirectory config helper
This ends up in either $HOME/.fex-emu/ or $XDG_DATA_HOME/.fex-emu/
2021-03-31 23:46:07 -07:00
Ryan Houdek 1ab6726498 Moves Config directory finding to FEXCore
This will be necessary in the next commit
2021-03-31 23:46:07 -07:00
Ryan Houdek c21acd0d45 Merge pull request #916 from Sonicadvance1/fix_execve_again
Fix execve again
2021-03-31 23:45:06 -07:00
Ryan Houdek e637751112 ptrace behaviour is now changed 2021-03-31 23:38:14 -07:00
Ryan Houdek fac6377ed3 Fixes a crash in openat 2021-03-31 23:37:57 -07:00
Ryan Houdek 4f3cb933cb Fixes execve 2021-03-31 23:37:56 -07:00
Ryan Houdek b2825bd848 Adds handler in ELFLoader to determine ELF type. 2021-03-31 23:37:56 -07:00
Ryan Houdek 83eb7f4ad3 Return EPERM in ptrace. A Feral launcher is attempting to use this to ensure it can't ptrace. 2021-03-31 23:37:56 -07:00
Ryan Houdek c3854c211a Fix finding home directory when we have no environment variables. 2021-03-31 23:37:56 -07:00
Ryan Houdek 61cd3eb3ce Fix crash in ELFDB if it tried loading something that wasn't an ELF 2021-03-31 23:37:56 -07:00
Ryan Houdek 490352f568 Merge pull request #921 from Sonicadvance1/update_imgui
Updates external imgui which fixes keypad enter in FEXConfig
2021-03-31 23:24:42 -07:00
Ryan Houdek 4004d5a3b7 Disables seccomp
FEX doesn't support seccomp in userspace and allowing these through causes chromium secure sandbox to break.
Disabling these with EINVAL allows FEX to behave as if seccomp isn't enabled in the kernel config

Necessary to get the Civ6 launcher further
2021-03-31 23:23:40 -07:00
Ryan Houdek 761447467e Updates external imgui which fixes keypad enter in FEXConfig 2021-03-30 15:39:11 -07:00
Ryan Houdek 3f38ed94ca Merge pull request #922 from Sonicadvance1/disable_oomscore
Disables gvisor test proc_pid_oomscore_test
2021-03-30 15:38:33 -07:00
Ryan Houdek ad877d4088 Disables gvisor test proc_pid_oomscore_test
Depending on runner this passes or fails.
The test expects to be able to open `/proc/self/oom_score_adj` as writable to adjust the oom score.
This is expected to fail on all three of our runners, but sometimes the solidrun board manages to open it.
Disable as it is a flake. Could be a kernel bug on the solid run board
2021-03-30 15:29:52 -07:00
Ryan Houdek 92fc6b1909 Merge pull request #920 from lioncash/insert
Config: Make lookup map assignment match behavior of comment
2021-03-30 14:17:37 -07:00
Lioncache 4b11dcb72a Config: Make lookup map assignment match behavior of comment
emplace() only performs an insertion or assignment if the key doesn't
already exist within the map, which is at odds with what the comment
above the line indicates should happen.

insert_or_assign() better models what the comment indicates.
2021-03-30 11:51:52 -04:00
Lioncache dd7c78dcd7 Config: Perform lookups by character overloads where applicable
These are marginally better to perform than the string equivalents
2021-03-30 11:41:57 -04:00
Lioncache a20ef403c1 Config: std::move strings within MergeEnvironmentVariables()
Same behavior, minus potential heap duplicate allocations
2021-03-30 11:40:42 -04:00
Ryan Houdek be06511d3d Merge pull request #919 from FEX-Emu/skmp/fix-build
Build: Fix for ubuntu 20.04
2021-03-30 03:20:43 -07:00
Stefanos Kornilios Mitsis Poiitidis b7a60337e8 Build: Fix for ubuntu 20.04 2021-03-30 13:08:19 +03:00
Stefanos Kornilios Mitsis Poiitidis 315b8e9c1e Merge pull request #917 from FEX-Emu/skmp/lightweight-docs-3
Docs: Add tags, unittest Readme.md, generated SourceOutline.md
2021-03-30 12:52:18 +03:00
Ryan Houdek 155a2b0194 Merge pull request #918 from lioncash/fallthrough
CMakeLists: Flag unannotated implicit fallthrough as an error
2021-03-30 02:37:07 -07:00
Stefanos Kornilios Mitsis Poiitidis 8c874f4540 Docs: Commit generated SourceOutline.md, update Readme.md 2021-03-30 12:21:18 +03:00
Stefanos Kornilios Mitsis Poiitidis 942b8549c6 Docs: Add tags to the source code 2021-03-30 12:21:18 +03:00
Ryan Houdek fa577b4527 Merge pull request #910 from FEX-Emu/skmp/lightweight-docs
Docs: lightweight doc autogen
2021-03-30 02:17:41 -07:00
Stefanos Kornilios Mitsis Poiitidis 15537f8c8c Scripts: Also add unittests folder for doc outline 2021-03-30 11:58:18 +03:00
Lioncache 434a74a88a CMakeLists: Flag unannotated implicit fallthrough as an error
Prevents a class of sneaky logic bugs from slipping through into the
codebase.

This also resolves a case of such a bug within the Decorder's ReadData()
where all 3 byte reads would be performed as if they were a 4 byte read.
2021-03-30 04:56:12 -04:00
Stefanos Kornilios Mitsis Poiitidis 1fdd6fbb14 Scripts: Add a documentation comment in changelog_generator.py 2021-03-30 11:55:39 +03:00
Stefanos Kornilios Mitsis Poiitidis 458bebf598 Scripts: Add a documentation comment in doc_outline_generator.py 2021-03-30 11:53:11 +03:00
Stefanos Kornilios Mitsis Poiitidis 2b10b9792b Scripts: Fix generate_release, don't use markdown for changelogs 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 0dd02da57c Docs: Update outline scripts 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 309139e203 Scripts: Add generate_release script 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 350c33ada4 Docs: Update outline script 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis bdd35e5743 docs: Update outline generator 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis ceb7082e37 Docs/Versioning: Add Changelog Generator 2021-03-30 11:23:00 +03:00
Stefanos Kornilios Mitsis Poiitidis 2d6dc80039 Docs: Adds tag-based outline generator w/ glossary support 2021-03-30 11:22:56 +03:00
Ryan Houdek c92df21627 Merge pull request #915 from lioncash/fallthrough
OpcodeDispatcher: Add missing break for UD2 in INTOp()
2021-03-29 22:56:19 -07:00
Lioncache 539de0492e OpcodeDispatcher: Add missing break for UD2 in INTOp()
Fixes a case where the reason gets overwritten.
2021-03-29 23:42:26 -04:00
Ryan Houdek efd41de8ea Merge pull request #914 from lioncash/optional
Context: Make use of std::optional with GetFilenameHash()
2021-03-29 16:44:08 -07:00
Lioncache 99ae862875 Context: Eliminate undefined behavior in LoadEntryList()
This breaks the strict-aliasing rule. We can use std::memcpy here to
make this code well-defined.
2021-03-29 12:17:36 -04:00
Lioncache b81ea43601 Context: Make use of std::optional with GetFilenameHash()
Same behavior, minus the need for an out parameter. While we're at it,
we can tidy up some of the file handling code.
2021-03-29 12:17:32 -04:00
Ryan Houdek fcd4974cff Merge pull request #912 from Sonicadvance1/config_docs_improvements
Configuration option improvements
2021-03-29 04:47:45 -07:00
Stefanos Kornilios Mitsis Poiitidis 8550c9f2dd Merge pull request #913 from Sonicadvance1/fix_32bit_mask
Adds a 32-bit mask for multiblock RIP calculation
2021-03-29 14:35:32 +03:00
Ryan Houdek 15dd70487f Removes legacy difference between FEX config env variable and enum names 2021-03-29 00:26:11 -07:00
Ryan Houdek 8fad0c8fdb Adds a 32-bit mask for multiblock RIP calculation
In the case of overflow then it'll just mask this result.
The rest of the logic already matches what 32-bit expects.

Fixes #703
2021-03-28 23:39:53 -07:00
Ryan Houdek b6c0a9f2a0 Configuration option improvements
This commit does three things that are bounded to each other
1) Moves configs from ConfigValues.inl to Config.json
2) Uses the json to generate a man file with the options inside of it
3) Generates the code required for FEXLoader to automatically parse defined options

Moving the configuration options to a parseable format was required to generate the man pages.
Can't really include an inl file in the man page
Man page gets generated and installed through the regular cmake install process
`man FEX` to get the man page

With this change, FEXLoader's argument parsing now will automatically be generated from the json.
This way whatever is in the json file matches what is in the man page, in the json, AND what is returned in FEXLoader --help.
It will stay in sync now.

FEXConfig is the only application that stays out of sync for now as it requires some more thought to plan out.

A minor improvement that this brings as well is that every boolean option that has a long argument also gains the inversion of that property.
aka, `--gdb` also gains `--no-gdb` The use case for this is minor but boolean arguments should always allow negated variants.
2021-03-28 23:28:41 -07:00
Ryan Houdek b5a6dd031b Merge pull request #911 from Sonicadvance1/std_size_life
Replaces FEX's usage of classic C array element count with std::size
2021-03-28 05:58:55 -07:00
Ryan Houdek 650877adca Replaces FEX's usage of classic C array element count with std::size
Just learned that this is a c++17 feature and it's a nice little
cleanup.

Only one remains in our codebase which can't easily be replaced
2021-03-28 05:39:35 -07:00
Ryan Houdek 432b6fccf2 Merge pull request #907 from Sonicadvance1/replace_glfw
Replaces FEXConfig usage of GLFW with SDL2
2021-03-27 05:53:00 -07:00
Ryan Houdek 47f3c983ee Merge pull request #905 from Sonicadvance1/removes_warnings_aarch64
Removes remaining warnings from AArch64 JIT.
2021-03-27 05:52:51 -07:00
Ryan Houdek 9565247ada Merge pull request #904 from Sonicadvance1/stop_compiling_twice
Stop compiling FEXCore twice
2021-03-27 05:52:39 -07:00
Ryan Houdek c59b9fbede Merge pull request #903 from Sonicadvance1/fix_silent_logs
Disables silent logging on unit tests
2021-03-27 05:52:28 -07:00
Ryan Houdek 7c96893579 Updates README with SDL2 requirement 2021-03-27 05:26:00 -07:00
Ryan Houdek 8ebef457fc Replaces FEXConfig usage of GLFW with SDL2
Sometimes GLFW3 can just consume a full CPU core, spinning on nothing.

Fixes #906
2021-03-27 05:23:58 -07:00
Ryan Houdek 9eee879be0 Removes remaining warnings from AArch64 JIT.
There are a couple of warnings remaining but need restructuring to solve.
I believe @phire is going to fix the one warning about a virtual destructor in a upcoming PR.
Fixes #679
2021-03-26 18:49:40 -07:00
Ryan Houdek 64affa8c8e Stop compiling FEXCore twice
This now creates an object binary then static and shared libraries from that.
This stops us outputting warnings twice in FEXCore.
Also improves compilation time even with ccache enabled.
Went from 12s to 8s compilation WITH ccache.
I'm sure without ccache it is even better.
2021-03-26 18:27:23 -07:00
Ryan Houdek dc041bdf0e Disables silent logging on unit tests
We need these for our CI artifacts
2021-03-26 18:04:17 -07:00
Ryan Houdek 51c43a0761 Adds option to disable silent logging 2021-03-26 18:00:23 -07:00
Ryan Houdek aa67142e6f Merge pull request #902 from Azkali/patch-1
Dockerfile add arm64 compatibility
2021-03-25 22:36:41 -07:00
The Great Wizard Azkali ae569da895 Update Dockerfile 2021-03-26 06:30:17 +01:00
The Great Wizard Azkali d77acdf474 Dockerfile add arm64 compatibility
Python3, Python3-dev and linux-headers prevent from building on arm64
2021-03-26 02:18:13 +01:00
Ryan Houdek 264ec44276 Merge pull request #892 from Sonicadvance1/install_thunks
Allows installing of FEXThunks in our data directory
2021-03-25 14:50:54 -07:00
Ryan Houdek da49eb3394 Allows installing of FEXThunks in our data directory
This is necessary for building FEX packages that contain some initial thunk libs.
Gives an initial foothold for a default location for the host and guest thunk folders
2021-03-24 03:27:38 -07:00
Ryan Houdek 5447ec3ec8 Merge pull request #884 from Sonicadvance1/fix_empty_string_crash
Fixes crash on empty path config string
2021-03-24 02:34:20 -07:00
Ryan Houdek ff3974c4a1 Merge pull request #890 from Sonicadvance1/merged_environment_variables
Fixes environment variable configuration from multiple layers
2021-03-24 02:34:14 -07:00
Ryan Houdek 358f12e074 Merge pull request #891 from Sonicadvance1/installed_global_app_configs
Implements support for installing global application profiles
2021-03-24 02:34:00 -07:00
Ryan Houdek c5648ac84c Implements support for installing global application profiles
These need to be used sparingly, we don't want to proactively crush user options.
2021-03-23 22:34:01 -07:00
Ryan Houdek 30e1f871ad Fixes environment variable configuration from multiple layers
This was a missing feature that I had skipped previously.
Before this commit, each layer would overwrite all previous environment variables if they had anything defined.

After this commit, the layers will now merge in priority order.
Meaning if a higher priority layer has the same environment variable defined, it will overwrite that specific variable
Ending up with a superset of all the layer's environment variables now.
2021-03-23 22:01:09 -07:00
Ryan Houdek 5c624684b7 Merge pull request #887 from Sonicadvance1/FEXBashLoader
Adds a FEXBash helper program
2021-03-23 21:37:45 -07:00
Ryan Houdek ef11b534ef Adds a FEXBash helper program
This is a helper program to execute a bash command through FEX.
Works around an edge case of
eg:
FEXInterpreter /bin/sh -c "echo A"
versus
FEXBash "echo A"
Argument expansion ends up being a pain point
This isn't currently used but will be very shortly
2021-03-23 21:27:30 -07:00
Ryan Houdek cd32eeb428 Adds FEXInterpreter as a build target
Makes it easier to have FEXInterpreter around
2021-03-23 21:27:29 -07:00
Ryan Houdek 43984c810d Merge pull request #886 from Sonicadvance1/namespace_perm
Returns EPERM on clone with namespace
2021-03-23 20:56:20 -07:00
Ryan Houdek a50689aa8b Merge pull request #885 from Sonicadvance1/cpuid_leaf
Adds support for CPUID leafs
2021-03-23 20:55:33 -07:00
Scott Mansell 311d6d4385 Merge pull request #883 from phire/thread_cleanup
Cleanup threads when they exit
2021-03-24 16:22:15 +13:00
Ryan Houdek a19c59f5b2 Merge pull request #888 from Sonicadvance1/default_silentlogs
Default to silent logging
2021-03-23 20:13:07 -07:00
Scott Mansell 6a8f68022f Join threads after fork.
Also, add some comments to document things and
rename function to be clear about it's purpose.
2021-03-24 15:41:09 +13:00
Ryan Houdek fdcb6206da Default to silent logging
Useful for debugging but we need to be silent by default now.
Without this anything that checks stdout output of applications from bash would fail

Steam with lspci, lsusb, uname, etc
2021-03-23 19:30:42 -07:00
Ryan Houdek 6ea977ba34 Remove proc_pid_uid_gid_map from known failures list
Now instead of crashing with namespacing, it skips the tests if it gets EPERM.
2021-03-23 19:28:12 -07:00
Ryan Houdek 9b35cb4408 Returns EPERM on clone with namespace
Notably this allows applications to work that don't require the namespace but check up front if they are able to clone with it.
Civ 6's launcher checks this as an example
2021-03-23 19:28:06 -07:00
Ryan Houdek afa869e1f2 Adds Config.h generated file
Gives us the install path and FEXInterpreter locations
2021-03-23 19:12:18 -07:00
Ryan Houdek 43ddba4e82 Adds support for CPUID leafs
Leafs come from ECX but only some CPUID functions support this.
This adds the initial infrastructure but doesn't yet add support for the CPUID functions to consume the leaf.
2021-03-23 18:13:14 -07:00
Ryan Houdek e3380957ba Fixes crash on empty path config string
If the config exists as an empty string then this would crash
2021-03-23 18:02:56 -07:00
Scott Mansell 21687f4dc5 Also destroy the parent thread 2021-03-24 03:14:23 +13:00
Scott Mansell 41a1a9d400 Cleanup threads when they exit 2021-03-24 03:14:23 +13:00
Stefanos Kornilios Mitsis Poiitidis 82c2b49a2e Merge pull request #878 from Sonicadvance1/fexconfig_actual_default
Changes FEXConfig to actually load default configuration
2021-03-22 18:24:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fbc49c1648 Merge pull request #876 from Sonicadvance1/QOL_config
Implements a couple quality of life configuration option handling
2021-03-22 18:23:52 +02:00
Stefanos Kornilios Mitsis Poiitidis d2f636893d Merge pull request #880 from Sonicadvance1/github_label_testing
Support disabling unit tests based on runner label
2021-03-22 18:18:08 +02:00
Ryan Houdek e5a9bd4d3e Fixes incorrect result on cmpxchg reg, reg failure when regsize = 32bit 2021-03-22 09:03:48 -07:00
Ryan Houdek ca24ea1d36 Fixes missing HandledLock on XCHG 2021-03-22 09:03:48 -07:00
Ryan Houdek 998e53aa70 Only disabled unaligned atomics test on ARMv8.0 2021-03-22 09:03:48 -07:00
Ryan Houdek de890e7387 Have unit tests check for runner label 2021-03-22 09:03:48 -07:00
Ryan Houdek 042a71be96 Get the runner's label for testing disabled tests later 2021-03-22 09:03:48 -07:00
Ryan Houdek 0a61741596 special case logging option stdout/stderr 2021-03-21 02:26:56 -07:00
Ryan Houdek 91a5a625ac Implements a couple quality of life configuration option handling
This takes the changes from #829 and moves it to the correct location to be picked up from any loader.
This also fixes #873.
Expands the config paths. Anything that has ~ or is relative will be converted to an absolute path.
2021-03-21 02:25:30 -07:00
Stefanos Kornilios Mitsis Poiitidis f68593ea97 Merge pull request #877 from Sonicadvance1/fix_absolute_linker
Adds some additional logic for finding absolute soft linked linker
2021-03-21 11:08:46 +02:00
Stefanos Kornilios Mitsis Poiitidis e15730782b Merge pull request #879 from Sonicadvance1/update_readme
Updates readme with some more direct information
2021-03-21 11:07:35 +02:00
Ryan Houdek fc27893f20 Updates readme with some more direct information
Some of this taken from Stef's recent changes. Some taken from the wiki.
combination of the two best bits
2021-03-21 01:53:51 -07:00
Ryan Houdek 5b8bf8dd24 Adds some additional logic for finding absolute soft linked linker
Ubuntu's x86_64 rootfs does a soft link from `/lib64/ld-linux-x86-64.so.2` to an absolute path of `/lib/x86_64-linux-gnu/ld-2.32.so`
Our frontend ELF loader logic would resolve this symlink to be just `/lib/x86_64-linux-gnu/ld-2.32.so` which would then fail.
Adds some additional logic to our frontend that if the symlink ends up being a softlink to an absolute address then it'll first resolve that symlink.
Then combines rootfs + absolute path. If that still fails then it'll fall down the regular linker path, which could still find the linker on the host side.

This is all a bit of a kludge to just work around Ubuntu doing an absolute address rather than a relative one.

Fixes #869
2021-03-21 01:11:47 -07:00
Ryan Houdek 0c48f1409e Changes FEXConfig to actually load default configuration
This was previously just some default values that I put in as placeholder
Now that we actually have defaults configured somewhere, use those.

Fixed #874
2021-03-21 00:48:51 -07:00
Ryan Houdek 8c1f2eb0b3 Merge pull request #870 from FEX-Emu/skmp/lock-validation
OpDisp: Validate LOCK handling, add missing segment offsets
2021-03-20 15:53:06 -07:00
Ryan Houdek 07613456d2 Merge pull request #872 from FEX-Emu/skmp/x87-fixes
X87: Init on X87FNSAVE, fix FNINIT
2021-03-20 15:52:54 -07:00
Ryan Houdek 493bb3bef3 Merge pull request #875 from phire/test-no-multibock
Explictly as for --no-multiblock in tests
2021-03-20 15:52:44 -07:00
Scott Mansell cf8572e090 Explictly as for --no-multiblock in tests
We switched the default over a while back, so we haven't
been getting test coverage with multiblock off
2021-03-21 04:57:47 +13:00
Stefanos Kornilios Mitsis Poiitidis d120f12c93 X87: Init on X87FNSAVE, fix FNINIT 2021-03-19 17:04:51 +02:00
Stefanos Kornilios Mitsis Poiitidis 695322d446 OpDisp: Validate LOCK handling, add missing segment offsets 2021-03-19 16:32:29 +02:00
Stefanos Kornilios Mitsis Poiitidis d34cde12ae Merge pull request #850 from Sonicadvance1/struct_verifier
libclang based Struct verifier written in python
2021-03-19 10:06:34 +02:00
Stefanos Kornilios Mitsis Poiitidis a7dc9d00d2 Merge pull request #864 from Sonicadvance1/cpuid_15h
Implements CPUID 15h
2021-03-19 10:02:02 +02:00
Ryan Houdek dbde2e400e Merge pull request #863 from Sonicadvance1/enable_invariant_tsc
Enables Invariant TSC CPUID bit
2021-03-18 12:39:08 -07:00
Ryan Houdek bdeefa7bf0 Merge pull request #861 from Sonicadvance1/add_version_config
Implements a --version argument
2021-03-18 12:39:00 -07:00
Ryan Houdek 51e513df14 Merge pull request #865 from FEX-Emu/skmp/fix-emufiles-lock
EmulatedFiles: Lock around FDToNameMap accesses
2021-03-18 04:11:58 -07:00
Stefanos Kornilios Mitsis Poiitidis 0106b362d6 EmulatedFiles: Lock around FDToNameMap accesses 2021-03-18 12:45:27 +02:00
Ryan Houdek c7f265158f Implements CPUID 15h
ARMv8 allows you query the cycle counter register very from userspace which allows us to emulate this easily.

My IceLake device returns 1.5Ghz through this interface
My AMD Zen+ device doesn't support this but with measurements has 3Ghz TSC
My Snapdragon 865 device has a TSC frequency of 19.20Mhz
2021-03-18 01:42:44 -07:00
Ryan Houdek fbf5325bdd Enables Invariant TSC CPUID bit
I keep forgetting to enable this bit. Fixes #855
2021-03-18 00:15:48 -07:00
Ryan Houdek 9b2bd8df28 Implements a --version argument
Nicer for the user to determine what their version is
2021-03-17 00:00:57 -07:00
Ryan Houdek 2a074204bc Merge pull request #859 from Sonicadvance1/improve_fexinterpreter_msg
Adds an installer message to the FEXInterpreter install
2021-03-16 23:13:12 -07:00
Ryan Houdek d79fb37135 Merge pull request #858 from Sonicadvance1/remove_testharness
Removes TestHarness
2021-03-16 23:13:03 -07:00
Ryan Houdek d01b40c5aa Adds an installer message to the FEXInterpreter install
Otherwise it doesn't feel like FEXInterpreter is installing
2021-03-16 22:36:58 -07:00
Ryan Houdek e871a43387 Merge pull request #857 from phire/deletething
Remove old and unused HostCore
2021-03-16 22:30:30 -07:00
Ryan Houdek ff45b37904 Merge pull request #856 from Sonicadvance1/remove_libcap
Removes libcap-dev dependency
2021-03-16 22:30:21 -07:00
Ryan Houdek 62aab57ef7 Removes TestHarness
This has been completely deprecated for favour of the TestHarnessRunner instead.
Is unused.
2021-03-16 22:14:49 -07:00
Scott Mansell 208fa0d1fc Remove old and unused HostCore
The thing we are actually using is Source/CommonCore/HostFactory.cpp
2021-03-17 18:09:49 +13:00
Ryan Houdek 634d4fead0 Removes libcap-dev dependency
We are only using this for syscalls getcap and setcap.
These are only passthrough pointers and don't need to be handled from a library.

Noticed this while writing user documentation
2021-03-16 22:08:15 -07:00
Ryan Houdek 0bdddedfe6 Merge pull request #853 from FEX-Emu/skmp/dont-lse-though-syscalls
RCLSE: Invalidate around OP_SYSCALLs, Syscalls might read the context
2021-03-16 17:54:54 -07:00
Stefanos Kornilios Mitsis Poiitidis 971740991b RCLSE: Invalidate around OP_SYSCALLs, Syscalls might read the context and need the latest version of it
Fixes steam w/ all optimization passes enabled
2021-03-17 02:42:32 +02:00
Ryan Houdek cb672e035c Merge pull request #852 from Sonicadvance1/fix_fexconfig_idle
Fixes FEXConfig trying to update full refresh
2021-03-16 00:53:40 -07:00
Ryan Houdek 4f64ba582c Patch up the remaining 32bit structs 2021-03-15 15:51:59 -07:00
Ryan Houdek fbe2583a04 Fix truncating to not set the log files to 21MB 2021-03-15 15:24:49 -07:00
Ryan Houdek befe9dcbae Adds struct verifier to github yaml workflow file
This way CI tests this on each commit
2021-03-15 15:24:49 -07:00
Ryan Houdek aacb0f6891 Adds new struct_verifier ctest to cmake
Currently only testing 32bit syscall struct definitions
2021-03-15 15:24:49 -07:00
Ryan Houdek 040cc746e7 Fixes FEXConfig trying to update full refresh
This burns a decent amount of CPU time just idling and we don't have active screen elements to matter here
It's still responsive because it'll update as the window gets any events
2021-03-15 09:46:49 -07:00
Ryan Houdek e0c5840f2c Adds libclang based struct verifier
This requires multiarch on the targets to work.
Will run a header through multiple architectures and ensure that the struct packing all works
2021-03-15 06:58:06 -07:00
Ryan Houdek 48027f1d2e Moves BUILD_TESTS check up the list
This way Source/ can have its own tests
2021-03-15 06:58:06 -07:00
Ryan Houdek 9dc0717cdb Attributes a few 32bit syscall types
This will be necessary for struct verification in the next commit
rusage needed to be updated to match the real rusage. Has unions with two named types
2021-03-15 06:58:06 -07:00
Ryan Houdek 8de7dd8fd0 Merge pull request #848 from Sonicadvance1/remove_warnings
Remove most warnings from FEX
2021-03-15 06:57:42 -07:00
Ryan Houdek 79e9477b14 Removes warnings from RAPass 2021-03-14 13:27:40 -07:00
Ryan Houdek 0b8f29000e Removes warnings from ConstProp 2021-03-14 13:27:39 -07:00
Ryan Houdek e530e3676f Removes warnings from IRDumper 2021-03-14 13:27:39 -07:00
Ryan Houdek 0c5e5dbdf2 Removes warnings from OpcodeDispatcher 2021-03-14 13:27:39 -07:00
Ryan Houdek d8591a8f14 Removes warnings from x86 JITCore 2021-03-14 13:27:39 -07:00
Ryan Houdek 941107120c Removes warnings from x86 BranchOps 2021-03-14 13:27:39 -07:00
Ryan Houdek ffc17c90c5 Removes warnings from InterpreterOps 2021-03-14 13:27:39 -07:00
Ryan Houdek 5285f4baa6 Adds CMake option to enable -Werror
We aren't error free so can't be default enabled
2021-03-14 13:27:39 -07:00
Ryan Houdek 037781c4d0 Merge pull request #849 from Sonicadvance1/fix_kernel_version
Fixes uname version and wraps /proc/version
2021-03-14 11:09:05 -07:00
Scott Mansell 4d2221a456 Merge pull request #831 from phire/one_dispatcher_to_rule_them_all
Unify all four dispatchers
2021-03-15 05:42:21 +13:00
Scott Mansell 466dc03c19 Remove stray cmake changes 2021-03-15 04:49:09 +13:00
Scott Mansell 5e1e09a43d Fix review issues 2021-03-15 04:41:14 +13:00
Ryan Houdek 3a0b475cfa Wrap /proc/version where we were leaking host kernel information
drops the correct FEX version in to the string as well
2021-03-12 20:29:35 -08:00
Ryan Houdek 4ece56ac4d Fixes uname version string
This wasn't following the correct format and it never actually had a working FEX_VERSION define.
Now it pulls in the GIT_DESCRIBE_STRING and brings in the correct date + time format
2021-03-12 20:29:35 -08:00
Ryan Houdek ea7af248e8 Ensure programs can pull in git_version.h 2021-03-12 20:29:35 -08:00
Scott Mansell 42b186a76c Don't generate a default vixl::codebuffer 2021-03-12 12:47:29 +13:00
Scott Mansell bdca7109a4 Force XBYAK64
Because XBYAK's autoconfig doesn't quite do the right thing
2021-03-11 23:46:07 +13:00
Scott Mansell 074bedf25c Automatically use correct ContextBackup type 2021-03-11 23:46:07 +13:00
Scott Mansell bdb199b6d6 Unify all four dispatchers 2021-03-11 23:46:07 +13:00
Stefanos Kornilios Mitsis Poiitidis e9afabc4cb Merge pull request #827 from FEX-Emu/skmp/thunks-json
Thunks: Add Thunk json, thunk guest folder
2021-03-11 11:05:23 +02:00
Ryan Houdek da1532f3b8 Merge pull request #836 from Sonicadvance1/fexcore_config_nodefault_value
Adds config value option constructor without default initializer
2021-03-11 01:04:58 -08:00
Ryan Houdek c73d2b2bca Merge pull request #843 from Sonicadvance1/locked_not
Adds support for locked NOT
2021-03-11 00:47:01 -08:00
Ryan Houdek da43db8787 Merge pull request #842 from Sonicadvance1/locked_adc_sbb
Adds support for locked ADC and SBB
2021-03-11 00:46:53 -08:00
Ryan Houdek ab8aff6e0d Adds unittests for locked not 2021-03-11 00:38:44 -08:00
Ryan Houdek 8a7caa82b2 Adds support for locked NOT
This was trivial
2021-03-11 00:38:12 -08:00
Ryan Houdek 2706c9aaae Adds unittests for locked ADC and SBB 2021-03-11 00:24:03 -08:00
Ryan Houdek 96f2a461f8 Adds support for locked ADC and SBB
These were fairly straightforward
2021-03-11 00:22:35 -08:00
Ryan Houdek cf966eb2e0 Adds config value option constructor without default initializer
This has the expected behaviour that the configuration option will be available and not default initialized.
To enforce this fact, it will assert if it tries to get a value that doesn't exist yet.
2021-03-10 23:52:19 -08:00
Ryan Houdek 6b860cd68e Merge pull request #839 from Sonicadvance1/more_32bit_syscalls
Adds a couple new 32bit syscalls found while running Steam
2021-03-10 21:34:31 -08:00
Ryan Houdek 6b4ccc46a8 Merge pull request #838 from Sonicadvance1/fix_cmpxchg_32bit_reg
Fixes an edge case of 32bit cmpxchg <reg>, <reg>
2021-03-10 21:34:24 -08:00
Ryan Houdek 8acb18b2b9 Merge pull request #837 from Sonicadvance1/fix_logmanager
Fixes crash with large strings through LogManager
2021-03-10 21:34:17 -08:00
Ryan Houdek 832a634f64 Merge pull request #841 from phire/vixl_cmake_improvements
Improve vixl cmakefiles
2021-03-10 21:20:57 -08:00
Scott Mansell ec196b47e2 Improve Vixl cmakefiles 2021-03-11 16:43:55 +13:00
Ryan Houdek 4dc2cf2b32 Adds more cmpxchg <reg>, <reg> unit tests
This will catch the previous fix
2021-03-10 17:47:54 -08:00
Ryan Houdek e264194453 Fixes an edge case of 32bit cmpxchg <reg>, <reg>
The comparison was using the wrong values which means this cmpxchg would fail
2021-03-10 17:47:54 -08:00
Ryan Houdek ddce28df11 Merge pull request #834 from Sonicadvance1/cleanup_clang_tidy_iwyu
Moves clang-tidy arguments to root cmakelists
2021-03-10 17:47:19 -08:00
Ryan Houdek fc5bf5d261 Merge pull request #833 from Sonicadvance1/add_perf_warning
Adds x86-64 host performance warning
2021-03-10 17:47:07 -08:00
Ryan Houdek 45e5e58c7e Merge pull request #832 from Sonicadvance1/cleanup_error_interpreter
Extends error message about not being able to find guest interpreter
2021-03-10 17:46:52 -08:00
Ryan Houdek 0dc7fd75e1 Adds a couple new 32bit syscalls found while running Steam
These are straightforward to implement. Only a few dozen more syscalls missing on 32bit
2021-03-10 17:43:21 -08:00
Ryan Houdek 10fbf60259 Converts the printf usages in FEXCore to LogManager
Now that we can safely handle long strings these will now work
2021-03-10 16:38:46 -08:00
Ryan Houdek 4e55f29589 Fixes crash with large strings through LogManager
va_list types have an edge where after operating on it, then it is undefined behaviour to do anything other than va_end on the object.
So in order to call vsnprintf on it multiple times we must create copies with va_copy.
Easy enough to work around by doing a few more copies of the va_list
2021-03-10 16:36:33 -08:00
Ryan Houdek 1ddfa3e7ed Moves clang-tidy arguments to root cmakelists
This way we don't need to redeclare the arguments twice
Also moves IWYU lower so it doesn't hit any external projects other than FEXCore

Still not running these since everything needs to be cleaned up anyway
2021-03-10 12:38:16 -08:00
Ryan Houdek ec3883181b Adds x86-64 host performance warning
Lets users know that x86-64 isn't our optimal target but can still be used if passed in a new cmake argument.
Easy enough just pass in -DENABLE_X86_HOST_DEBUG=True to cmake.

Closes #776
2021-03-10 11:36:19 -08:00
Ryan Houdek e2ec645855 Extends error message about not being able to find guest interpreter
Passes the result back up to the frontend as well which allows us to early exit correctly.
Also ensures that we return ENOEXEC on these error cases so if someone is waiting on a return value, they don't just get zero
Fixes #757
2021-03-10 11:11:36 -08:00
Stefanos Kornilios Mitsis Poiitidis f0d3004363 Thunks: Add Thunk json, thunk guest folder 2021-03-10 15:51:34 +02:00
Stefanos Kornilios Mitsis Poiitidis e708b83dbb Merge pull request #826 from Sonicadvance1/disable_rcpc
Disables RCPC on ARM64 JIT
2021-03-10 12:31:17 +02:00
Ryan Houdek a9d3851684 Disables RCPC on ARM64 JIT
This is currently bugged on Snapdragon 865.
Seems to only affect its prime core. Smells like errata.
2021-03-10 02:13:16 -08:00
Ryan Houdek 7e4ba77d72 Merge pull request #825 from FEX-Emu/skmp/fix-sigchld
Signals: SA_NOCLDSTOP only blocks CLD_CONTINUED/STOPPED/TRAPPED
2021-03-10 01:50:04 -08:00
Stefanos Kornilios Mitsis Poiitidis 449ea3e80b Signals: SA_NOCLDSTOP only blocks CLD_CONTINUED/STOPPED/TRAPPED 2021-03-10 11:25:40 +02:00
Stefanos Kornilios Mitsis Poiitidis 8fd91a5b7a Merge pull request #821 from Sonicadvance1/update_vixl
Updates vixl submodule
2021-03-10 09:12:40 +02:00
Ryan Houdek e3d6db2eb5 Merge pull request #823 from FEX-Emu/skmp/fix-secondaryalu-atomics
OpDisp: Add atomic logic for SecondaryALUOp
2021-03-09 18:15:18 -08:00
Stefanos Kornilios Mitsis Poiitidis d930de30ef OpDisp: Add atomic logic for SecondaryALUOp 2021-03-09 18:32:13 +02:00
Scott Mansell b36378369a Merge pull request #822 from FEX-Emu/skmp/fix-scm-bug
ConstProp: Restrict imm code motion around selects to matching sizes, fixes dav1d
2021-03-09 22:12:50 +13:00
Stefanos Kornilios Mitsis Poiitidis cbf4c9a687 ConstProp: Restrict imm code motion around selects to matching sizes, fixes dav1d 2021-03-09 11:06:16 +02:00
Ryan Houdek bcebe7d47a Updates vixl submodule
Works around the no enum enum conversion warnings
2021-03-08 23:57:58 -08:00
Ryan Houdek e9b057bc8b Merge pull request #817 from phire/SeparateThreadAndState
Separate thread and state
2021-03-08 23:41:56 -08:00
Scott Mansell afda855caf disptacher fixes/cleanups 2021-03-09 20:28:53 +13:00
Scott Mansell 6ed36e3304 Refactor syscalls to take Frame
The vast majoirty of syscalls don't need anything in thread or frame.
So lets save an indirection for all those syscalls.

Most of the syscalls which do need Thread (or CTX via
Thread are in Thread.cpp or Memory.cpp
These have all been modifiy to fetch Thread from Frame
2021-03-09 20:28:50 +13:00
Scott Mansell 35c01038ba Seperate per-thread and per-frame state 2021-03-09 20:27:44 +13:00
Scott Mansell 43a087d554 Allow both interpeter dispatchers to be built
This commit doesn't include cmakefile changes to actually do this.
2021-03-09 20:25:14 +13:00
Ryan Houdek a1fed323df Merge pull request #820 from Sonicadvance1/remove_config_duplication
Deduplicates some configuration data
2021-03-08 23:20:04 -08:00
Ryan Houdek 0bcd1c7df5 Merge pull request #816 from Sonicadvance1/cpuid_80000005
Implements CPUID 0x8000'0005 for L1 cacheline information
2021-03-08 23:15:03 -08:00
Ryan Houdek a3ba1ae199 Deduplicates some configuration data
Configuration mapping was duplicated between three different tables.
Additionally default configuration values were strewn about. Making it confusing as to what the default value would end up being

Adds a new ConfigValues.inl header that defines a few things right next to each other.
Defines the enum name as usual.
Defines the JSON config option name.
Defines the Environment config option name
Defines the default value that the configuration should be
2021-03-08 22:53:41 -08:00
Ryan Houdek a2278cfb1e Adds StrConv type for enum conversion 2021-03-08 22:50:42 -08:00
Scott Mansell 98cb54e3e4 Merge pull request #812 from phire/both_sides
Allow both ARM64 and X86_64 jits to be compiled at the same time
2021-03-09 14:59:11 +13:00
Scott Mansell cb17127ccb Remove last traces of vixl simulator mode 2021-03-08 16:30:25 +13:00
Scott Mansell 93b6dbdc2b Fix Arm64JitCore instantiation on x86 2021-03-08 16:30:25 +13:00
Scott Mansell bea0adc3d1 Allow both jitcores to be enabled simultaneously 2021-03-08 16:30:25 +13:00
Scott Mansell 9c339c85c6 Cmake: allow independant control of both jits 2021-03-08 16:30:25 +13:00
Scott Mansell 9cf7545f8d Abstract mcontext out of JITs
This will help with unification of common code later
2021-03-08 16:30:25 +13:00
Ryan Houdek c657165e56 Implements CPUID 0x8000'0005 for L1 cacheline information
@skmp commented about this in the L2 cacheline information PR.
Quickly implement it since it also has cacheline information.

Fairly trivial since every piece of information is just some 8bit variables.
2021-03-06 08:32:47 -08:00
Ryan Houdek 05ae7d3c5e Merge pull request #814 from Sonicadvance1/cacheline_cpuid
Implements CPUID 0x8000'0006 for cacheline information
2021-03-06 08:23:03 -08:00
Ryan Houdek f94529ebb2 Implements CPUID 0x8000'0006 for cacheline information
This exposes some cache information including cacheline information.

Fills out this data structure as well just incase we hit another cacheline size bug
2021-03-06 07:30:25 -08:00
Ryan Houdek 852ab4e245 Merge pull request #815 from Sonicadvance1/sse4_1_movnt
Implements MOVNTDQA
2021-03-06 07:25:56 -08:00
Ryan Houdek bb6e7360a2 Merge pull request #813 from Sonicadvance1/disable_python_dev_check
Disabled cmake check for python development
2021-03-06 07:25:16 -08:00
Ryan Houdek 50ae18cbaf Adds MOVNTDQA unit test 2021-03-06 06:23:27 -08:00
Ryan Houdek 1adbf66fee Implements MOVNTDQA
Found this when running Steamlink under FEX on my laptop.
Turns out the Iris video driver now uses this unconditionally in a couple of locations.
Makes sense there since all the Iris targets are expected to support SSE4.1 atm.
2021-03-06 06:23:18 -08:00
Ryan Houdek 8222acdf20 Disabled cmake check for python development
We only need the interpreter
2021-03-06 06:06:59 -08:00
Stefanos Kornilios Mitsis Poiitidis 3978097f15 Merge pull request #705 from FEX-Emu/skmp/mmap-cache-flushing
Flush IR/Code cache on mmap, mmunmap & mprot
2021-03-03 12:59:29 +02:00
Stefanos Kornilios Mitsis Poiitidis 613e5b66ba SMC: Expand --smc-checks options to 'none', 'mman', 'full' 2021-03-03 12:49:23 +02:00
Ryan Houdek e95d1bd761 Merge pull request #811 from phire/remove-operator-names
X64_64 JIt: Remove -fno-operator-names
2021-03-02 18:55:53 -08:00
Scott Mansell ec95610587 X64_64 JIt: Remove -fno-operator-names
Some IDEs didn't recognize this flag and threw
phantom errors when these files were open.

Lets lean towards a smoother developer experience
2021-03-03 15:47:58 +13:00
Ryan Houdek b57cf78cc7 Merge pull request #810 from phire/fix_include
Add missing unordered_map include
2021-03-02 18:00:39 -08:00
Scott Mansell f5ac5d1667 Add missing unordered_map include
It's really unhelpful that libstdc++ includes this by default
2021-03-03 14:06:16 +13:00
Stefanos Kornilios Mitsis Poiitidis e92578817d SMC: Flush IR/Code cache on mmap, mmunmap & mprot 2021-02-26 13:22:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 5b4696976d Merge pull request #805 from FEX-Emu/skmp/asm-thunks
Thunks: Convert guest thunk to asm, to avoid issues with older gcc versions
2021-02-26 12:57:14 +02:00
Stefanos Kornilios Mitsis Poiitidis aecea294d5 Merge pull request #792 from Sonicadvance1/implemented_unaligned_memory_ops
Implements unaligned atomic memory ops for ARMv8.1+
2021-02-26 12:50:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 70a98cd465 Thunks: Convert guest thunk to asm, to avoid issues with older gcc verions 2021-02-26 12:43:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 2071bab238 Merge pull request #803 from Sonicadvance1/faccessat2
Implements support for faccessat2 in syscallhandler
2021-02-25 23:58:11 +02:00
Ryan Houdek 9746e0175e Change unsupported faccessat2 to use safe syscall unsupported 2021-02-25 13:45:24 -08:00
Ryan Houdek d78a3784fb Guards SYS_faccessat2 define for new enough glibc defines 2021-02-25 13:44:40 -08:00
Ryan Houdek 7dd5252d68 Implements support for faccessat2 in syscallhandler
Only exists if the host kernel is >= 5.8.0.
2021-02-24 09:52:30 -08:00
Ryan Houdek ac92a103df Adds support for faccessat2 to FileManager 2021-02-24 09:51:29 -08:00
Ryan Houdek 74a00fee0f Updates syscall definitions enums 2021-02-24 09:51:03 -08:00
Ryan Houdek ebe447f86a Merge pull request #801 from Sonicadvance1/fix_cmpxchg_flags
Fixes CMPXCHG flags being incorrect aside from ZF
2021-02-24 09:01:01 -08:00
Ryan Houdek ebac071ae1 Adds cmpxchg to register zext and flags unit test 2021-02-23 23:22:24 -08:00
Ryan Houdek 68db163efc Adds cmpxchg to memory zext and flag unit test 2021-02-23 23:22:00 -08:00
Ryan Houdek 6c469c21db Fixes Zext and flags behaviour of CMPXCHG to register 2021-02-23 23:21:28 -08:00
Ryan Houdek 9798522bc5 Fixes Zext behaviour of CMPXCHG to memory 2021-02-23 23:21:05 -08:00
Ryan Houdek 55ddc5b6e4 Adds unit tests to ensure cmpxchg flag correctness 2021-02-23 22:02:39 -08:00
Ryan Houdek 373f4275f6 Fixes CMPXCHG flags being incorrect aside from ZF
Almost everything only checks ZF but we had the arguments reversed for
the rest of the comparison flag results.
2021-02-23 22:01:52 -08:00
Scott Mansell 6fee5b9ca2 Merge pull request #789 from FEX-Emu/skmp/add-cacheinfo-cpuid
CPUID: Add cache information, function 0x2
2021-02-24 03:06:20 +13:00
Stefanos Kornilios Mitsis Poiitidis aaa63a2409 Merge pull request #730 from FEX-Emu/skmp/ir-cache
AOTIR: Initial implementation
2021-02-23 13:10:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 6d0d3eaf04 AOTIR: Review feedback 2021-02-23 12:23:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 075cd423ed AOTIR: Review feedback 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 8e06966ddc AOTIR: Make OP_REMOVECODEENTRY relocatable 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis bdcf3ff606 AOTIR: Merge IRLists, RALists and DebugDataLists to LocalIRCache; cleanups and fixups 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b29121c71a AOTIR: Rename AOTCache to AOTIRCache 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis a5ee4bdf3c AOTIR: Fix double free issues 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis abb2e09c20 AOTIR: Add support for 32-bit process 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 56d10dcbbb IR: Make OP_THUNK use an inline sha256 hash of the thunk name, update thunk scripts 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 1f919949cc JIT: Fix arm64 build 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 63507f5ece IR: Fix int64_t parsing 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis dbc8b8e8e9 AOTIR: Append optimization flags to fileid 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 681deca4eb IR: Make entrypoint implicit, Add InlineEntrypointOffset, make ValidateCode Entrypoint-relative 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b3d12fbb7b AOTIR: Cleanup interface 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fafc987eab AOTIR: Rename Generate to Capture 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis d825ed2b03 AOTIR: Reduce map lookup 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis f2f62c5f2c AOTIR: Implement EntrypointOffset for aarch64 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 3e1e093b24 AOTIR: Store in ~/.fex-emu/aotir, per so 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 71dceb5339 Fix relocation support for FinishOp 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis be40d1604a Improve aotir format 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 47d4481a75 AOTIR: Support relocations via new ir op 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fc30efe040 AOTIR: Add hashing of code bytes 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 57680a8614 AOTIR: Add --aotir-generate and --aotir-load to FEXLoader & Config 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 66bdface49 AOTIR: Make copies for insertions, only cache RA'd blocks 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 816c4656df IR: Add AOT Cache 2021-02-23 12:08:57 +02:00
Scott Mansell 7b97c3fa23 Merge pull request #755 from Sonicadvance1/host_uname
Pulls uname nodename from host system
2021-02-23 23:07:29 +13:00
Stefanos Kornilios Mitsis Poiitidis 8642406b8f CPUID: Add cache information, function 0x2 2021-02-23 11:01:15 +02:00
Stefanos Kornilios Mitsis Poiitidis e2599db3ed Merge pull request #793 from Sonicadvance1/assert_on_missing_lock
Adds assert checks on missing LOCK support
2021-02-23 10:51:30 +02:00
Ryan Houdek bb0850eab8 Adds assert checks on missing LOCK support
These usually don't get hit, but Geekbench4 DOES manage to hit LOCK on
BTS.

Which we just don't support right now.
2021-02-22 23:20:18 -08:00
Stefanos Kornilios Mitsis Poiitidis 1af541475a Merge pull request #795 from Sonicadvance1/atomic_bittest_ops
Implements BTC, BTR, BTS atomic variants
2021-02-23 09:08:39 +02:00
Stefanos Kornilios Mitsis Poiitidis 6669bd2d4a Merge pull request #794 from Sonicadvance1/fix_log_move_fail
Pass log moves on buildbot stage failure
2021-02-23 08:59:21 +02:00
Ryan Houdek d3e98fedb4 Pulls uname nodename from host system
Hardcoding FEXCore as the nodename is an annoyance.
Pull the host's nodename instead.

Fixes #600
2021-02-22 22:39:05 -08:00
Ryan Houdek 8a39f4b25c Implements BTC, BTR, BTS atomic unit tests
This just takes the regular non-atomic unit tests and changes them to
have lock prefixes.
These are all handled as byte sized atomics so there aren't any
alignment problems.
2021-02-22 22:29:31 -08:00
Ryan Houdek 232eeff483 Implements BTC, BTR, BTS atomic variants
BTS specifically is being used for threading related tasks. Which could
be why our threading has been unstable.
2021-02-22 22:27:10 -08:00
Ryan Houdek 12ef69650f Pass log moves on buildbot stage failure
Passes the log movement stage to clean up the output.
Really it is only the first unit test stage that is failing, but each
subsequent log run will have failed to move since the previous stage
didn't run.

Makes it easier to scan over a failure as only the first unit test
failure step.
2021-02-22 18:42:28 -08:00
Ryan Houdek f1349becb4 Disables the unaligned atomic memory op tests
These don't run on the armv8.0 runner
2021-02-22 18:33:17 -08:00
Ryan Houdek f848957e17 Implements unit tests for the new unaligned atomics
Tests all the unaligned atomic ops we support now
2021-02-22 18:25:25 -08:00
Ryan Houdek 55fc47e0b7 Passes SIGBUS handler to unaligned atomic op handler for ARMv8.1
ARMv8.0 still unsupported here.

This allows us to support unaligned atomic ops for every atomic x86 op
that we support.
2021-02-22 18:22:10 -08:00
Ryan Houdek 4f40166902 Implements unaligned atomic memory ops for ARMv8.1+
This takes the existing unaligned CAS handling code and makes it more
robust to handle both true unaligned CAS and unaligned atomic memory
ops.

Specifically it inserts some lambda helpers to calculate the Desired and
Expected values inside of the CAS loops to account for CAS and memory
operation.
2021-02-22 18:17:53 -08:00
Stefanos Kornilios Mitsis Poiitidis fba626c547 Merge pull request #790 from FEX-Emu/skmp/workaround-exit-group
Threading: Workaround exit_group bug
2021-02-22 17:50:59 +02:00
Stefanos Kornilios Mitsis Poiitidis 9b360e66dc Threading: Workaround exit_group bug 2021-02-22 14:24:42 +02:00
Stefanos Kornilios Mitsis Poiitidis 3a6fd00154 Merge pull request #785 from FEX-Emu/skmp/fix-sar8-16
OpDisp: Imm SAR OpSize < 32 needs sign extension
2021-02-18 22:33:46 +02:00
Stefanos Kornilios Mitsis Poiitidis 3bc12de017 OpDisp: Imm SAR OpSize < 32 needs sign extension 2021-02-17 08:17:20 +02:00
Stefanos Kornilios Mitsis Poiitidis 5ae6a64800 Merge pull request #778 from FEX-Emu/skmp/x87-round-trunc
x87: Round, Truncate & Precision control
2021-02-17 05:27:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 6615d5972f x87: Support for FCW, rounding, truncation, precision control 2021-02-17 05:16:53 +02:00
Stefanos Kornilios Mitsis Poiitidis d075b689a5 Merge pull request #784 from FEX-Emu/skmp/fix-cvt-cvtt
OpDisp: Fix cvt/cvtt mapping to dispatcher handlers
2021-02-17 05:14:44 +02:00
Stefanos Kornilios Mitsis Poiitidis eb60bda1d3 OpDisp: Fix cvt/cvtt mapping to dispatcher handlers 2021-02-16 11:19:28 +02:00
Stefanos Kornilios Mitsis Poiitidis 8c5d88816f Merge pull request #781 from FEX-Emu/skmp/hackfix-frem
X87: FPREM needs to set C2 to 0 to indicate finished iteration
2021-02-16 09:54:37 +02:00
Ryan Houdek 73e9a35e36 Merge pull request #783 from FEX-Emu/skmp/fix-selects
Syscalls/x86-32: Fix selects to remarshal the fd_sets
2021-02-15 23:26:33 -08:00
Stefanos Kornilios Mitsis Poiitidis 568fb3643d Syscalls/x86-32: Fix selects to actually remarshal the sets after running the host syscall 2021-02-16 02:35:38 +02:00
Stefanos Kornilios Mitsis Poiitidis 82367f49e2 X87: FPERM needs to set C2 to 0 to indicate finished iteration 2021-02-16 02:34:52 +02:00
Stefanos Kornilios Mitsis Poiitidis 9f2d49026e Merge pull request #768 from FEX-Emu/skmp/fix-csgo
Fixes for Counter Strike Global Offensive
2021-02-15 03:47:14 +02:00
Stefanos Kornilios Mitsis Poiitidis bed00d11f1 Tests: Disable pr57275.c because it uses VMOVAPS 2021-02-14 21:36:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 9f0a3dab87 Merge pull request #769 from FEX-Emu/skmp/jit-fallbacks
x87: Call C++ handlers instead of forcing interpreter
2021-02-14 21:27:32 +02:00
Stefanos Kornilios Mitsis Poiitidis 07bd2244c7 IR: Remove ShouldInterpret & related logic 2021-02-12 14:21:47 +02:00
Stefanos Kornilios Mitsis Poiitidis 077093eacc JIT/x64: Use xmm0 as temp instead of xmm11, add xmm11 to RA list 2021-02-12 14:15:22 +02:00
Stefanos Kornilios Mitsis Poiitidis 428111b44b JIT/arm64: fallback support 2021-02-12 01:23:15 +02:00
Stefanos Kornilios Mitsis Poiitidis b564927825 IprOps: Add BCDLOAD, BCDSTORE fallbacks 2021-02-11 19:24:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b0a046dfcb JIT/IPR: Add fallbacks for almost all x87 ops 2021-02-11 18:52:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 8f5bbde4dd JIT/x86: Implement VBSL 2021-02-11 18:51:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 7dc0c275cf Add some basic infrastructure + a few fpr ops 2021-02-11 17:20:29 +02:00
Stefanos Kornilios Mitsis Poiitidis 4e31c1d1ce RA: Remove erratic assert 2021-02-11 14:31:53 +02:00
Stefanos Kornilios Mitsis Poiitidis df1c38908c JIT: Support edge case in CreateElementPair 2021-02-11 14:31:32 +02:00
Stefanos Kornilios Mitsis Poiitidis fc4b875ff1 Frontend: Only return error if invalid op in first block 2021-02-11 14:31:02 +02:00
Ryan Houdek 40b891ca69 Merge pull request #766 from phire/brk_fix
BRK: don't munmap if mmap failed
2021-02-08 21:32:02 -08:00
Scott Mansell 3d6461f0ca BRK: don't munmap if mmap failed 2021-02-09 17:29:44 +13:00
Ryan Houdek c11b629189 Merge pull request #764 from phire/fix_cpuinfo
Fix the number of physical cpus reported
2021-02-08 16:13:09 -08:00
Scott Mansell 46ec7c406f emulated-cpuid: add some comments 2021-02-09 13:03:49 +13:00
Ryan Houdek 2a0b9ae681 Merge pull request #763 from phire/bigger_spills
arm64: Fix blocks with over 256 spill slots.
2021-02-08 15:57:15 -08:00
Scott Mansell a28acf5505 Fix the number of physical cpus reported. 2021-02-09 12:52:27 +13:00
Scott Mansell c4dda4a3f3 arm64: Fix blocks with over 256 spill slots.
With multiblock and -n4000, at least one block in geekbench5's camera
benchmark was spilling 324 values. This meant the required stack
adjustment was 5184 bytes, which is larger than can fit into a 12bit
immediate.

Now, it's probally a bug that our RA is spilling that many values, and
we probally should be packing our spill slots closer together, but we
still shouldn't fail to run in this edge case.

So this commit adds a fallback path which uses a temp register to load
the stack adjustment.
2021-02-09 12:08:32 +13:00
Stefanos Kornilios Mitsis Poiitidis b40131ede3 Merge pull request #759 from phire/fix_mul_op
Fix power of 2 OP_MUL optimisation (Fixes Geekbench result upload)
2021-02-07 08:08:37 +02:00
Scott Mansell 5b0ade0bd8 Fix power of 2 OP_MUL optimisation
The new left-shift that replaced the multiply takes
it's size from arg[0]. If arg[0] was 64bit and the original
OP_MUL was 32bit, then the original code would truncate
upper bits, and the replacement code wouldn't.

This bug broke the SSL cert checking code in geekbench,
causing it to fail to upload results.
Fixes #647
2021-02-07 18:49:47 +13:00
Stefanos Kornilios Mitsis Poiitidis fcf1c2572b Merge pull request #756 from Sonicadvance1/python_check
Adds python version check
2021-02-06 14:20:49 +02:00
Ryan Houdek 69891867b1 Adds python version check 2021-02-05 10:36:52 -08:00
Ryan Houdek abefae2b58 Merge pull request #754 from Sonicadvance1/fix_artifacts_on_failure
Always upload artifact results even on failure
2021-02-03 16:08:51 -08:00
Ryan Houdek 143cb6d0a5 Always upload artifact results even on failure
This was an oversight of the original commit
2021-02-03 15:59:55 -08:00
Ryan Houdek e0042cc3fc Merge pull request #752 from Sonicadvance1/upload_artifacts
Uploads testing artifacts to github
2021-02-03 09:40:57 -08:00
Ryan Houdek 6ccdfd11a3 Uploads testing artifacts to github 2021-02-02 06:56:06 -08:00
463 changed files with 30192 additions and 10790 deletions

No files matched your search

+79 -2
View File
@@ -13,18 +13,22 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_FORCE32BITALLOCATOR: 1
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64], [self-hosted, ARMv8.0], [self-hosted, ARMv8.2]]
arch: [[self-hosted, x64], [self-hosted, ARMv8.0], [self-hosted, ARMv8.2], [self-hosted, ARMv8.4]]
fail-fast: false
steps:
- uses: actions/checkout@v2
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
# Need to update submodules
run: git submodule update --init --depth 1
@@ -45,7 +49,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -59,27 +63,100 @@ jobs:
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_64
- name: GCC64 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
- name: gcc target tests 32
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_32
- name: GCC32 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+9
View File
@@ -30,3 +30,12 @@
shallow = true
path = External/fex-gcc-target-tests-bins
url = https://github.com/FEX-Emu/fex-gcc-target-tests-bins.git
[submodule "External/jemalloc"]
path = External/jemalloc
url = https://github.com/FEX-Emu/jemalloc.git
[submodule "External/fmt"]
path = External/fmt
url = https://github.com/fmtlib/fmt.git
[submodule "External/drm-headers"]
path = External/drm-headers
url = https://github.com/FEX-Emu/drm-headers.git
+143 -32
View File
@@ -12,6 +12,8 @@ option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -27,14 +29,6 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_IWYU)
find_program(IWYU_EXE "iwyu")
if (IWYU_EXE)
message(STATUS "IWYU enabled")
set(CMAKE_CXX_INCLUDE_WHAT_YOU_USE "${IWYU_EXE}")
endif()
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -67,8 +61,8 @@ endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address)
link_libraries(-fno-omit-frame-pointer -fsanitize=address)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
link_libraries(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
endif()
if (ENABLE_TSAN)
@@ -83,6 +77,14 @@ set (CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" Be warned: FEX isn't optimized for x86_64 hosts!\n"
" Support for x86_64 hosts is only for debugging and convenience!\n"
" Don't expect amazing performance or optimal code generation!\n"
" Pass -DENABLE_X86_HOST_DEBUG=True to bypass this message!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
@@ -91,24 +93,31 @@ endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
set(_M_ARM_64 1)
add_definitions(-D_M_ARM_64=1)
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
if(CMAKE_BUILD_TYPE MATCHES DEBUG)
add_definitions(-DVIXL_DEBUG=1)
endif()
endif()
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
message(FATAL_ERROR "FEX doesn't support getting compiled with GCC!")
endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
pkg_check_modules(XXHASH libxxhash REQUIRED)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
add_subdirectory(External/cpp-optparse/)
include_directories(External/cpp-optparse/)
add_subdirectory(External/fmt/)
add_subdirectory(External/imgui/)
include_directories(External/imgui/)
@@ -147,32 +156,115 @@ if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
endif()
if(_M_ARM_64)
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
# Disable some Werror that can add frustration when developing
add_compile_options(-Wno-error=unused-variable)
endif()
endif()
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=native")
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
endif()
endif()
endif()
if (ENABLE_IWYU)
find_program(IWYU_EXE "iwyu")
if (IWYU_EXE)
message(STATUS "IWYU enabled")
set(CMAKE_CXX_INCLUDE_WHAT_YOU_USE "${IWYU_EXE}")
endif()
endif()
if (ENABLE_CLANG_FORMAT)
find_program(CLANG_TIDY_EXE "clang-tidy")
if (NOT CLANG_TIDY_EXE)
message(FATAL_ERROR "Couldn't find clang-tidy")
endif()
set(CLANG_TIDY_FLAGS
"-checks=*"
"-fuchsia*"
"-bugprone-macro-parentheses"
"-clang-analyzer-core.*"
"-cppcoreguidelines-pro-type-*"
"-cppcoreguidelines-pro-bounds-array-to-pointer-decay"
"-cppcoreguidelines-pro-bounds-pointer-arithmetic"
"-cppcoreguidelines-avoid-c-arrays"
"-cppcoreguidelines-avoid-magic-numbers"
"-cppcoreguidelines-pro-bounds-constant-array-index"
"-cppcoreguidelines-no-malloc"
"-cppcoreguidelines-special-member-functions"
"-cppcoreguidelines-owning-memory"
"-cppcoreguidelines-macro-usage"
"-cppcoreguidelines-avoid-goto"
"-google-readability-function-size"
"-google-readability-namespace-comments"
"-google-readability-braces-around-statements"
"-google-build-using-namespace"
"-hicpp-*"
"-llvm-namespace-comment"
"-llvm-include-order" # Messes up with case sensitivity
"-llvmlibc-*"
"-misc-unused-parameters"
"-modernize-loop-convert"
"-modernize-use-auto"
"-modernize-avoid-c-arrays"
"-modernize-use-nodiscard"
"readability-*"
"-readability-function-size"
"-readability-implicit-bool-conversion"
"-readability-braces-around-statements"
"-readability-else-after-return"
"-readability-magic-numbers"
"-readability-named-parameter"
"-readability-uppercase-literal-suffix"
"-cert-err34-c"
"-cert-err58-cpp"
"-bugprone-exception-escape"
)
string(REPLACE ";" "," CLANG_TIDY_FLAGS "${CLANG_TIDY_FLAGS}")
set(CMAKE_CXX_CLANG_TIDY ${CLANG_TIDY_EXE} "${CLANG_TIDY_FLAGS}")
endif()
add_compile_options(-Wall)
add_subdirectory(External/FEXCore)
add_subdirectory(Source/)
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/Config.h)
if (BUILD_TESTS)
include(CTest)
enable_testing()
message(STATUS "Unit tests are enabled")
endif()
add_subdirectory(External/FEXCore)
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
@@ -183,16 +275,35 @@ if (BUILD_THUNKS)
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
CMAKE_ARGS "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
install(
CODE "MESSAGE(\"-- Installing: host-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkHostsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Host
)"
DEPENDS host-libs
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}" "-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
install(
CODE "MESSAGE(\"-- Installing: guest-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkGuestsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
DEPENDS guest-libs
)
endif()
+25
View File
@@ -0,0 +1,25 @@
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS *.json)
file(GLOB GEN_CONFIG_SOURCES CONFIGURE_DEPENDS *.json.in)
# Any application configuration json file gets installed
foreach(CONFIG_SRC ${CONFIG_SOURCES})
install(FILES ${CONFIG_SRC}
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
endforeach()
# Any configuration file json file that needs to be generated
# First generate then install it
foreach(GEN_CONFIG_SRC ${GEN_CONFIG_SOURCES})
# Get the filename only component
get_filename_component(CONFIG_NAME ${GEN_CONFIG_SRC} NAME_WE)
# Configure it
configure_file(
${GEN_CONFIG_SRC}
${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json)
# Then install the configured json
install(
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
endforeach()
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"Env": "STEAM_GAME_LAUNCH_SHELL=@CMAKE_INSTALL_PREFIX@/bin/FEXBash"
}
}
+2 -1
View File
@@ -4,7 +4,8 @@ FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-10 llvm-10 nasm ninja-build libnuma-dev \
libcap-dev libglfw3-dev libepoxy-dev
libcap-dev libglfw3-dev libepoxy-dev python3-dev \
python3 linux-headers-generic
COPY . /opt/FEX
+12 -13
View File
@@ -4,6 +4,17 @@ project(${PROJECT_NAME}
VERSION 0.01
LANGUAGES CXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
set(_M_ARM_64 1)
endif()
set(ENABLE_JIT_X86_64 ${_M_X86_64} CACHE BOOL "Enable the x86_64 JIT")
set(ENABLE_JIT_ARM64 ${_M_ARM_64} CACHE BOOL "Enable the ARM64 JIT")
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_JITSYMBOLS "Enable visibility of JITSymbols in profiling tools" FALSE)
@@ -17,23 +28,11 @@ set(CMAKE_INCLUDE_CURRENT_DIR ON)
include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -fno-operator-names -mcx16")
set(CMAKE_REQUIRED_DEFINITIONS "-fno-operator-names")
message(STATUS "Enabling x86-64 JIT")
set(ENABLE_JIT 1)
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
message(STATUS "Enabling AArch64 JIT")
set(_M_ARM_64 1)
set(ENABLE_JIT 1)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
endif()
endif()
set(CMAKE_CXX_STANDARD 20)
+467
View File
@@ -0,0 +1,467 @@
import datetime
import json
import sys
def print_header():
header = '''#ifndef OPT_BASE
#define OPT_BASE(type, group, enum, json, default)
#endif
#ifndef OPT_BOOL
#define OPT_BOOL(group, enum, json, default) OPT_BASE(bool, group, enum, json, default)
#endif
#ifndef OPT_UINT8
#define OPT_UINT8(group, enum, json, default) OPT_BASE(uint8_t, group, enum, json, default)
#endif
#ifndef OPT_INT32
#define OPT_INT32(group, enum, json, default) OPT_BASE(int32_t, group, enum, json, default)
#endif
#ifndef OPT_UINT32
#define OPT_UINT32(group, enum, json, default) OPT_BASE(uint32_t, group, enum, json, default)
#endif
#ifndef OPT_UINT64
#define OPT_UINT64(group, enum, json, default) OPT_BASE(uint64_t, group, enum, json, default)
#endif
#ifndef OPT_STR
#define OPT_STR(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#endif
#ifndef OPT_STRARRAY
#define OPT_STRARRAY(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#endif
'''
output_file.write(header)
def print_tail():
tail = '''#undef OPT_BASE
#undef OPT_BOOL
#undef OPT_UINT8
#undef OPT_INT32
#undef OPT_UINT32
#undef OPT_UINT64
#undef OPT_STR
#undef OPT_STRARRAY
'''
output_file.write(tail)
def print_config(type, group_name, json_name, default_value):
output_file.write("OPT_{0} ({1}, {2}, {3}, {4})\n".format(type.upper(), group_name.upper(), json_name.upper(), json_name, default_value))
def print_options(options):
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
print_config(
op_vals["Type"],
op_group,
op_key,
default)
output_file.write("\n")
def print_unnamed_options(options):
output_file.write("// Unnamed configuration options\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
print_config(
op_vals["Type"],
op_group,
op_key.upper(), # KEY is the enum here, there is no json configuration for these
default)
output_file.write("\n")
def print_man_option(short, long, desc, default):
if (short != None):
output_man.write(".It Fl {0} , ".format(short))
else:
output_man.write(".It ")
output_man.write("Fl Fl {0}=".format(long))
output_man.write("\n");
# Print description
for line in desc:
output_man.write(".Pp\n")
output_man.write("{0}\n".format(line))
output_man.write(".Pp\n")
output_man.write("\\fBdefault:\\fR {0}\n".format(default))
output_man.write(".Pp\n\n")
def print_man_env_option(name, desc, default):
output_man.write("\\fBFEX_{0}\\fR\n".format(name))
# Print description
for line in desc:
output_man.write(".Pp\n")
output_man.write("{0}\n".format(line))
output_man.write(".Pp\n")
output_man.write("\\fBdefault:\\fR {0}\n".format(default))
output_man.write(".Pp\n\n")
def print_man_options(options):
output_man.write(".Sh OPTIONS\n")
output_man.write(".Bl -tag -width -indent\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
default = op_vals["Default"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_option(
short,
long,
op_vals["Desc"],
default
)
output_man.write(".El\n")
def print_man_environment(options):
output_man.write(".Sh ENVIRONMENT\n")
output_man.write(".Bl -tag -width -indent\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_env_option(
op_key.upper(),
op_vals["Desc"],
default
)
print_man_environment_tail()
output_man.write(".El\n")
def print_man_environment_tail():
# Additional environment variables that live outside of the normal loop
print_man_env_option(
"FEX_APP_CONFIG_LOCATION",
[
"Allows the user to override where FEX looks for configuration files",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/",
"This will override the full path",
],
"''")
print_man_env_option(
"FEX_APP_CONFIG",
[
"Allows the user to override where FEX looks for only the application config file",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/Config.json",
"This will override this file location",
"One must be careful with this option as it will override any applications that load with execve as well"
"If you need to support applications that execve then use FEX_APP_CONFIG_LOCATION instead"
],
"''")
print_man_env_option(
"FEX_APP_DATA_LOCATION",
[
"Allows the user to override where FEX looks for data files",
"By default FEX will look in {$HOME, $XDG_DATA_HOME}/.fex-emu/",
"This will override the full path",
"This is the folder where FEX stores generated files like IR cache"
],
"''")
def print_man_header():
header ='''.Dd {0}
.Dt FEX
.Os Linux
.Sh NAME
.Nm FEXLoader
.Nm FEXInterpreter
.Nm FEXBash
.Nd Fast x86-64 and x86 emulation.
.Sh SYNOPSIS
.Nm
.Op options
.Op Ar --
.Ar Application
<args> ...
.Pp
.Nm FEXInterpreter
.Ar Application
<args> ...
.Pp
.Nm FEXBash
.Ar <args> ...
.Sh DESCRIPTION
FEX allows you to run x86 and x86-64 binaries on an AArch64 host, similar to qemu-user and box86.
It has native support for a rootfs overlay, so you don't need to chroot, as well as some thunklibs so it can forward things like GL to the host.
FEX presents a Linux 5.0 interface to the guest, and supports both AArch64 and x86-64 as hosts.
FEX is very much work in progress, so expect things to change.
'''
output_man.write(header.format(datetime.datetime.now().strftime("%d-%m-%Y")))
def print_man_tail():
tail ='''.Sh FILES
.Bl -tag -width "$prefix/share/fex-emu/GuestThunks" -compact
.It Pa $XDG_HOME_DIR/.fex-emu
Default FEX user configuration directory
.It Pa $prefix/share/fex-emu/AppConfig
System level application configuration files
.It Pa $prefix/share/fex-emu/GuestThunks
guest-side thunk data libraries
.It Pa $prefix/lib/fex-emu/HostThunks
host-side thunks for guest communication
.El
'''
output_man.write(tail)
def print_config_option(type, group_name, json_name, default_value, short, choices, desc):
if (type == "bool"):
# Bool gets some special handling to add an inverted case
output_argloader.write("{0}Group".format(group_name))
options = ""
AddedArg = False
if (short != None):
AddedArg = True
options += "\"-{0}\"".format(short)
if (AddedArg):
options += ", "
options += "\"--{0}\"".format(json_name.lower())
output_argloader.write(".add_option({0})".format(options))
output_argloader.write("\n")
output_argloader.write("\t.action(\"store_true\")\n")
output_argloader.write("\t.dest(\"{0}\")\n".format(json_name));
# help
output_argloader.write("\t.help(\n")
desc_line_ender = ""
if (len(desc) > 1):
desc_line_ender = "\\n"
for line in desc:
output_argloader.write("\t\t\"{0}{1}\"\n".format(line, desc_line_ender))
output_argloader.write("\t)\n")
output_argloader.write("\t.set_default({0});\n\n".format(default_value));
output_argloader.write("{0}Group".format(group_name))
output_argloader.write(".add_option(\"--no-{0}\")\n".format(json_name.lower()))
# Inverted case sets the bool to false
output_argloader.write("\t.action(\"store_false\")\n")
output_argloader.write("\t.dest(\"{0}\");\n".format(json_name));
else:
output_argloader.write("{0}Group".format(group_name))
options = ""
AddedArg = False
if (short != None):
AddedArg = True
options += "\"-{0}\"".format(short)
if (AddedArg):
options += ", "
options += "\"--{0}\"".format(json_name.lower())
output_argloader.write(".add_option({0})".format(options))
output_argloader.write("\n")
output_argloader.write("\t.dest(\"{0}\")\n".format(json_name));
if (choices != None):
output_argloader.write("\t.choices({\n")
for choice in choices:
output_argloader.write("\t\t\"{0}\",\n".format(choice))
output_argloader.write("\t})\n")
# help
output_argloader.write("\t.help(\n")
desc_line_ender = ""
if (len(desc) > 1):
desc_line_ender = "\\n"
for line in desc:
output_argloader.write("\t\t\"{0}{1}\"\n".format(line, desc_line_ender))
output_argloader.write("\t)\n")
output_argloader.write("\t.set_default({0});\n".format(default_value));
output_argloader.write("\n");
def print_argloader_options(options):
output_argloader.write("#ifdef BEFORE_PARSE\n")
output_argloader.write("#undef BEFORE_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = "\"" + op_vals["TextDefault"] + "\""
short = None
choices = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if ("Choices" in op_vals):
choices = op_vals["Choices"]
print_config_option(
op_vals["Type"],
op_group,
op_key,
default,
short,
choices,
op_vals["Desc"])
output_argloader.write("\n")
output_argloader.write("#endif\n")
def print_parse_argloader_options(options):
output_argloader.write("#ifdef AFTER_PARSE\n")
output_argloader.write("#undef AFTER_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
output_argloader.write("if (Options.is_set_by_user(\"{0}\")) {{\n".format(op_key))
value_type = op_vals["Type"]
NeedsString = False
conversion_func = "std::to_string"
if ("ArgumentHandler" in op_vals):
NeedsString = True
conversion_func = "FEX::Handler::{0}".format(op_vals["ArgumentHandler"])
if (value_type == "str"):
NeedsString = True
conversion_func = ""
if (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
output_argloader.write("\tfor (auto iter = Array.begin(); iter != Array.end(); ++iter) {\n")
output_argloader.write("\t\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, *iter);\n".format(op_key.upper()))
output_argloader.write("\t}\n")
else:
if (NeedsString):
output_argloader.write("\tstd::string UserValue = Options[\"{0}\"];\n".format(op_key))
else:
output_argloader.write("\t{0} UserValue = Options.get(\"{1}\");\n".format(value_type, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, {1}(UserValue));\n".format(op_key.upper(), conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def check_for_duplicate_options(options):
short_map = []
long_map = []
# Spin through all the items and see if we have a duplicate option
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
long_invert = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if (op_vals["Type"] == "bool"):
long_invert = "no-" + long
# Check for short key duplication
if (short != None):
if (short in short_map):
raise Exception("Short config '{0}' for option '{1}' has duplicate entry!".format(short, op_key))
else:
short_map.append(short)
# Check for long key duplication
if (long in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long))
else:
long_map.append(long)
# Check for long key duplication
if (long_invert != None):
if (long_invert in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long_invert))
else:
long_map.append(long_invert)
if (len(sys.argv) < 5):
sys.exit()
output_filename = sys.argv[2]
output_man_page = sys.argv[3]
output_argumentloader_filename = sys.argv[4]
json_file = open(sys.argv[1], "r")
json_text = json_file.read()
json_file.close()
json_object = json.loads(json_text)
options = json_object["Options"]
unnamed_options = json_object["UnnamedOptions"]
check_for_duplicate_options(options)
# Generate config include file
output_file = open(output_filename, "w")
print_header()
print_options(options)
print_unnamed_options(unnamed_options)
print_tail()
output_file.close()
# Generate man file
output_man = open(output_man_page, "w")
print_man_header()
print_man_options(options)
print_man_environment(options)
print_man_tail()
output_man.close()
# Generate argument loader code
output_argloader = open(output_argumentloader_filename, "w")
print_argloader_options(options);
print_parse_argloader_options(options);
output_argloader.close()
+25 -19
View File
@@ -108,10 +108,10 @@ def print_ir_sizes(ops, defines):
output_file.write("[[maybe_unused]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
output_file.write("std::string_view const& GetName(IROps Op);\n")
output_file.write("uint8_t GetArgs(IROps Op);\n")
output_file.write("FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("bool HasSideEffects(IROps Op);\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) std::string_view const& GetName(IROps Op);\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) uint8_t GetArgs(IROps Op);\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) bool HasSideEffects(IROps Op);\n")
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
@@ -277,7 +277,7 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write("\tusing IRPair = Wrapper<T>;\n\n")
output_file.write("\tIRPair<IROp_Header> AllocateRawOp(size_t HeaderSize) {\n")
output_file.write("\t\tauto Op = reinterpret_cast<IROp_Header*>(Data.Allocate(HeaderSize));\n")
output_file.write("\t\tauto Op = reinterpret_cast<IROp_Header*>(DualListData.DataAllocate(HeaderSize));\n")
output_file.write("\t\tmemset(Op, 0, HeaderSize);\n")
output_file.write("\t\tOp->Op = IROps::OP_DUMMY;\n")
output_file.write("\t\treturn IRPair<IROp_Header>{Op, CreateNode(Op)};\n")
@@ -286,7 +286,7 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tT *AllocateOrphanOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(Data.Allocate(Size));\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn Op;\n")
@@ -295,25 +295,25 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tIRPair<T> AllocateOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(Data.Allocate(Size));\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpSize(OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(Data.Begin());\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Size;\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElements(OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(Data.Begin());\n")
output_file.write("\t\tLogMan::Throw::A(HeaderOp->HasDest, \"Op %s has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\tLOGMAN_THROW_A(HeaderOp->HasDest, \"Op %s has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\treturn HeaderOp->Size / HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(Data.Begin());\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->HasDest;\n")
output_file.write("\t}\n\n")
@@ -387,11 +387,14 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write(") {\n")
output_file.write("\t\tauto Op = AllocateOp<IROp_%s, IROps::OP_%s>();\n" % (op_key, op_key.upper()))
output_file.write("\t\tOp.first->Header.NumArgs = %d;\n" % (SSAArgs))
if (SSAArgs != 0):
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
for i in range(0, SSAArgs):
output_file.write("\t\tOp.first->Header.Args[%d] = ssa%d->Wrapped(ListDataBegin);\n" % (i, i))
if (SSAArgs != 0):
for i in range(0, SSAArgs):
output_file.write("\t\tOp.first->Header.Args[%d] = ssa%d->Wrapped(ListData.Begin());\n" % (i, i))
output_file.write("\t\tssa%d->AddUse();\n" % (i))
if (HasArgs):
@@ -399,11 +402,6 @@ def print_ir_allocator_helpers(ops, defines):
data_name = op_vals["Args"][i]
output_file.write("\t\tOp.first->%s = %s;\n" % (data_name, data_name))
if (HasFixedDestSize):
output_file.write("\t\tOp.first->Header.Size = %d;\n" % FixedDestSize)
if (HasDestSize):
output_file.write("\t\tOp.first->Header.Size = %s;\n" % DestSize)
if (HasDest):
# We can only infer a size if we have arguments
if not (HasFixedDestSize or HasDestSize):
@@ -412,10 +410,18 @@ def print_ir_allocator_helpers(ops, defines):
if (SSAArgs != 0):
for i in range(0, SSAArgs):
output_file.write("\t\tuint8_t Size%d = GetOpSize(ssa%s);\n" % (i, i))
for i in range(0, SSAArgs):
output_file.write("\t\tInferSize = std::max(InferSize, Size%d);\n" % (i))
output_file.write("\t\tOp.first->Header.Size = InferSize;\n")
output_file.write("\t\tOp.first->Header.NumArgs = %d;\n" % (SSAArgs))
if (HasFixedDestSize):
output_file.write("\t\tOp.first->Header.Size = %d;\n" % FixedDestSize)
if (HasDestSize):
output_file.write("\t\tOp.first->Header.Size = %s;\n" % DestSize)
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / (%s);\n" % NumElements)
if (HasDest):
@@ -499,7 +505,7 @@ def print_ir_parser_allocator_helpers(ops, defines):
if (SSAArgs != 0):
for i in range(0, SSAArgs):
output_file.write("\t\tOp.first->Header.Args[%d] = ssa%d->Wrapped(ListData.Begin());\n" % (i, i))
output_file.write("\t\tOp.first->Header.Args[%d] = ssa%d->Wrapped(DualListData.ListBegin());\n" % (i, i))
output_file.write("\t\tssa%d->AddUse();\n" % (i))
if (HasArgs):
+130 -94
View File
@@ -1,49 +1,4 @@
if (ENABLE_CLANG_FORMAT)
find_program(CLANG_TIDY_EXE "clang-tidy")
set(CLANG_TIDY_FLAGS
"-checks=*"
"-fuchsia*"
"-bugprone-macro-parentheses"
"-clang-analyzer-core.*"
"-cppcoreguidelines-pro-type-*"
"-cppcoreguidelines-pro-bounds-array-to-pointer-decay"
"-cppcoreguidelines-pro-bounds-pointer-arithmetic"
"-cppcoreguidelines-avoid-c-arrays"
"-cppcoreguidelines-avoid-magic-numbers"
"-cppcoreguidelines-pro-bounds-constant-array-index"
"-cppcoreguidelines-no-malloc"
"-cppcoreguidelines-special-member-functions"
"-cppcoreguidelines-owning-memory"
"-cppcoreguidelines-macro-usage"
"-cppcoreguidelines-avoid-goto"
"-google-readability-function-size"
"-google-readability-namespace-comments"
"-google-readability-braces-around-statements"
"-google-build-using-namespace"
"-hicpp-*"
"-llvm-namespace-comment"
"-llvm-include-order" # Messes up with case sensitivity
"-misc-unused-parameters"
"-modernize-loop-convert"
"-modernize-use-auto"
"-modernize-avoid-c-arrays"
"-modernize-use-nodiscard"
"readability-*"
"-readability-function-size"
"-readability-implicit-bool-conversion"
"-readability-braces-around-statements"
"-readability-else-after-return"
"-readability-magic-numbers"
"-readability-named-parameter"
"-readability-uppercase-literal-suffix"
"-cert-err34-c"
"-cert-err58-cpp"
"-bugprone-exception-escape"
)
string(REPLACE ";" "," CLANG_TIDY_FLAGS "${CLANG_TIDY_FLAGS}")
set(CMAKE_CXX_CLANG_TIDY ${CLANG_TIDY_EXE} "${CLANG_TIDY_FLAGS}")
endif()
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (SRCS
Common/Paths.cpp
@@ -130,6 +85,11 @@ set (SRCS
Interface/Core/X86Tables.cpp
Interface/Core/X86DebugInfo.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64_stubs.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/X86Tables/BaseTables.cpp
@@ -154,6 +114,7 @@ set (SRCS
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
@@ -161,61 +122,63 @@ set (SRCS
Interface/IR/Passes/StaticRegisterAllocationPass.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Utils/ELFLoader.cpp
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/ELFContainer.cpp
Utils/ELFSymbolDatabase.cpp
Utils/LogManager.cpp
Utils/Threads.cpp
)
if (_M_X86_64)
list(APPEND SRCS Interface/Core/Interpreter/x86_64Dispatcher.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp
Interface/Core/Interpreter/Arm64Dispatcher.cpp)
Interface/Core/ArchHelpers/Arm64.cpp)
endif()
set (JIT_LIBS )
if (ENABLE_JIT)
if (_M_X86_64)
add_definitions(-D_M_X86_64=1)
if (NOT FORCE_AARCH64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
endif()
endif()
if(_M_ARM_64)
add_definitions(-D_M_ARM_64=1)
add_definitions(-DVIXL_INCLUDE_TARGET_AARCH64=1)
add_definitions(-DVIXL_CODE_BUFFER_MMAP=1)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
list(APPEND JIT_LIBS vixl)
endif()
set(DEFINES )
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
endif()
if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
Interface/Core/JIT/x86_64/ALUOps.cpp
Interface/Core/JIT/x86_64/AtomicOps.cpp
Interface/Core/JIT/x86_64/BranchOps.cpp
Interface/Core/JIT/x86_64/ConversionOps.cpp
Interface/Core/JIT/x86_64/EncryptionOps.cpp
Interface/Core/JIT/x86_64/FlagOps.cpp
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
list(APPEND DEFINES -DJIT_X86_64)
endif()
if (ENABLE_JIT_ARM64)
list(APPEND DEFINES -DJIT_ARM64)
list(APPEND SRCS
Interface/Core/JIT/Arm64/JIT.cpp
Interface/Core/JIT/Arm64/ALUOps.cpp
Interface/Core/JIT/Arm64/AtomicOps.cpp
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
Interface/Core/JIT/Arm64/VectorOps.cpp)
endif()
if (ENABLE_JITSYMBOLS)
add_definitions(-DENABLE_JITSYMBOLS=1)
list(APPEND DEFINES -DENABLE_JITSYMBOLS=1)
endif()
# Generate IR include file
@@ -251,21 +214,64 @@ add_custom_command(
set_source_files_properties(${OUTPUT_IR_NAME} PROPERTIES
GENERATED TRUE)
# Create teh target
# Create the target
add_custom_target(IR_INC
DEPENDS "${OUTPUT_NAME}"
DEPENDS "${OUTPUT_IR_DOC}")
# Generate the configuration include file
set(OUTPUT_CONFIG_FOLDER "${CMAKE_BINARY_DIR}/include/FEXCore/Config")
set(OUTPUT_CONFIG_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigValues.inl")
set(OUTPUT_CONFIG_OPTION_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigOptions.inl")
set(INPUT_CONFIG_NAME "${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
add_custom_command(
OUTPUT "${OUTPUT_CONFIG_NAME}"
OUTPUT "${OUTPUT_CONFIG_OPTION_NAME}"
OUTPUT "${OUTPUT_MAN_NAME}"
DEPENDS "${INPUT_CONFIG_NAME}"
DEPENDS CREATE_CONFIG_FOLDER
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py" "${INPUT_CONFIG_NAME}" "${OUTPUT_CONFIG_NAME}" "${OUTPUT_MAN_NAME}"
"${OUTPUT_CONFIG_OPTION_NAME}"
)
set_source_files_properties(${OUTPUT_CONFIG_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_MAN_NAME} PROPERTIES
GENERATED TRUE)
# Create the target
add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_CONFIG_NAME}"
DEPENDS "${OUTPUT_CONFIG_OPTION_NAME}"
DEPENDS "${OUTPUT_MAN_NAME}")
# Install the man page
install(FILES ${OUTPUT_MAN_NAME} DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
check_cxx_compiler_flag(-fdiagnostics-color=always GCC_COLOR)
check_cxx_compiler_flag(-fcolor-diagnostics CLANG_COLOR)
function(AddLibrary Name Type)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} pthread rt ${JIT_LIBS} ${LINUX_LIBS} dl)
add_dependencies(${Name} CONFIG_INC)
target_link_libraries(${Name} pthread vixl dl fmt::fmt xxhash FEX_jemalloc)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
@@ -275,9 +281,16 @@ function(AddLibrary Name Type)
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
target_compile_definitions(${Name} PRIVATE ${DEFINES})
target_compile_options(${Name}
PRIVATE
-Wno-trigraphs -Wall)
-Wall
-Werror=implicit-fallthrough
-Wno-trigraphs
-ffunction-sections
)
if (GCC_COLOR)
target_compile_options(${Name}
@@ -291,6 +304,29 @@ function(AddLibrary Name Type)
endif()
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} pthread vixl dl fmt::fmt xxhash FEX_jemalloc)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${Name}
PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed"
)
endif()
endfunction()
AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
+7 -6
View File
@@ -1,5 +1,6 @@
#pragma once
#include "Common/MathUtils.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
@@ -16,16 +17,16 @@ struct BitSet final {
ElementType *Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LogMan::Throw::A((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(malloc(AllocateSize));
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LogMan::Throw::A((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(realloc(Memory, AllocateSize));
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
void Free() {
free(Memory);
FEXCore::Allocator::free(Memory);
Memory = nullptr;
}
bool Get(T Element) {
@@ -60,7 +61,7 @@ struct BitSetView final {
ElementType *Memory;
void GetView(BitSet<T> &Set, uint64_t ElementOffset) {
LogMan::Throw::A((ElementOffset % MinimumSize) == 0,
LOGMAN_THROW_A((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
+19 -11
View File
@@ -6,10 +6,13 @@
#include <sys/stat.h>
namespace FEXCore::Paths {
std::string CachePath;
std::string EntryCache;
std::unique_ptr<std::string> CachePath;
std::unique_ptr<std::string> EntryCache;
void InitializePaths() {
CachePath = std::make_unique<std::string>();
EntryCache = std::make_unique<std::string>();
char const *HomeDir = getenv("HOME");
if (!HomeDir) {
@@ -22,29 +25,34 @@ namespace FEXCore::Paths {
char *XDGDataDir = getenv("XDG_DATA_DIR");
if (XDGDataDir) {
CachePath = XDGDataDir;
*CachePath = XDGDataDir;
}
else {
if (HomeDir) {
CachePath = HomeDir;
*CachePath = HomeDir;
}
}
CachePath += "/.fex-emu/";
EntryCache = CachePath + "/EntryCache/";
*CachePath += "/.fex-emu/";
*EntryCache = *CachePath + "/EntryCache/";
// Ensure the folder structure is created for our Data
if (!std::filesystem::exists(EntryCache) &&
!std::filesystem::create_directories(EntryCache)) {
LogMan::Msg::D("Couldn't create EntryCache directory: '%s'", EntryCache.c_str());
if (!std::filesystem::exists(*EntryCache) &&
!std::filesystem::create_directories(*EntryCache)) {
LogMan::Msg::D("Couldn't create EntryCache directory: '%s'", EntryCache->c_str());
}
}
void ShutdownPaths() {
CachePath.reset();
EntryCache.reset();
}
std::string GetCachePath() {
return CachePath;
return *CachePath;
}
std::string GetEntryCachePath() {
return EntryCache;
return *EntryCache;
}
}
+1
View File
@@ -3,6 +3,7 @@
namespace FEXCore::Paths {
void InitializePaths();
void ShutdownPaths();
std::string GetCachePath();
std::string GetEntryCachePath();
}
@@ -96,6 +96,7 @@ extFloat80_t
switch ( roundingMode ) {
case softfloat_round_near_even:
if ( !(sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF )) ) break;
__attribute__((fallthrough));
case softfloat_round_near_maxMag:
if ( exp == 0x3FFE ) goto mag1;
break;
+12 -5
View File
@@ -77,7 +77,7 @@ struct X80SoftFloat {
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_round_near_even, false);
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
@@ -173,19 +173,26 @@ struct X80SoftFloat {
}
operator int16_t() const {
return extF80_to_i32(*this, softfloat_round_near_even, false);
auto rv = extF80_to_i32(*this, softfloat_roundingMode, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
return INT16_MIN;
} else {
return rv;
}
}
operator int32_t() const {
return extF80_to_i32(*this, softfloat_round_near_even, false);
return extF80_to_i32(*this, softfloat_roundingMode, false);
}
operator int64_t() const {
return extF80_to_i64(*this, softfloat_round_near_even, false);
return extF80_to_i64(*this, softfloat_roundingMode, false);
}
operator uint64_t() const {
return extF80_to_ui64(*this, softfloat_round_near_even, false);
return extF80_to_ui64(*this, softfloat_roundingMode, false);
}
void operator=(const float rhs) {
+7
View File
@@ -38,4 +38,11 @@ namespace FEXCore::StrConv {
*Result = Value;
return true;
}
template <typename T,
typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]] static bool Conv(std::string_view Value, T *Result) {
*Result = static_cast<T>(std::stoull(std::string(Value), nullptr, 0));
return true;
}
}
+251 -93
View File
@@ -4,106 +4,123 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Config/Config.h>
#include <filesystem>
#include <pwd.h>
#include <map>
#include <sys/sysinfo.h>
#include <unistd.h>
namespace FEXCore::Config {
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
switch (Option) {
case FEXCore::Config::CONFIG_MULTIBLOCK:
CTX->Config.Multiblock = Config != 0;
break;
case FEXCore::Config::CONFIG_MAXBLOCKINST:
CTX->Config.MaxInstPerBlock = Config;
break;
case FEXCore::Config::CONFIG_DEFAULTCORE:
CTX->Config.Core = static_cast<FEXCore::Config::ConfigCore>(Config);
break;
case FEXCore::Config::CONFIG_VIRTUALMEMSIZE:
CTX->Config.VirtualMemSize = Config;
break;
case FEXCore::Config::CONFIG_SINGLESTEP:
CTX->Config.RunningMode = Config != 0 ? FEXCore::Context::CoreRunningMode::MODE_SINGLESTEP : FEXCore::Context::CoreRunningMode::MODE_RUN;
break;
case FEXCore::Config::CONFIG_GDBSERVER:
Config != 0 ? CTX->StartGdbServer() : CTX->StopGdbServer();
break;
case FEXCore::Config::CONFIG_IS64BIT_MODE:
CTX->Config.Is64BitMode = Config != 0;
break;
case FEXCore::Config::CONFIG_TSO_ENABLED:
CTX->Config.TSOEnabled = Config != 0;
break;
case FEXCore::Config::CONFIG_SMC_CHECKS:
CTX->Config.SMCChecks = Config != 0;
break;
case FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS:
CTX->Config.ABILocalFlags = Config != 0;
break;
case FEXCore::Config::CONFIG_ABI_NO_PF:
CTX->Config.ABINoPF = Config != 0;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
CTX->Config.ValidateIRarser = Config != 0;
break;
default: LogMan::Msg::A("Unknown configuration option");
char const* FindUserHomeThroughUID() {
auto passwd = getpwuid(geteuid());
if (passwd) {
return passwd->pw_dir;
}
return nullptr;
}
const char *GetHomeDirectory() {
char const *HomeDir = getenv("HOME");
// Try to get home directory from uid
if (!HomeDir) {
HomeDir = FindUserHomeThroughUID();
}
// try the PWD
if (!HomeDir) {
HomeDir = getenv("PWD");
}
// Still doesn't exit? You get local
if (!HomeDir) {
HomeDir = ".";
}
return HomeDir;
}
std::string GetConfigDirectory(bool Global) {
std::string ConfigDir;
if (Global) {
ConfigDir = GLOBAL_DATA_DIRECTORY;
}
else {
char const *HomeDir = GetHomeDirectory();
char const *ConfigXDG = getenv("XDG_CONFIG_HOME");
char const *ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (ConfigOverride) {
// Config override completely overrides the config directory
ConfigDir = ConfigOverride;
}
else {
ConfigDir = ConfigXDG ? ConfigXDG : HomeDir;
ConfigDir += "/.fex-emu/";
}
// Ensure the folder structure is created for our configuration
if (!std::filesystem::exists(ConfigDir) &&
!std::filesystem::create_directories(ConfigDir)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigDir.c_str());
// Let's go local in this case
return "./";
}
}
return ConfigDir;
}
std::string GetConfigFileLocation() {
std::string ConfigFile{};
const char *AppConfig = getenv("FEX_APP_CONFIG");
if (AppConfig) {
// App config environment variable overwrites only the config file
ConfigFile = AppConfig;
}
else {
ConfigFile = GetConfigDirectory(false) + "Config.json";
}
return ConfigFile;
}
std::string GetApplicationConfig(const std::string &Filename, bool Global) {
std::string ConfigFile = GetConfigDirectory(Global);
if (!Global &&
!std::filesystem::exists(ConfigFile) &&
!std::filesystem::create_directories(ConfigFile)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigFile.c_str());
// Let's go local in this case
return "./";
}
ConfigFile += "AppConfig/" + Filename + ".json";
return ConfigFile;
}
std::string GetDataDirectory() {
std::string DataDir{};
char const *HomeDir = GetHomeDirectory();
char const *DataXDG = getenv("XDG_DATA_HOME");
char const *DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (DataOverride) {
// Data override will override the complete directory
DataDir = DataOverride;
}
else {
DataDir = DataXDG ?: HomeDir;
DataDir += "/.fex-emu/";
}
return DataDir;
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, std::string const &Config) {
switch (Option) {
case CONFIG_ROOTFSPATH:
CTX->Config.RootFSPath = Config;
break;
case CONFIG_THUNKLIBSPATH:
CTX->Config.ThunkLibsPath = Config;
break;
case FEXCore::Config::CONFIG_DUMPIR:
CTX->Config.DumpIR = Config;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
}
uint64_t GetConfig(FEXCore::Context::Context *CTX, ConfigOption Option) {
switch (Option) {
case FEXCore::Config::CONFIG_MULTIBLOCK:
return CTX->Config.Multiblock;
break;
case FEXCore::Config::CONFIG_MAXBLOCKINST:
return CTX->Config.MaxInstPerBlock;
break;
case FEXCore::Config::CONFIG_DEFAULTCORE:
return CTX->Config.Core;
break;
case FEXCore::Config::CONFIG_VIRTUALMEMSIZE:
return CTX->Config.VirtualMemSize;
break;
case FEXCore::Config::CONFIG_SINGLESTEP:
return CTX->Config.RunningMode == FEXCore::Context::CoreRunningMode::MODE_SINGLESTEP ? 1 : 0;
case FEXCore::Config::CONFIG_GDBSERVER:
return CTX->GetGdbServerStatus();
break;
case FEXCore::Config::CONFIG_IS64BIT_MODE:
return CTX->Config.Is64BitMode;
break;
case FEXCore::Config::CONFIG_TSO_ENABLED:
return CTX->Config.TSOEnabled;
break;
case FEXCore::Config::CONFIG_SMC_CHECKS:
return CTX->Config.SMCChecks;
break;
case FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS:
return CTX->Config.ABILocalFlags;
break;
case FEXCore::Config::CONFIG_ABI_NO_PF:
return CTX->Config.ABINoPF;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
return CTX->Config.ValidateIRarser;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
return 0;
}
@@ -137,6 +154,7 @@ namespace FEXCore::Config {
private:
void MergeConfigMap(const LayerOptions &Options);
void MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value);
};
void MetaLayer::Load() {
@@ -151,10 +169,56 @@ namespace FEXCore::Config {
}
}
void MetaLayer::MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
if (MetaEnvironment == OptionMap.end()) {
// Doesn't exist, just insert
OptionMap.insert_or_assign(Option, Value);
return;
}
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
std::unordered_map<std::string, std::string> LookupMap;
const auto AddToMap = [&LookupMap](FEXCore::Config::LayerValue const &Value) {
for (const auto &EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == std::string::npos) {
// Broken environment variable
// Skip
continue;
}
auto Key = std::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = std::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
// Add the key to the map, overwriting whatever previous value was there
LookupMap.insert_or_assign(std::move(Key), std::move(Value));
}
};
AddToMap(MetaEnvironment->second);
AddToMap(Value);
// Now with the two layers merged in the map
// Add all the values to the option
Erase(Option);
for (auto &Val : LookupMap) {
// Set will emplace multiple options in to its list
Set(Option, Val.first + "=" + Val.second);
}
}
void MetaLayer::MergeConfigMap(const LayerOptions &Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto &it : Options) {
OptionMap.insert_or_assign(it.first, it.second);
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV) {
MergeEnvironmentVariables(it.first, it.second);
}
else {
OptionMap.insert_or_assign(it.first, it.second);
}
}
}
@@ -177,8 +241,91 @@ namespace FEXCore::Config {
}
}
std::string ExpandPath(std::string PathName) {
if (PathName.empty()) {
return {};
}
std::filesystem::path Path{PathName};
// Expand home if it exists
if (Path.is_relative()) {
std::string Home = getenv("HOME") ?: "";
// Home expansion only works if it is the first character
// This matches bash behaviour
if (PathName.at(0) == '~') {
PathName.replace(0, 1, Home);
return PathName;
}
// Expand relative path to absolute
Path = std::filesystem::absolute(Path);
// Only return if it exists
if (std::filesystem::exists(Path)) {
return Path;
}
}
return {};
}
void ReloadMetaLayer() {
Meta->Load();
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THREADS)) {
FEX_CONFIG_OPT(Cores, THREADS);
if (Cores == 0) {
// When the number of emulated CPU cores is zero then auto detect
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THREADS, std::to_string(get_nprocs_conf()));
}
}
auto ExpandPathIfExists = [](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
}
};
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
std::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
if (std::filesystem::exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKHOSTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKHOSTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKHOSTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKGUESTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKGUESTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKGUESTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKCONFIG)) {
FEX_CONFIG_OPT(PathName, THUNKCONFIG);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKCONFIG, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_OUTPUTLOG)) {
FEX_CONFIG_OPT(PathName, OUTPUTLOG);
if (PathName() != "stdout" && PathName() != "stderr") {
ExpandPathIfExists(FEXCore::Config::CONFIG_OUTPUTLOG, PathName());
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_SINGLESTEP)) {
// Single stepping also enforces single instruction size blocks
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, std::to_string(1u));
}
}
void AddLayer(std::unique_ptr<FEXCore::Config::Layer> _Layer) {
@@ -201,6 +348,10 @@ namespace FEXCore::Config {
Meta->Set(Option, Data);
}
void Erase(ConfigOption Option) {
Meta->Erase(Option);
}
void EraseSet(ConfigOption Option, std::string Data) {
Meta->EraseSet(Option, Data);
}
@@ -240,8 +391,14 @@ namespace FEXCore::Config {
}
}
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template int8_t Value<int8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int8_t Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template int16_t Value<int16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int16_t Default);
template uint16_t Value<uint16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint16_t Default);
template int32_t Value<int32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int32_t Default);
template uint32_t Value<uint32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint32_t Default);
template int64_t Value<int64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int64_t Default);
template uint64_t Value<uint64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint64_t Default);
// Constructor
@@ -258,5 +415,6 @@ namespace FEXCore::Config {
*List = **Value;
}
}
template void Value<std::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, std::list<std::string> *List);
}
+259
View File
@@ -0,0 +1,259 @@
{
"Options": {
"CPU": {
"Core": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigCore::CONFIG_IRJIT",
"TextDefault": "irjit",
"ShortArg": "c",
"Choices": [ "irint", "irjit", "host" ],
"ArgumentHandler": "CoreHandler",
"Desc": [
"Which CPU core to use",
"host only exists on x86_64",
"[irint, irjit, host]"
]
},
"Multiblock": {
"Type": "bool",
"Default": "true",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation"
]
},
"MaxInst": {
"Type": "int32",
"Default": "5000",
"ShortArg": "n",
"Desc": [
"Maximum number of instruction to store in a block"
]
},
"Threads": {
"Type": "uint32",
"Default": "1",
"ShortArg": "T",
"Desc": [
"Number of physical hardware threads to tell the process we have.",
"0 will auto detect."
]
}
},
"Emulation": {
"RootFS": {
"Type": "str",
"Default": "",
"ShortArg": "R",
"Desc": [
"Which Root filesystem prefix to use",
"This can be a filesystem path",
"\teg: ~/RootFS/Debian_x86_64",
"Or this can be a name of a rootfs",
"If the named rootfs exists in the FEX data folder then it will use that one",
"\teg: $HOME/.fex-emu/RootFS/<RootFS name>/",
"Or if you have XDG_DATA_HOME the config will search in that directory",
"\teg: $XDG_DATA_HOME/.fex-emu/RootFS/<RootFS name>/"
]
},
"ThunkHostLibs": {
"Type": "str",
"Default": "",
"ShortArg": "t",
"Desc": [
"Folder to find the host-side thunking libraries."
]
},
"ThunkGuestLibs": {
"Type": "str",
"Default": "",
"ShortArg": "j",
"Desc": [
"Folder to find the guest-side thunking libraries."
]
},
"ThunkConfig": {
"Type": "str",
"Default": "",
"ShortArg": "k",
"Desc": [
"A json file specifying where to overlay the thunks."
]
},
"Env": {
"Type": "strarray",
"Default": "",
"ShortArg": "E",
"Desc": [
"Adds an environment variable to the emulated environment."
]
}
},
"Debug": {
"SingleStep": {
"Type": "bool",
"Default": "false",
"ShortArg": "S",
"Desc": [
"Single stepping configuration."
]
},
"GdbServer": {
"Type": "bool",
"Default": "false",
"ShortArg": "G",
"Desc": [
"Enables the GDB server."
]
},
"DumpIR": {
"Type": "str",
"Default": "no",
"Desc": [
"Folder to dump the IR in to.",
"[no, stdout, stderr, <Folder>]"
]
},
"DumpGPRs": {
"Type": "bool",
"Default": "false",
"ShortArg": "g",
"Desc": [
"When the test harness ends, print the GPR state."
]
},
"O0": {
"Type": "bool",
"Default": "false",
"ShortArg": "O0",
"Desc": [
"Disables optimizations passes for debugging."
]
},
"Force32BitAllocator": {
"Type": "bool",
"Default": "false",
"Desc": [
"Forces use of the 32-bit allocator on 32-bit applications",
"Used to work around ulimit problems of CI runner",
"Potentially useful for debugging memory problems",
"32-bit allocator is always used if your host kernel is older than 4.17"
]
}
},
"Logging": {
"SilentLog": {
"Type": "bool",
"Default": "true",
"ShortArg": "s",
"Desc": [
"Disables logging"
]
},
"OutputLog": {
"Type": "str",
"Default": "stderr",
"ShortArg": "o",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, <Filename>]"
]
}
},
"Hacks": {
"SMCChecks": {
"Type": "uint8",
"Default": "FEXCore::Config::CONFIG_SMC_MMAN",
"TextDefault": "mman",
"ArgumentHandler": "SMCCheckHandler",
"Desc": [
"Checks code for modification before execution.",
"\tnone: No checks",
"\tmman: Invalidate on mmap, mprotect, munmap",
"\tfull: Validate code before every run (slow)"
]
},
"TSOEnabled": {
"Type": "bool",
"Default": "true",
"Desc": [
"Controls TSO IR ops.",
"Highly likely to break any multithreaded application if disabled."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
"Desc": [
"When enabled enables an optimization around flags.",
"Assumes flags are not used across cals.",
"Hand-written assembly can violate this assumption."
]
},
"ABINoPF": {
"Type": "bool",
"Default": "false",
"Desc": [
"When enabled enables an optimization around parity flag calculation.",
"Removes the calculation of the parity flag from GPR instructions.",
"Assuming no uses rely on it"
]
},
"ParanoidTSO": {
"Type": "bool",
"Default": "false",
"Desc": [
"Makes TSO operations even more strict.",
"Forces vector loadstores to also become atomic."
]
}
},
"Misc": {
"AOTIRCapture": {
"Type": "bool",
"Default": "false",
"Desc": [
"Captures IR and generates an AOT IR cache.",
"Captures both the loaded executable and libraries it loads."
]
},
"AOTIRGenerate": {
"Type": "bool",
"Default": "false",
"Desc": [
"Scans file for executable code and generates an AOT IR cache.",
"Does not run the executable."
]
},
"AOTIRLoad": {
"Type": "bool",
"Default": "false",
"Desc": [
"Loads an AOT IR cache for the loaded executable."
]
}
}
},
"UnnamedOptions": {
"Misc": {
"IS_INTERPRETER": {
"Type": "bool",
"Default": "false"
},
"INTERPRETER_INSTALLED": {
"Type": "bool",
"Default": "false"
},
"APP_FILENAME": {
"Type": "str",
"Default": ""
},
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
}
}
}
}
+42 -29
View File
@@ -14,6 +14,10 @@ namespace FEXCore::Context {
IR::InstallOpcodeHandlers(Mode);
}
void ShutdownStaticTables() {
FEXCore::Paths::ShutdownPaths();
}
FEXCore::Context::Context *CreateNewContext() {
return new FEXCore::Context::Context{};
}
@@ -23,6 +27,9 @@ namespace FEXCore::Context {
}
void DestroyContext(FEXCore::Context::Context *CTX) {
if (CTX->ParentThread) {
CTX->DestroyThread(CTX->ParentThread);
}
delete CTX;
}
@@ -47,6 +54,9 @@ namespace FEXCore::Context {
CTX->Step();
}
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
Thread->CTX->CompileBlock(Thread->CurrentFrame, GuestRIP);
}
FEXCore::Context::ExitReason RunUntilExit(FEXCore::Context::Context *CTX) {
return CTX->RunUntilExit();
@@ -65,11 +75,11 @@ namespace FEXCore::Context {
}
void GetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, &CTX->ParentThread->State.State, sizeof(FEXCore::Core::CPUState));
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void SetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(&CTX->ParentThread->State.State, State, sizeof(FEXCore::Core::CPUState));
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
void Pause(FEXCore::Context::Context *CTX) {
@@ -84,10 +94,6 @@ namespace FEXCore::Context {
CTX->CustomCPUFactory = std::move(Factory);
}
void SetFallbackCPUBackendFactory(FEXCore::Context::Context *CTX, CustomCPUFactoryType Factory) {
CTX->FallbackCPUFactory = std::move(Factory);
}
bool AddVirtualMemoryMapping([[maybe_unused]] FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
return false;
}
@@ -123,20 +129,12 @@ namespace FEXCore::Context {
CTX->StopThread(Thread);
}
void DeleteForkedThreads(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : CTX->Threads) {
if (DeadThread == Thread) {
continue;
}
void DestroyThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->DestroyThread(Thread);
}
// Setting running to false ensures that when they are shutdown we won't send signals to kill them
DeadThread->State.RunningEvents.Running = false;
}
// We now only have one thread
CTX->IdleWaitRefCount = 1;
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
@@ -147,8 +145,31 @@ namespace FEXCore::Context {
CTX->SyscallHandler = Handler;
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, [[maybe_unused]] uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function);
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function, Leaf);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->AOTIRLoader = CacheReader;
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter) {
CTX->AOTIRWriter = CacheWriter;
}
void FinalizeAOTIRCache(FEXCore::Context::Context *CTX) {
CTX->FinalizeAOTIRCache();
}
void WriteFilesWithCode(FEXCore::Context::Context *CTX, std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
CTX->WriteFilesWithCode(Writer);
}
void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name) {
return CTX->AddNamedRegion(Base, Length, Offset, Name);
}
void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length) {
return CTX->RemoveNamedRegion(Base, Length);
}
namespace Debug {
@@ -163,10 +184,6 @@ namespace Debug {
return CTX->GetRuntimeStatsForThread(Thread);
}
FEXCore::Core::CPUState GetCPUState(FEXCore::Context::Context *CTX) {
return CTX->GetCPUState();
}
bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
return CTX->GetDebugDataForRIP(RIP, Data);
}
@@ -183,10 +200,6 @@ namespace Debug {
// void SetIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir) {
// CTX->SetIRForRIP(RIP, ir);
// }
FEXCore::Core::ThreadState *GetThreadState(FEXCore::Context::Context *CTX) {
return CTX->GetThreadState();
}
}
}
+134 -42
View File
@@ -1,4 +1,5 @@
#pragma once
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
@@ -6,13 +7,24 @@
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Event.h>
#include <stdint.h>
#include <functional>
#include <istream>
#include <map>
#include <memory>
#include <mutex>
#include <optional>
#include <ostream>
#include <set>
#include <shared_mutex>
#include <unordered_map>
#include <queue>
namespace FEXCore {
class ThunkHandler;
@@ -21,7 +33,8 @@ class GdbServer;
class SiganlDelegator;
namespace CPU {
class JITCore;
class Arm64JITCore;
class X86JITCore;
}
namespace HLE {
class SyscallHandler;
@@ -30,6 +43,8 @@ class SyscallHandler;
namespace FEXCore::IR {
class RegisterAllocationPass;
class RegisterAllocationData;
class IRListView;
namespace Validation {
class IRValidation;
}
@@ -41,36 +56,75 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct AOTIRInlineEntry {
uint64_t GuestHash;
uint64_t GuestLength;
/* RAData followed by IRData */
uint8_t InlineData[0];
IR::RegisterAllocationData *GetRAData();
IR::IRListView *GetIRData();
};
struct AOTIRInlineIndexEntry {
uint64_t GuestStart;
uint64_t DataOffset;
};
struct AOTIRInlineIndex {
uint64_t Count;
uint64_t DataBase;
AOTIRInlineIndexEntry Entries[0];
AOTIRInlineEntry *Find(uint64_t GuestStart);
AOTIRInlineEntry *GetInlineEntry(uint64_t DataOffset);
};
struct AOTIRCaptureCacheEntry {
std::unique_ptr<std::ostream> Stream;
std::map<uint64_t, uint64_t> Index;
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData);
};
struct Context {
friend class FEXCore::HLE::SyscallHandler;
friend class FEXCore::CPU::JITCore;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
#ifdef JIT_X86_64
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::IR::Validation::IRValidation;
struct {
bool Multiblock {false};
bool BreakOnFrontendFailure {true};
int64_t MaxInstPerBlock {-1LL};
uint64_t VirtualMemSize {1ULL << 36};
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
FEXCore::Config::ConfigCore Core {FEXCore::Config::CONFIG_INTERPRETER};
bool GdbServer {false};
std::string RootFSPath;
std::string ThunkLibsPath;
bool Is64BitMode {true};
bool TSOEnabled {true};
bool SMCChecks {false};
bool ABILocalFlags {false};
bool ABINoPF {false};
std::string DumpIR;
uint64_t VirtualMemSize{1ULL << 36};
// this is for internal use
bool ValidateIRarser { false };
FEX_CONFIG_OPT(Multiblock, MULTIBLOCK);
FEX_CONFIG_OPT(SingleStepConfig, SINGLESTEP);
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(ABINoPF, ABINOPF);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRGenerate, AOTIRGENERATE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(Core, CORE);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(DumpIR, DUMPIR);
} Config;
using IntCallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
IntCallbackReturn InterpreterCallbackReturn;
FEXCore::HostFeatures HostFeatures;
@@ -93,9 +147,32 @@ namespace FEXCore::Context {
std::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
CustomCPUFactoryType CustomCPUFactory;
CustomCPUFactoryType FallbackCPUFactory;
std::function<void(uint64_t ThreadId, FEXCore::Context::ExitReason)> CustomExitHandler;
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
};
std::unordered_map<std::string, AOTIRCacheEntry> AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ostream>(const std::string&)> AOTIRWriter;
std::unordered_map<std::string, AOTIRCaptureCacheEntry> AOTIRCaptureCache;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
std::map<uint64_t, AddrToFileEntry> AddrToFile;
std::map<std::string, std::string> FilesWithCode;
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
@@ -108,7 +185,7 @@ namespace FEXCore::Context {
bool InitCore(FEXCore::CodeLoader *Loader);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
void Pause();
void Run();
@@ -119,7 +196,7 @@ namespace FEXCore::Context {
void StopThread(FEXCore::Core::InternalThreadState *Thread);
void SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event);
bool GetGdbServerStatus() { return (bool)DebugServer; }
bool GetGdbServerStatus() const { return DebugServer != nullptr; }
void StartGdbServer();
void StopGdbServer();
void HandleCallback(uint64_t RIP);
@@ -128,25 +205,30 @@ namespace FEXCore::Context {
static void RemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState
static void RemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveCodeEntry(Frame->Thread, GuestRIP);
}
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(uint64_t Thread);
FEXCore::Core::CPUState GetCPUState();
bool GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data);
bool FindHostCodeForRIP(uint64_t RIP, uint8_t **Code);
// XXX:
// bool FindIRForRIP(uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir);
// void SetIRForRIP(uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir);
FEXCore::Core::ThreadState *GetThreadState();
void LoadEntryList();
std::tuple<FEXCore::IR::IRListView *, FEXCore::IR::RegisterAllocationData *, uint64_t, uint64_t, uint64_t, uint64_t> GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<FEXCore::IR::IRListView<true> *, FEXCore::IR::RegisterAllocationData *, uint64_t, uint64_t> GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<void *, FEXCore::IR::IRListView<true> *, FEXCore::Core::DebugData *, FEXCore::IR::RegisterAllocationData *, bool> CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<void *, FEXCore::IR::IRListView *, FEXCore::Core::DebugData *, FEXCore::IR::RegisterAllocationData *, bool, uint64_t, uint64_t> CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
bool LoadAOTIRCache(int streamfd);
void FinalizeAOTIRCache();
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer);
// Used for thread creation from syscalls
void InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread);
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
@@ -155,38 +237,48 @@ namespace FEXCore::Context {
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
void RunThread(FEXCore::Core::InternalThreadState *Thread);
void DestroyThread(FEXCore::Core::InternalThreadState *Thread);
void CleanupAfterFork(FEXCore::Core::InternalThreadState *ExceptForThread);
std::vector<FEXCore::Core::InternalThreadState*> *const GetThreads() { return &Threads; }
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
#if ENABLE_JITSYMBOLS
FEXCore::JITSymbols Symbols;
#endif
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
private:
void WaitForIdleWithTimeout();
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void NotifyPause();
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length);
FEXCore::CodeLoader *LocalLoader{};
// Entry Cache
bool GetFilenameHash(std::string const &Filename, std::string &Hash);
void AddThreadRIPsToEntryList(FEXCore::Core::InternalThreadState *Thread);
void SaveEntryList();
std::set<uint64_t> EntryList;
std::vector<uint64_t> InitLocations;
uint64_t StartingRIP;
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
std::atomic<bool> AOTIRCaptureCacheWriteoutFlusing;
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
void AOTIRCaptureCacheWriteoutQueue_Flush();
void AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn);
bool StartPaused = false;
FEXCore::Config::Value<std::string> AppFilename{FEXCore::Config::CONFIG_APP_FILENAME, ""};
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
};
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::InternalThreadState *Thread, FEXCore::HLE::SyscallArguments *Args);
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args);
}
File diff suppressed because it is too large. Load diff
+24 -2
View File
@@ -9,6 +9,28 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t CASAL_MASK = 0x3F'E0'FC'00;
constexpr uint32_t CASAL_INST = 0x08'E0'FC'00;
bool HandleCASPAL(void *_mcontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_mcontext, void *_info, uint32_t Instr);
constexpr uint32_t ATOMIC_MEM_MASK = 0x3B200C00;
constexpr uint32_t ATOMIC_MEM_INST = 0x38200000;
constexpr uint32_t LDAXP_MASK = 0xBF'FF'80'00;
constexpr uint32_t LDAXP_INST = 0x88'7F'80'00;
constexpr uint32_t STLXP_MASK = 0xBF'E0'80'00;
constexpr uint32_t STLXP_INST = 0x88'20'80'00;
// Load ops are 4 bits
// Acquire and release bits are independent on the instruction
constexpr uint32_t ATOMIC_ADD_OP = 0b0000;
constexpr uint32_t ATOMIC_CLR_OP = 0b0001;
constexpr uint32_t ATOMIC_EOR_OP = 0b0010;
constexpr uint32_t ATOMIC_SET_OP = 0b0011;
constexpr uint32_t ATOMIC_SMAX_OP = 0b0100;
constexpr uint32_t ATOMIC_SMIN_OP = 0b0101;
constexpr uint32_t ATOMIC_UMAX_OP = 0b0110;
constexpr uint32_t ATOMIC_UMIN_OP = 0b0111;
constexpr uint32_t ATOMIC_SWAP_OP = 0b1000;
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr);
}
@@ -0,0 +1,231 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include "aarch64/cpu-aarch64.h"
namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
if (SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
SetCPUFeatures(Features);
if (!SupportsAtomics) {
WARN_ONCE("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
#endif
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
if (Is64Bit && ((~Constant)>> 16) == 0) {
movn(Reg, (~Constant) & 0xFFFF);
return;
}
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
}
}
}
void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
{x25, x26},
{x27, x28},
{x29, x30},
}};
for (auto &RegPair : CalleeSaved) {
stp(RegPair.first, RegPair.second, PairOffset);
}
// Additionally we need to store the lower 64bits of v8-v15
// Here's a fun thing, we can use two ST4 instructions to store everything
// We just need a single sub to sp before that
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v8, v9, v10, v11},
{v12, v13, v14, v15},
}};
uint32_t VectorSaveSize = sizeof(uint64_t) * 8;
sub(sp, sp, VectorSaveSize);
// SP supporting move
// We just saved x19 so it is safe
add(x19, sp, 0);
MemOperand QuadOffset(x19, 32, PostIndex);
for (auto &RegQuad : FPRs) {
st4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
}
void Arm64Emitter::PopCalleeSavedRegisters() {
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v12, v13, v14, v15},
{v8, v9, v10, v11},
}};
MemOperand QuadOffset(sp, 32, PostIndex);
for (auto &RegQuad : FPRs) {
ld4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
{x23, x24},
{x21, x22},
{x19, x20},
}};
for (auto &RegPair : CalleeSaved) {
ldp(RegPair.first, RegPair.second, PairOffset);
}
}
void Arm64Emitter::SpillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
void Arm64Emitter::FillStaticRegs() {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
void Arm64Emitter::PushDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
sub(sp, sp, SPOffset);
int i = 0;
for (auto RA : RAFPR)
{
str(RA.Q(), MemOperand(sp, i * 8));
i+=2;
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
str(RA, MemOperand(sp, i * 8));
i++;
}
#endif
str(lr, MemOperand(sp, i * 8));
}
void Arm64Emitter::PopDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
int i = 0;
for (auto RA : RAFPR)
{
ldr(RA.Q(), MemOperand(sp, i * 8));
i+=2;
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
ldr(RA, MemOperand(sp, i * 8));
i++;
}
#endif
ldr(lr, MemOperand(sp, i * 8));
add(sp, sp, SPOffset);
}
void Arm64Emitter::ResetStack() {
if (SpillSlots == 0)
return;
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
}
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetBuffer()->GetOffsetAddress<uint64_t>(GetCursorOffset());
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
nop();
}
}
}
@@ -0,0 +1,77 @@
#pragma once
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
// All but x29 are caller saved
const std::array<aarch64::Register, 16> SRA64 = {
x4, x5, x6, x7, x8, x9, x10, x11,
x12, x18, x17, x16, x15, x14, x13, x29
};
// All are callee saved
const std::array<aarch64::Register, 9> RA64 = {
x20, x21, x22, x23, x24, x25, x26, x27,
x19
};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA64Pair = {{
{x20, x21},
{x22, x23},
{x24, x25},
{x26, x27},
}};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA32Pair = {{
{w20, w21},
{w22, w23},
{w24, w25},
{w26, w27},
}};
// All are caller saved
const std::array<aarch64::VRegister, 16> SRAFPR = {
v16, v17, v18, v19, v20, v21, v22, v23,
v24, v25, v26, v27, v28, v29, v30, v31
};
// v8..v15 = (lower 64bits) Callee saved
const std::array<aarch64::VRegister, 12> RAFPR = {
/*v0, v1, v2, v3,*/v4, v5, v6, v7, // v0 ~ v3 are used as temps
v8, v9, v10, v11, v12, v13, v14, v15
};
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(size_t size);
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs();
void FillStaticRegs();
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void ResetStack();
void Align16B();
uint32_t SpillSlots{};
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
};
}
@@ -0,0 +1,26 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore::ArchHelpers::Arm64 {
#ifndef _M_ARM_64
// These are stub implementations that exist only to allow instantiating the arm64 jit
// on non arm platforms.
// Obvously such a configuration can't do the actual arm64-specific stuff
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASPAL Not Implemented");
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASAL Not Implemented");
}
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleAtomicMemOp Not Implemented");
}
#endif
}
@@ -0,0 +1,212 @@
#pragma once
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/UContext.h>
#include <signal.h>
#include <string.h>
#include <ucontext.h>
#include <stdint.h>
#include <type_traits>
namespace FEXCore::ArchHelpers::Context {
struct X86ContextBackup {
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
static constexpr int RedZoneSize = 128;
};
struct ArmContextBackup {
// Host State
uint64_t GPRs[31];
uint64_t PrevSP;
uint64_t PrevPC;
uint64_t PState;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
// Arm64 doesn't have a red zone
static constexpr int RedZoneSize = 0;
};
static inline mcontext_t* GetMContext(void* ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
return &_context->uc_mcontext;
}
#ifdef _M_ARM_64
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->sp;
}
static inline uint64_t GetPc(void* ucontext) {
return GetMContext(ucontext)->pc;
}
static inline void SetSp(void* ucontext, uint64_t val) {
GetMContext(ucontext)->sp = val;
}
static inline void SetPc(void* ucontext, uint64_t val) {
GetMContext(ucontext)->pc = val;
}
static inline uint64_t GetState(void* ucontext) {
return GetMContext(ucontext)->regs[28];
}
static inline void SetState(void* ucontext, uint64_t val) {
GetMContext(ucontext)->regs[28] = val;
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
return GetMContext(ucontext)->regs[id];
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
GetMContext(ucontext)->regs[id] = val;
}
constexpr uint32_t FPR_MAGIC = 0x46508001U;
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
using ContextBackup = ArmContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
memcpy(&Backup->GPRs[0], &_mcontext->regs[0], 31 * sizeof(uint64_t));
Backup->PrevSP = ArchHelpers::Context::GetSp(ucontext);
Backup->PrevPC = ArchHelpers::Context::GetPc(ucontext);
Backup->PState = _mcontext->pstate;
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Backup->FPRs[0], 32 * sizeof(__uint128_t));
HostState->FPCR = Backup->FPCR;
HostState->FPSR = Backup->FPSR;
// Restore GPRs and other state
_mcontext->pstate = Backup->PState;
ArchHelpers::Context::SetPc(ucontext, Backup->PrevPC);
ArchHelpers::Context::SetSp(ucontext, Backup->PrevSP);
memcpy(&_mcontext->regs[0], &Backup->GPRs[0], 31 * sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
#endif
#ifdef _M_X86_64
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_RSP];
}
static inline uint64_t GetPc(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_RIP];
}
static inline void SetSp(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_RSP] = val;
}
static inline void SetPc(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_RIP] = val;
}
static inline uint64_t GetState(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_R14];
}
static inline void SetState(void* ucontext, uint64_t val) {
GetMContext(ucontext)->gregs[REG_R14] = val;
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not impelented for x86 host");
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
ERROR_AND_DIE("Not impelented for x86 host");
}
using ContextBackup = X86ContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
memcpy(&Backup->GPRs[0], &_mcontext->gregs[0], sizeof(X86ContextBackup::GPRs));
// Copy the FPRState
memcpy(&Backup->FPRState, _mcontext->fpregs, sizeof(X86ContextBackup::FPRState));
// XXX: Save 256bit and 512bit AVX register state
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
auto _mcontext = GetMContext(ucontext);
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Backup->GPRs[0], sizeof(X86ContextBackup::GPRs));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Backup->FPRState, sizeof(X86ContextBackup::FPRState));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
}
}
#endif
} // namespace FEXCore::ArchHelpers::Context
+199 -6
View File
@@ -1,11 +1,42 @@
/*
$info$
tags: opcodes|cpuid
desc: Handles presented capability bits for guest cpu
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "git_version.h"
#include <cstring>
#ifdef _M_X86_64
#include <cpuid.h>
#endif
namespace FEXCore {
//#define CPUID_AMD
#ifdef _M_ARM_64
static uint32_t GetCycleCounterFrequency() {
uint64_t Result{};
__asm("mrs %[Res], CNTFRQ_EL0"
: [Res] "=r" (Result));
return Result;
}
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x15) {
__cpuid(0x15, eax, ebx, ecx, edx);
if (eax && ebx && ecx) {
return ecx * ebx / eax;
}
}
return 0;
}
#endif
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h() {
FEXCore::CPUID::FunctionResults Res{};
@@ -59,7 +90,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h() {
(0 << 16) | // Reserved
(0 << 17) | // Process-context identifiers
(1 << 18) | // Prefetching from memory mapped device
(0 << 19) | // SSE4.1
(1 << 19) | // SSE4.1
(0 << 20) | // SSE4.2
(0 << 21) | // X2APIC
(1 << 22) | // MOVBE
@@ -93,7 +124,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h() {
(1 << 16) | // Page Attribute Table
(1 << 17) | // 36bit page size extension
(0 << 18) | // Processor serial number
(0 << 19) | // CLFLUSH
(1 << 19) | // CLFLUSH
(0 << 20) | // Reserved
(0 << 21) | // Debug store
(0 << 22) | // Thermal monitor and software controled clock
@@ -109,6 +140,31 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h() {
return Res;
}
// 2: Cache and TLB information
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_02h() {
FEXCore::CPUID::FunctionResults Res{};
// returns default values from i7 model 1Ah
Res.eax = 0x1 | // Number of iterations needed for all descriptors
(0x5A << 8) |
(0x03 << 16) |
(0x55 << 24);
Res.ebx = 0xE4 |
(0xB2 << 8) |
(0xF0 << 16) |
(0 << 24);
Res.ecx = 0; // null descriptors
Res.edx = 0x2C |
(0x21 << 8) |
(0xCA << 16) |
(0x09 << 24);
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // Always running APIC
@@ -226,6 +282,18 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h() {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h() {
FEXCore::CPUID::FunctionResults Res{};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1;
Res.ecx = FrequencyHz;
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h() {
FEXCore::CPUID::FunctionResults Res{};
@@ -350,10 +418,120 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h() {
return Res;
}
// L1 Cache and TLB identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0005h() {
FEXCore::CPUID::FunctionResults Res{};
// L1 TLB Information for 2MB and 4MB pages
Res.eax =
(64 << 0) | // Number of TLB instruction entries
(255 << 8) | // instruction TLB associativity type (full)
(64 << 16) | // Number of TLB data entries
(255 << 24); // data TLB associativity type (full)
// L1 TLB Information for 4KB pages
Res.ebx =
(64 << 0) | // Number of TLB instruction entries
(255 << 8) | // instruction TLB associativity type (full)
(64 << 16) | // Number of TLB data entries
(255 << 24); // data TLB associativity type (full)
// L1 data cache identifiers
Res.ecx =
(64 << 0) | // L1 data cache size line in bytes
(1 << 8) | // L1 data cachelines per tag
(8 << 16) | // L1 data cache associativity
(32 << 24); // L1 data cache size in KB
// L1 instruction cache identifiers
Res.edx =
(64 << 0) | // L1 instruction cache line size in bytes
(1 << 8) | // L1 instruction cachelines per tag
(4 << 16) | // L1 instruction cache associativity
(64 << 24); // L1 instruction cache size in KB
return Res;
}
// L2 Cache identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0006h() {
FEXCore::CPUID::FunctionResults Res{};
// L2 TLB Information for 2MB and 4MB pages
Res.eax =
(1024 << 0) | // Number of TLB instruction entries
(6 << 12) | // instruction TLB associativity type
(1536 << 16) | // Number of TLB data entries
(3 << 28); // data TLB associativity type
// L2 TLB Information for 4KB pages
Res.ebx =
(1024 << 0) | // Number of TLB instruction entries
(6 << 12) | // instruction TLB associativity type
(1536 << 16) | // Number of TLB data entries
(5 << 28); // data TLB associativity type
// L2 cache identifiers
Res.ecx =
(64 << 0) | // cacheline size
(1 << 8) | // cachelines per tag
(6 << 12) | // cache associativity
(512 << 16); // L2 cache size in KB
// L3 cache identifiers
Res.edx =
(64 << 0) | // cacheline size
(1 << 8) | // cachelines per tag
(6 << 12) | // cache associativity
(16 << 18); // L2 cache size in KB
return Res;
}
// Advanced power management
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0007h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // APIC timer not affected by p-state
Res.edx =
(1 << 8); // Invariant TSC
return Res;
}
// Virtual and physical address sizes
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax =
(48 << 0) | // PhysAddrSize = 48-bit
(48 << 8) | // LinAddrSize = 48-bit
(0 << 16); // GuestPhysAddrSize == PhysAddrSize
Res.ebx =
(0 << 2) | // XSaveErPtr: Saving and restoring error pointers
(0 << 1) | // IRPerf: Instructions retired count support
(0 << 0); // CLZERO support
uint32_t CoreCount = Cores() - 1;
Res.ecx =
(0 << 16) | // PerfTscSize: Performance timestamp count size
(0 << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
// TLB 1GB page identifiers
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0019h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax =
(0xF << 28) | // L1 DTLB associativity for 1GB pages
(64 << 16) | // L1 DTLB entry count for 1GB pages
(0xF << 12) | // L1 ITLB associativity for 1GB pages
(64 << 0); // L1 ITLB entry count for 1GB pages
Res.ebx =
(0 << 28) | // L2 DTLB associativity for 1GB pages
(0 << 16) | // L2 DTLB entry count for 1GB pages
(0 << 12) | // L2 ITLB associativity for 1GB pages
(0 << 0); // L2 ITLB entry count for 1GB pages
return Res;
}
@@ -366,7 +544,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
CTX = ctx;
RegisterFunction(0, std::bind(&CPUIDEmu::Function_0h, this));
RegisterFunction(1, std::bind(&CPUIDEmu::Function_01h, this));
// 2: Cache and TLB information
RegisterFunction(2, std::bind(&CPUIDEmu::Function_02h, this));
// 3: Serial Number(previously), now reserved
// 4: Deterministic cache parameters for each level
// 5: Monitor/mwait
@@ -383,7 +561,11 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x12: Intel SGX capability enumeration
// 0x13: Reserved
// 0x14: Intel Processor trace
// 0x15: Timestamp counter information
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
RegisterFunction(0x15, std::bind(&CPUIDEmu::Function_15h, this));
#endif
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
@@ -398,12 +580,23 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// Processor brand string continued
RegisterFunction(0x8000'0004, std::bind(&CPUIDEmu::Function_8000_0004h, this));
// 0x8000'0005: L1 Cache and TLB identifiers
#ifdef CPUID_AMD
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_8000_0005h, this));
#else
// This is full reserved on Intel platforms
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_Reserved, this));
#endif
// 0x8000'0006: L2 Cache identifiers
RegisterFunction(0x8000'0006, std::bind(&CPUIDEmu::Function_8000_0006h, this));
// Advanced power management information
RegisterFunction(0x8000'0007, std::bind(&CPUIDEmu::Function_8000_0007h, this));
// 0x8000'0008: Virtual and physical address sizes
// Virtual and physical address sizes
RegisterFunction(0x8000'0008, std::bind(&CPUIDEmu::Function_8000_0008h, this));
// 0x8000'000A: SVM Revision
// 0x8000'0019: TLB 1GB page identifiers
// TLB 1GB page identifiers
RegisterFunction(0x8000'0019, std::bind(&CPUIDEmu::Function_8000_0019h, this));
// 0x8000'001A: Performance optimization identifiers
// 0x8000'001B: Instruction based sampling identifiers
// 0x8000'001C: Lightweight profiling capabilities
+16 -3
View File
@@ -3,6 +3,8 @@
#include <unordered_map>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore {
namespace Context {
@@ -22,16 +24,21 @@ private:
public:
void Init(FEXCore::Context::Context *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function) {
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, [[maybe_unused]] uint32_t Leaf) {
auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end())
if (Handler == FunctionHandlers.end()) {
#ifndef NDEBUG
LogMan::Msg::E("Unhandled CPU ID function, 0x%x", Function);
#endif
return Function_Reserved();
}
return Handler->second();
}
private:
FEXCore::Context::Context *CTX;
FEX_CONFIG_OPT(Cores, THREADS);
using FunctionHandler = std::function<FEXCore::CPUID::FunctionResults()>;
void RegisterFunction(uint32_t Function, FunctionHandler Handler) {
@@ -43,15 +50,21 @@ private:
// Functions
FEXCore::CPUID::FunctionResults Function_0h();
FEXCore::CPUID::FunctionResults Function_01h();
FEXCore::CPUID::FunctionResults Function_02h();
FEXCore::CPUID::FunctionResults Function_06h();
FEXCore::CPUID::FunctionResults Function_07h();
FEXCore::CPUID::FunctionResults Function_15h();
FEXCore::CPUID::FunctionResults Function_8000_0000h();
FEXCore::CPUID::FunctionResults Function_8000_0001h();
FEXCore::CPUID::FunctionResults Function_8000_0002h();
FEXCore::CPUID::FunctionResults Function_8000_0003h();
FEXCore::CPUID::FunctionResults Function_8000_0004h();
FEXCore::CPUID::FunctionResults Function_8000_0005h();
FEXCore::CPUID::FunctionResults Function_8000_0006h();
FEXCore::CPUID::FunctionResults Function_8000_0007h();
FEXCore::CPUID::FunctionResults Function_8000_0008h();
FEXCore::CPUID::FunctionResults Function_8000_0009h();
FEXCore::CPUID::FunctionResults Function_8000_0019h();
FEXCore::CPUID::FunctionResults Function_Reserved();
};
}
+17 -13
View File
@@ -5,6 +5,12 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore {
static void* ThreadHandler(void *Arg) {
FEXCore::CompileService *This = reinterpret_cast<FEXCore::CompileService*>(Arg);
This->ExecutionThread();
return nullptr;
}
CompileService::CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ParentThread {Thread} {
@@ -16,9 +22,7 @@ namespace FEXCore {
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
WorkerThread = std::thread([this]() {
ExecutionThread();
});
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
}
void CompileService::Initialize() {
@@ -30,7 +34,7 @@ namespace FEXCore {
ShuttingDown = true;
// Kick the working thread
StartWork.NotifyAll();
WorkerThread.join();
WorkerThread->join(nullptr);
}
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread) {
@@ -57,9 +61,7 @@ namespace FEXCore {
}
}
LogMan::Throw::A(CompileThreadData->IRLists.size() == 0, "Compile service must never have IRLists");
LogMan::Throw::A(CompileThreadData->RALists.size() == 0, "Compile service must never have RALists");
LogMan::Throw::A(CompileThreadData->DebugData.size() == 0, "Compile service must never have DebugData");
LOGMAN_THROW_A(CompileThreadData->LocalIRCache.size() == 0, "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
@@ -94,7 +96,7 @@ namespace FEXCore {
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->State.ThreadManager.TID.load());
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
pthread_setname_np(pthread_self(), ThreadName);
while (true) {
@@ -122,15 +124,15 @@ namespace FEXCore {
// If we had a work item then work on it
if (Item) {
// Make sure it's not in lookup cache by accident
LogMan::Throw::A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
LOGMAN_THROW_A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->State.State.rip = Item->RIP;
CompileThreadData->CurrentFrame->State.rip = Item->RIP;
auto [CodePtr, IRList, DebugData, RAData, Generated] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
LogMan::Throw::A(Generated == true, "Compile Service doesn't have IR Cache");
LOGMAN_THROW_A(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
@@ -141,6 +143,8 @@ namespace FEXCore {
Item->IRList = IRList;
Item->DebugData = DebugData;
Item->RAData = RAData;
Item->StartAddr = StartAddr;
Item->Length = Length;
GCArray.emplace_back(Item);
Item->ServiceWorkDone.NotifyAll();
+8 -3
View File
@@ -2,6 +2,7 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <memory>
#include <thread>
@@ -31,9 +32,11 @@ class CompileService final {
// Outgoing
void *CodePtr{};
FEXCore::IR::IRListView<true> *IRList{};
FEXCore::IR::IRListView *IRList{};
FEXCore::IR::RegisterAllocationData *RAData{};
FEXCore::Core::DebugData *DebugData{};
uint64_t StartAddr;
uint64_t Length;
// Communication
Event ServiceWorkDone{};
@@ -43,12 +46,14 @@ class CompileService final {
WorkItem *CompileCode(uint64_t RIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread);
// Public for threading
void ExecutionThread();
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
void ExecutionThread();
std::thread WorkerThread;
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::unique_ptr<FEXCore::Core::InternalThreadState> CompileThreadData;
std::mutex QueueMutex{};
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,365 @@
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <bit>
#include <cmath>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
auto Buffer = GetBuffer();
DispatchPtr = Buffer->GetOffsetAddress<CPUBackend::AsmDispatch>(GetCursorOffset());
// while (true) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// Ptr();
// }
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
Literal l_VirtualMemory {VirtualMemorySize};
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_L1Ptr {Thread->LookupCache->GetL1Pointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
Literal l_ExitFunctionLink {config.ExitFunctionLink};
Literal l_ExitFunctionLinkThis {config.ExitFunctionLinkThis};
// Push all the register we need to save
PushCalleeSavedRegisters();
// Push our memory base to the correct register
// Move our thread pointer to the correct register
// This is passed in to parameter 0 (x0)
mov(STATE, x0);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
AbsoluteLoopTopAddressFillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled) {
FillStaticRegs();
}
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
aarch64::Label FullLookup{};
aarch64::Label CallBlock{};
aarch64::Label LoopTop{};
aarch64::Label ExitSpillSRA{};
aarch64::Label ThreadPauseHandler{};
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
auto RipReg = x2;
// L1 Cache
ldr(x0, &l_L1Ptr);
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
ldp(x3, x0, MemOperand(x0));
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
b(&CallBlock);
}
// L1C check failed, do a full lookup
bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, &l_PagePtr);
// Mask the address by the virtual address size so we can check for aliases
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, Thread->LookupCache->GetVirtualMemorySize() - 1);
}
else {
ldr(x3, &l_VirtualMemory);
and_(x3, RipReg, x3);
}
aarch64::Label NoBlock;
{
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
// Load the pointer from the offset
ldr(x0, MemOperand(x0, x1, Shift::LSL, 3));
// If page pointer is zero then we have no block
cbz(x0, &NoBlock);
// Steal the page offset
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x3, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x3, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, &l_L1Ptr);
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
// Jump to the block
if (!config.ExecuteBlocksWithCall) {
br(x3);
} else {
bind(&CallBlock);
mov(x0, STATE);
blr(x3);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, &l_CTX);
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
} else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
}
}
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
PopCalleeSavedRegisters();
// Return from the function
// LR is set to the correct return location now
ret();
}
{
ExitFunctionLinkerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
ldr(x3, &l_ExitFunctionLink);
blr(x3);
if (SRAEnabled)
FillStaticRegs();
br(x0);
}
// Need to create the block
{
bind(&NoBlock);
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileBlock);
if (SRAEnabled)
SpillStaticRegs();
// X2 contains our guest RIP
blr(x3); // { CTX, Frame, RIP}
if (SRAEnabled)
FillStaticRegs();
b(&LoopTop);
}
{
SignalHandlerReturnAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// Now to get back to our old location we need to do a fault dance
// We can't use SIGTRAP here since gdb catches it and never gives it to the application!
hlt(0);
}
{
ThreadPauseHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
if (SRAEnabled)
SpillStaticRegs();
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
// Call our sleep handler
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x2, &l_Sleep);
blr(x2);
PauseReturnInstruction = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// Fault to start running again
hlt(0);
}
{
// The expectation here is that a thunked function needs to call back in to the JIT in a reentrant safe way
// To do this safely we need to do some state tracking and register saving
//
// eg:
// JIT Call->
// Thunk->
// Thunk callback->
//
// The thunk callback needs to execute JIT code and when it returns, it needs to safely return to the thunk rather than JIT space
// This is handled by pushing a return address trampoline to the stack so when the guest address returns it hits our custom thunk return
// - This will safely return us to the thunk
//
// On return to the thunk, the thunk can get whatever its return value is from the thread context depending on ABI handling on its end
// When the thunk itself returns, it'll do its regular return logic there
// void ReentrantCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
CallbackPtr = Buffer->GetOffsetAddress<CPUBackend::JITCallback>(GetCursorOffset());
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
// First thing we need to move the thread state pointer back in to our register
mov(STATE, x0);
// Make sure to adjust the refcounter so we don't clear the cache now
LoadConstant(x0, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
ldr(w2, MemOperand(x0));
add(w2, w2, 1);
str(w2, MemOperand(x0));
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// load static regs
if (SRAEnabled)
FillStaticRegs();
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
place(&l_VirtualMemory);
place(&l_PagePtr);
place(&l_L1Ptr);
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
place(&l_ExitFunctionLink);
place(&l_ExitFunctionLinkThis);
FinalizeCode();
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
#endif
}
void Arm64Dispatcher::SpillSRA(void *ucontext) {
for(int i = 0; i < SRA64.size(); i++) {
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
// TODO: Also recover FPRs, not sure where the neon context is
// This is usually not needed
/*
for(int i = 0; i < SRAFPR.size(); i++) {
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
}
*/
}
#ifdef _M_ARM_64
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Dispatcher = std::make_unique<Arm64Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
}
#endif
}
@@ -0,0 +1,18 @@
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
protected:
void SpillSRA(void *ucontext) override;
};
}
@@ -0,0 +1,367 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Common/MathUtils.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <FEXCore/Core/X86Enums.h>
namespace FEXCore::CPU {
void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ArchHelpers::Context::ContextBackup);
// We need to back up behind the host's red zone
// We do this on the guest side as well
// (does nothing on arm hosts)
NewSP -= ArchHelpers::Context::ContextBackup::RedZoneSize;
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
ArchHelpers::Context::BackupContext(ucontext, Context);
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, ThreadState->CurrentFrame, sizeof(FEXCore::Core::CPUState));
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
SignalFrames.push(NewSP);
}
void Dispatcher::RestoreThreadState(void *ucontext) {
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uintptr_t NewSP = OldSP;
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(ThreadState->CurrentFrame, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
StoreThreadState(Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
uint64_t OldGuestSP = Frame->State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
if (GuestAction->sa_flags & SA_SIGINFO) {
if (SRAEnabled) {
if (!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
} else {
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
}
}
// Setup ucontext a bit
if (CTX->Config.Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::ucontext_t);
uint64_t UContextLocation = NewGuestSP;
NewGuestSP -= sizeof(siginfo_t);
uint64_t SigInfoLocation = NewGuestSP;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
siginfo_t *guest_siginfo = reinterpret_cast<siginfo_t*>(SigInfoLocation);
// We have extended float information
guest_uctx->uc_flags |= FEXCore::x86_64::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = &guest_uctx->__fpregs_mem;
#define COPY_REG(x) \
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x] = Frame->State.gregs[X86State::REG_##x];
COPY_REG(R8);
COPY_REG(R9);
COPY_REG(R10);
COPY_REG(R11);
COPY_REG(R12);
COPY_REG(R13);
COPY_REG(R14);
COPY_REG(R15);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
// Copy float registers
memcpy(guest_uctx->__fpregs_mem._st, Frame->State.mm, sizeof(Frame->State.mm));
memcpy(guest_uctx->__fpregs_mem._xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = Frame->State.FCW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
(Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] << 11) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] << 8) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] << 9) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] << 10) |
(Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] << 14);
// Copy over signal stack information
guest_uctx->uc_stack.ss_flags = GuestStack->ss_flags;
guest_uctx->uc_stack.ss_sp = GuestStack->ss_sp;
guest_uctx->uc_stack.ss_size = GuestStack->ss_size;
// siginfo_t
siginfo_t *HostSigInfo = reinterpret_cast<siginfo_t*>(info);
if (HostSigInfo->si_code == SI_USER) {
// If the signal was a user signal then we need to pass this struct through unaltered
// Guest might be doing something with it
*guest_siginfo = *HostSigInfo;
}
else {
guest_siginfo->si_signo = Signal;
switch (Signal) {
case SIGSEGV:
case SIGBUS:
guest_siginfo->si_code = HostSigInfo->si_code;
guest_siginfo->si_errno = HostSigInfo->si_errno;
// Macro expansion to get the si_addr
guest_siginfo->si_addr = HostSigInfo->si_addr;
break;
default: LogMan::Msg::D("Unhandled siginfo_t signal: %d", Signal); break;
}
}
Frame->State.gregs[X86State::REG_RSI] = SigInfoLocation;
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
// XXX: 32bit Support
NewGuestSP -= sizeof(FEXCore::x86::ucontext_t);
uint64_t UContextLocation = 0; // NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::siginfo_t);
uint64_t SigInfoLocation = 0; // NewGuestSP;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = UContextLocation;
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SigInfoLocation;
}
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
Frame->State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
if (CTX->Config.Is64BitMode) {
Frame->State.gregs[X86State::REG_RDI] = Signal;
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
}
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
LOGMAN_THROW_A(CTX->X86CodeGen.SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[X86State::REG_RSP] = NewGuestSP;
}
return true;
}
bool Dispatcher::HandleSIGILL(int Signal, void *info, void *ucontext) {
if (ArchHelpers::Context::GetPc(ucontext) == SignalHandlerReturnAddress) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
if (ArchHelpers::Context::GetPc(ucontext) == PauseReturnInstruction) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
return true;
}
return false;
}
bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = ThreadState->SignalReason.load();
auto Frame = ThreadState->CurrentFrame;
if (SignalReason == FEXCore::Core::SignalEvent::Pause) {
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
}
// Set the new PC
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
// Ref count our faults
// We use this to track if it is safe to clear cache
++SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::Stop) {
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the core and get out safely
ArchHelpers::Context::SetSp(ucontext, Frame->ReturningStackLocation);
// Our ref counting doesn't matter anymore
SignalHandlerRefCounter = 0;
// Set the new PC
if (SRAEnabled && IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), false)) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
}
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::Return) {
RestoreThreadState(ucontext);
// Ref count our faults
// We use this to track if it is safe to clear cache
--SignalHandlerRefCounter;
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
return false;
}
uint64_t Dispatcher::GetCompileBlockPtr() {
using ClassPtrType = void (FEXCore::Context::Context::*)(FEXCore::Core::CpuStateFrame *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast CompileBlockPtr;
CompileBlockPtr.ClassPtr = &FEXCore::Context::Context::CompileBlockJit;
return CompileBlockPtr.Data;
}
void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
for (auto iter = CodeBuffers.begin(); iter != CodeBuffers.end(); ++iter) {
auto [start, end] = *iter;
if (start == reinterpret_cast<uint64_t>(start_to_remove)) {
CodeBuffers.erase(iter);
return;
}
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher) {
return IsAddressInDispatcher(Address);
}
return false;
}
}
@@ -0,0 +1,85 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "Interface/Context/Context.h"
#include <stack>
namespace FEXCore::CPU {
struct DispatcherConfig {
bool ExecuteBlocksWithCall = false;
uintptr_t ExitFunctionLink = 0;
uintptr_t ExitFunctionLinkThis = 0;
bool StaticRegisterAssignment = false;
};
class Dispatcher {
public:
virtual ~Dispatcher() = default;
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
/**
* @name Dispatch Helper functions
* @{ */
uint64_t ThreadStopHandlerAddress{};
uint64_t ThreadStopHandlerAddressSpillSRA{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t AbsoluteLoopTopAddressFillSRA{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
uint64_t Start{};
uint64_t End{};
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
void RegisterCodeBuffer(uint8_t* start, size_t size) {
CodeBuffers.emplace_back(reinterpret_cast<uint64_t>(start),
reinterpret_cast<uint64_t>(start + size));
}
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
protected:
Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ThreadState {Thread} {}
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
private:
std::vector<std::tuple<uint64_t, uint64_t>> CodeBuffers; // Start, End
};
}
@@ -0,0 +1,320 @@
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE, nullptr, this) {
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
// r10, r11
//
// Callee Saved
// rbx, rbp, r12, r13, r14, r15
//
// 1St Argument: rdi <ThreadState>
// XMM:
// All temp
// while (true) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Bunch of exit state stuff
// x86-64 ABI has the stack aligned when /call/ happens
// Which means the destination has a misaligned stack at that point
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
mov(STATE, rdi);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)], rsp);
Label LoopTop;
Label FullLookup;
Label CallBlock;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
L(LoopTop);
AbsoluteLoopTopAddressFillSRA = AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
// L1 Cache
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
jne(FullLookup);
if (!config.ExecuteBlocksWithCall) {
jmp(qword[r13 + rax + 0]);
} else {
mov(rax, qword[r13 + rax + 0]);
jmp(CallBlock);
}
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
// Load page pointer
mov(rdi, qword [r13 + rax * 8]);
cmp(rdi, 0);
je(NoBlock);
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
cmp(rcx, rdx);
jne(NoBlock);
// Load the block pointer
mov(rax, qword [rdi + rax]);
cmp(rax, 0);
je(NoBlock);
// Update L1
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
mov(qword[r13 + rcx*8 + 8], rdx);
mov(qword[r13 + rcx*8 + 0], rax);
// Real block if we made it here
if (!config.ExecuteBlocksWithCall) {
jmp(rax);
} else {
L(CallBlock);
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
}
{
L(ExitBlock);
ThreadStopHandlerAddress = getCurr<uint64_t>();
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
// Block creation
{
L(NoBlock);
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
mov(rax, GetCompileBlockPtr());
call(rax);
// rdx already contains RIP here
jmp(LoopTop);
}
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
// {rdi, rsi, rdx}
mov(rdi, config.ExitFunctionLinkThis);
mov(rsi, STATE);
mov(rdx, rax); // rax is set at the block end
mov(rax, config.ExitFunctionLink);
call(rax);
jmp(rax);
}
{
// Pause handler
ThreadPauseHandlerAddress = getCurr<uint64_t>();
L(ThreadPauseHandler);
mov(rdi, reinterpret_cast<uintptr_t>(CTX));
mov(rsi, STATE);
mov(rax, reinterpret_cast<uint64_t>(SleepThread));
call(rax);
// XXX: Unsupported atm
PauseReturnInstruction = getCurr<uint64_t>();
ud2();
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
// First thing we need to move the thread state pointer back in to our register
mov(STATE, rdi);
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
add(dword [rax], 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
mov(rax, CTX->X86CodeGen.CallbackReturn);
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])]);
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], rsi);
// Back to the loop top now
jmp(LoopTop);
}
{
// Signal return handler
SignalHandlerReturnAddress = getCurr<uint64_t>();
ud2();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
// rsi = rsp
mov(rsp, rsi);
// Now jump back to the thunk
// XXX: XMM?
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
ready();
Start = reinterpret_cast<uint64_t>(getCode());
End = Start + getSize();
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(Start), End-Start, Name);
#endif
}
X86Dispatcher::~X86Dispatcher() {
}
#ifdef _M_X86_64
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
DispatcherConfig config;
config.ExecuteBlocksWithCall = true;
Dispatcher = std::make_unique<X86Dispatcher>(ctx, Thread, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Dispatcher->ReturnPtr;
}
#endif
}
@@ -0,0 +1,27 @@
#pragma once
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Utils/Allocator.h>
#define XBYAK64
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator, public Xbyak::Allocator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
virtual ~X86Dispatcher() override;
// Xbyak::Allocator
Xbyak::uint8 *alloc(size_t size) override { Size = size; return reinterpret_cast<uint8_t*>(FEXCore::Allocator::mmap(nullptr, size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0)); }
void free(Xbyak::uint8 *p) override { FEXCore::Allocator::munmap(p, Size); }
bool useProtect() const override { return false; }
private:
size_t Size{};
};
}
+126 -80
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: frontend|x86-meta-blocks
desc: Extracts instruction & block meta info, frontend multiblock logic
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/InternalThreadState.h"
@@ -8,6 +15,7 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/LogManager.h>
#include <set>
namespace FEXCore::Frontend {
using namespace FEXCore::X86Tables;
@@ -117,38 +125,29 @@ Decoder::Decoder(FEXCore::Context::Context *ctx)
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LogMan::Throw::A(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
LOGMAN_THROW_A(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
}
uint8_t Decoder::PeekByte(uint8_t Offset) {
uint8_t Decoder::PeekByte(uint8_t Offset) const {
uint8_t Byte = InstStream[InstructionSize + Offset];
return Byte;
}
uint64_t Decoder::ReadData(uint8_t Size) {
uint64_t Res{};
#define READ_DATA(x, y) \
case x: { \
y const *Data = reinterpret_cast<y const*>(&InstStream[InstructionSize]); \
Res = *Data; \
} \
break
switch (Size) {
case 0: return 0;
READ_DATA(1, uint8_t);
READ_DATA(2, uint16_t);
case 3: memcpy(&Res, &InstStream[InstructionSize], Size);
READ_DATA(4, uint32_t);
READ_DATA(8, uint64_t);
default:
LogMan::Msg::A("Unknown data size to read");
return 0;
if (Size == 0) {
return 0;
}
#undef READ_DATA
if (Size > sizeof(uint64_t)) {
LOGMAN_MSG_A("Unknown data size to read");
return 0;
}
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
#ifndef NDEBUG
for(size_t i = 0; i < Size; ++i) {
@@ -157,6 +156,7 @@ uint64_t Decoder::ReadData(uint8_t Size) {
#else
SkipBytes(Size);
#endif
return Res;
}
@@ -197,9 +197,9 @@ void Decoder::DecodeModRM_16(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
}
Operand->TypeSIB.Type = DecodedOperand::TYPE_SIB;
Operand->TypeSIB.Scale = 1;
Operand->TypeSIB.Offset = Literal;
Operand->Type = DecodedOperand::OpType::SIB;
Operand->Data.SIB.Scale = 1;
Operand->Data.SIB.Offset = Literal;
// Only called when ModRM.mod != 0b11
struct Encodings {
@@ -238,8 +238,8 @@ void Decoder::DecodeModRM_16(X86Tables::DecodedOperand *Operand, X86Tables::ModR
uint8_t LookupIndex = ModRM.mod << 3 | ModRM.rm;
auto it = Lookup[LookupIndex];
Operand->TypeSIB.Base = it.Base;
Operand->TypeSIB.Index = it.Index;
Operand->Data.SIB.Base = it.Base;
Operand->Data.SIB.Index = it.Index;
}
void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM) {
@@ -277,21 +277,21 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
// SIB
Operand->TypeSIB.Type = DecodedOperand::TYPE_SIB;
Operand->TypeSIB.Scale = 1 << SIB.scale;
Operand->Type = DecodedOperand::OpType::SIB;
Operand->Data.SIB.Scale = 1 << SIB.scale;
// The invalid encoding types are described at Table 1-12. "promoted nsigned is always non-zero"
Operand->TypeSIB.Index = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X ? 1 : 0, SIB.index, false, false, false, false, 0b100);
Operand->TypeSIB.Base = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
Operand->Data.SIB.Index = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X ? 1 : 0, SIB.index, false, false, false, false, 0b100);
Operand->Data.SIB.Base = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
uint64_t Literal {0};
LogMan::Throw::A(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
LOGMAN_THROW_A(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
Literal = ReadData(Displacement);
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
}
Operand->TypeSIB.Offset = Literal;
Operand->Data.SIB.Offset = Literal;
}
else if (ModRM.mod == 0) {
// Explained in Table 1-14. "Operand Addressing Using ModRM and SIB Bytes"
@@ -300,13 +300,13 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
uint32_t Literal;
Literal = ReadData(4);
Operand->TypeRIPLiteral.Type = DecodedOperand::TYPE_RIP_RELATIVE;
Operand->TypeRIPLiteral.Literal.u = Literal;
Operand->Type = DecodedOperand::OpType::RIPRelative;
Operand->Data.RIPLiteral.Value.u = Literal;
}
else {
// Register-direct addressing
Operand->TypeGPR.Type = DecodedOperand::TYPE_GPR_DIRECT;
Operand->TypeGPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
Operand->Type = DecodedOperand::OpType::GPRDirect;
Operand->Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
}
}
else {
@@ -318,9 +318,9 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
Displacement = DisplacementSize;
Operand->TypeGPRIndirect.Type = DecodedOperand::TYPE_GPR_INDIRECT;
Operand->TypeGPRIndirect.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
Operand->TypeGPRIndirect.Displacement = Literal;
Operand->Type = DecodedOperand::OpType::GPRIndirect;
Operand->Data.GPRIndirect.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
Operand->Data.GPRIndirect.Displacement = Literal;
}
}
@@ -344,7 +344,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
return false;
}
LogMan::Throw::A(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
LOGMAN_THROW_A(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
uint8_t DestSize{};
@@ -460,22 +460,25 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ||
HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RDX)) {
// Some instructions hardcode their destination as RAX
CurrentDest->TypeGPR.Type = DecodedOperand::TYPE_GPR;
CurrentDest->TypeGPR.HighBits = false;
CurrentDest->TypeGPR.GPR = HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ? FEXCore::X86State::REG_RAX : FEXCore::X86State::REG_RDX;
CurrentDest->Type = DecodedOperand::OpType::GPR;
CurrentDest->Data.GPR.HighBits = false;
CurrentDest->Data.GPR.GPR = HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ? FEXCore::X86State::REG_RAX : FEXCore::X86State::REG_RDX;
CurrentDest = &DecodeInst->Src[0];
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LogMan::Throw::A(!HasMODRM, "This instruction shouldn't have ModRM!");
LOGMAN_THROW_A(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
// This also means that the destination is always a GPR on these ones
// ADDITIONALLY:
// If there is a REX prefix then that allows extended GPR usage
CurrentDest->TypeGPR.Type = DecodedOperand::TYPE_GPR;
DecodeInst->Dest.TypeGPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100) || HasHighXMM;
CurrentDest->TypeGPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
CurrentDest->Type = DecodedOperand::OpType::GPR;
DecodeInst->Dest.Data.GPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100) || HasHighXMM;
CurrentDest->Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
return false;
}
uint8_t Bytes = Info->MoreBytes;
@@ -498,55 +501,63 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
ModRM.Hex = DecodeInst->ModRM;
// Decode the GPR source first
GPR.TypeGPR.Type = DecodedOperand::TYPE_GPR;
GPR.TypeGPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX) || HasHighXMM;
GPR.TypeGPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_R ? 1 : 0, ModRM.reg, GPR8Bit, HasREX, HasXMMGPR, HasMMGPR);
GPR.Type = DecodedOperand::OpType::GPR;
GPR.Data.GPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX) || HasHighXMM;
GPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_R ? 1 : 0, ModRM.reg, GPR8Bit, HasREX, HasXMMGPR, HasMMGPR);
if (GPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
return false;
// ModRM.mod == 0b11 == Register
// ModRM.Mod != 0b11 == Register-direct addressing
if (ModRM.mod == 0b11) {
NonGPR.TypeGPR.Type = DecodedOperand::TYPE_GPR;
NonGPR.TypeGPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX) || HasHighXMM;
NonGPR.TypeGPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, NonGPR8Bit, HasREX, HasXMMNonGPR, HasMMNonGPR);
NonGPR.Type = DecodedOperand::OpType::GPR;
NonGPR.Data.GPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX) || HasHighXMM;
NonGPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, NonGPR8Bit, HasREX, HasXMMNonGPR, HasMMNonGPR);
if (NonGPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
return false;
}
else {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&NonGPR, ModRM);
}
return true;
};
size_t CurrentSrc = 0;
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM) {
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SF_MOD_DST) {
ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest);
if (!ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest))
return false;
}
else {
ModRMOperand(DecodeInst->Dest, DecodeInst->Src[CurrentSrc], HasXMMDst, HasXMMSrc, HasMMDst, HasMMSrc, Is8BitDest, Is8BitSrc);
if (!ModRMOperand(DecodeInst->Dest, DecodeInst->Src[CurrentSrc], HasXMMDst, HasXMMSrc, HasMMDst, HasMMSrc, Is8BitDest, Is8BitSrc))
return false;
}
++CurrentSrc;
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_RAX)) {
DecodeInst->Src[CurrentSrc].TypeGPR.Type = DecodedOperand::TYPE_GPR;
DecodeInst->Src[CurrentSrc].TypeGPR.HighBits = false;
DecodeInst->Src[CurrentSrc].TypeGPR.GPR = FEXCore::X86State::REG_RAX;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = FEXCore::X86State::REG_RAX;
++CurrentSrc;
}
else if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_RCX)) {
DecodeInst->Src[CurrentSrc].TypeGPR.Type = DecodedOperand::TYPE_GPR;
DecodeInst->Src[CurrentSrc].TypeGPR.HighBits = false;
DecodeInst->Src[CurrentSrc].TypeGPR.GPR = FEXCore::X86State::REG_RCX;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = FEXCore::X86State::REG_RCX;
++CurrentSrc;
}
if (Bytes != 0) {
LogMan::Throw::A(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
LOGMAN_THROW_A(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].TypeLiteral.Size = Bytes;
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
uint64_t Literal {0};
Literal = ReadData(Bytes);
uint64_t Literal = ReadData(Bytes);
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT) ||
(DecodeFlags::GetSizeDstFlags(DecodeInst->Flags) == DecodeFlags::SIZE_64BIT && Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT64BIT)) {
@@ -559,15 +570,15 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
else {
Literal = static_cast<int32_t>(Literal);
}
DecodeInst->Src[CurrentSrc].TypeLiteral.Size = DestSize;
DecodeInst->Src[CurrentSrc].Data.Literal.Size = DestSize;
}
Bytes = 0;
DecodeInst->Src[CurrentSrc].TypeLiteral.Type = DecodedOperand::TYPE_LITERAL;
DecodeInst->Src[CurrentSrc].TypeLiteral.Literal = Literal;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LogMan::Throw::A(Bytes == 0, "Inst at 0x%lx: 0x%04x '%s' Had an instruction of size %d with %d remaining", DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
LOGMAN_THROW_A(Bytes == 0, "Inst at 0x%lx: 0x%04x '%s' Had an instruction of size %d with %d remaining", DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
}
@@ -592,7 +603,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return false;
}
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
LOGMAN_THROW_A(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
@@ -648,7 +659,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
3,
};
uint8_t Field = RegToField[ModRM.reg];
LogMan::Throw::A(Field != 255, "Invalid field selected!");
LOGMAN_THROW_A(Field != 255, "Invalid field selected!");
LocalOp = (Field << 3) | ModRM.rm;
return NormalOp(&SecondModRMTableOps[LocalOp], LocalOp);
@@ -682,7 +693,10 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
uint8_t Byte2 = ReadByte();
pp = Byte2 & 0b11;
map_select = Byte1 & 0b11111;
LogMan::Throw::A(map_select >= 1 && map_select <= 3, "We don't understand a map_select of: %d", map_select);
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::E("We don't understand a map_select of: %d", map_select);
return false;
}
}
uint16_t VEXOp = ReadByte();
@@ -730,6 +744,8 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->PC = PC;
for(;;) {
if (InstructionSize >= MAX_INST_SIZE)
return false;
uint8_t Op = ReadByte();
switch (Op) {
case 0x0F: {// Escape Op
@@ -880,7 +896,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
LogMan::Throw::A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
LOGMAN_THROW_A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
@@ -910,6 +926,10 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
if (DecodeInst->Dest.IsGPR()) {
assert(DecodeInst->Dest.Data.GPR.GPR != 255);
}
return true;
}
@@ -919,6 +939,7 @@ void Decoder::BranchTargetInMultiblockRange() {
// If the RIP setting is conditional AND within our symbol range then it can be considered for multiblock
uint64_t TargetRIP = 0;
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
bool Conditional = true;
switch (DecodeInst->OP) {
@@ -928,24 +949,33 @@ void Decoder::BranchTargetInMultiblockRange() {
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
LogMan::Throw::A(DecodeInst->Src[0].TypeNone.Type == DecodedOperand::TYPE_LITERAL, "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].TypeLiteral.Literal;
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
break;
}
case 0xE9:
case 0xEB: // Both are unconditional JMP instructions
LogMan::Throw::A(DecodeInst->Src[0].TypeNone.Type == DecodedOperand::TYPE_LITERAL, "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].TypeLiteral.Literal;
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
Conditional = false;
break;
case 0xE8: // Call - Immediate target, We don't want to inline calls
if (ExternalBranches) {
ExternalBranches->insert(DecodeInst->PC + DecodeInst->InstSize);
}
[[fallthrough]];
case 0xC2: // RET imm
case 0xC3: // RET
case 0xE8: // Call - Immediate target, We don't want to inline calls
default:
return;
break;
}
if (GPRSize == 4) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
TargetRIP &= 0xFFFFFFFFU;
}
// If the target RIP is within the symbol ranges then we are golden
if (TargetRIP >= SymbolMinAddress && TargetRIP < SymbolMaxAddress) {
// Update our conditional branch ranges before we return
@@ -965,6 +995,10 @@ void Decoder::BranchTargetInMultiblockRange() {
BlocksToDecode.find(TargetRIP) == BlocksToDecode.end()) {
BlocksToDecode.emplace(TargetRIP);
}
} else {
if (ExternalBranches) {
ExternalBranches->insert(TargetRIP);
}
}
}
@@ -988,10 +1022,13 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
// If we don't have symbols available then we become a bit optimistic about multiblock ranges
if (!SymbolAvailable) {
// If we don't have a symbol available then assume all branches are valid for multiblock
SymbolMaxAddress = ~0ULL;
SymbolMaxAddress = SectionMaxAddress;
SymbolMinAddress = EntryPoint;
}
DecodedMinAddress = EntryPoint;
DecodedMaxAddress = EntryPoint;
// Entry is a jump target
BlocksToDecode.emplace(PC);
@@ -1015,11 +1052,20 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
if (ErrorDuringDecoding) {
LogMan::Msg::D("Couldn't Decode something at 0x%lx, Started at 0x%lx", PC + PCOffset, PC);
LogMan::Throw::A(EntryPoint != (RIPToDecode + PCOffset), "Trying to execute invalid code");
if (Blocks.size() == 1) {
return false;
}
LOGMAN_THROW_A(Blocks.size() != 1, "Decode Error in entry block");
CurrentBlockDecoding.HasInvalidInstruction = true;
if (ErrorDuringDecoding && Blocks.size() != 1) {
ErrorDuringDecoding = false;
}
break;
}
DecodedMinAddress = std::min(DecodedMinAddress, RIPToDecode + PCOffset);
DecodedMaxAddress = std::max(DecodedMaxAddress, RIPToDecode + PCOffset + DecodeInst->InstSize);
++TotalInstructions;
++BlockNumberOfInstructions;
++DecodedSize;
+9 -2
View File
@@ -26,10 +26,15 @@ public:
Decoder(FEXCore::Context::Context *ctx);
bool DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
std::vector<DecodedBlocks> const *GetDecodedBlocks() {
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
}
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
private:
FEXCore::Context::Context *CTX;
@@ -38,7 +43,7 @@ private:
void BranchTargetInMultiblockRange();
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset);
uint8_t PeekByte(uint8_t Offset) const;
uint64_t ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) { InstructionSize += Size; }
bool NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op);
@@ -62,10 +67,12 @@ private:
uint64_t MaxCondBranchBackwards {~0ULL};
uint64_t SymbolMaxAddress {};
uint64_t SymbolMinAddress {~0ULL};
uint64_t SectionMaxAddress {~0ULL};
std::vector<DecodedBlocks> Blocks;
std::set<uint64_t> BlocksToDecode;
std::set<uint64_t> HasBlocks;
std::set<uint64_t> *ExternalBranches {nullptr};
// ModRM rm decoding
using DecodeModRMPtr = void (FEXCore::Frontend::Decoder::*)(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
+96 -100
View File
@@ -1,3 +1,10 @@
/*
$info$
tags: glue|gdbserver
desc: Provides a gdb interface to the guest state
$end_info$
*/
#include <cstdlib>
#include <cstdio>
#include <iomanip>
@@ -8,15 +15,19 @@
#include <optional>
#include "Common/NetStream.h"
#include "Common/SoftFloat.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <sys/types.h>
#include <sys/socket.h>
#include <netdb.h>
#include <string.h>
#include <cstring>
#include <fcntl.h>
#include <unistd.h>
#include <fmt/format.h>
#include <fstream>
#include <netdb.h>
#include <sys/socket.h>
#include <sys/types.h>
#include <unistd.h>
#include "GdbServer.h"
#include <FEXCore/Core/CodeLoader.h>
@@ -27,12 +38,12 @@ namespace FEXCore
void GdbServer::Break(int signal) {
std::lock_guard lk(sendMutex);
if (!CommsStream) {
return;
}
std::ostringstream ss;
ss << "S" << std::setfill('0') << std::setw(2) << std::hex << signal;
if (CommsStream)
SendPacket(*CommsStream, ss.str());
const auto str = fmt::format("S{:02x}", signal);
SendPacket(*CommsStream, str);
}
GdbServer::GdbServer(FEXCore::Context::Context *ctx) : CTX(ctx) {
@@ -58,7 +69,7 @@ GdbServer::GdbServer(FEXCore::Context::Context *ctx) : CTX(ctx) {
StartThread();
}
static int calculateChecksum(std::string &packet) {
static int calculateChecksum(const std::string &packet) {
unsigned char checksum = 0;
for (const char &c : packet) {
checksum += c;
@@ -92,11 +103,9 @@ static std::string encodeHex(unsigned char *data, size_t length) {
}
static std::string getThreadName(uint32_t ThreadID) {
std::fstream fs;
std::ostringstream ThreadFile;
ThreadFile << "/proc/" << getpid() << "/task/" << ThreadID << "/comm";
const auto ThreadFile = fmt::format("/proc/{}/task/{}/comm", getpid(), ThreadID);
std::fstream fs(ThreadFile, std::fstream::in | std::fstream::binary);
fs.open(ThreadFile.str(), std::fstream::in | std::fstream::binary);
if (fs.is_open()) {
std::string ThreadName;
fs >> ThreadName;
@@ -128,7 +137,7 @@ std::string GdbServer::ReadPacket(std::iostream &stream) {
switch(c) {
case '$': // start of packet
if (packet.size() != 0)
LogMan::Msg::E("Dropping unexpected data: \"%s\"", packet.c_str());
LogMan::Msg::EFmt("Dropping unexpected data: \"{}\"", packet);
// clear any existing data, must have been a mistake.
packet = std::string();
@@ -149,7 +158,7 @@ std::string GdbServer::ReadPacket(std::iostream &stream) {
if (calculateChecksum(packet) == expected_checksum) {
return packet;
} else {
LogMan::Msg::E("Received Invalid Packet: $%s#%02x %c%c", packet.c_str(), expected_checksum);
LogMan::Msg::EFmt("Received Invalid Packet: ${}#{:02x}", packet, expected_checksum);
}
break;
}
@@ -162,10 +171,10 @@ std::string GdbServer::ReadPacket(std::iostream &stream) {
return "";
}
static std::string escapePacket(std::string packet) {
static std::string escapePacket(const std::string& packet) {
std::ostringstream ss;
for(auto &c : packet) {
for(const auto &c : packet) {
switch (c) {
case '$':
case '#':
@@ -184,13 +193,11 @@ static std::string escapePacket(std::string packet) {
return ss.str();
}
void GdbServer::SendPacket(std::ostream &stream, std::string packet) {
auto escaped = escapePacket(packet);
std::ostringstream ss;
void GdbServer::SendPacket(std::ostream &stream, const std::string& packet) {
const auto escaped = escapePacket(packet);
const auto str = fmt::format("${}#{:02x}", escaped, calculateChecksum(escaped));
ss << '$' << escaped << '#';
ss << std::setfill('0') << std::setw(2) << std::hex << (int)calculateChecksum(escaped);
stream << ss.str() << std::flush;
stream << str << std::flush;
}
void GdbServer::SendACK(std::ostream &stream, bool NACK) {
@@ -211,7 +218,7 @@ void GdbServer::SendACK(std::ostream &stream, bool NACK) {
}
}
struct __attribute__((packed)) GDBContextDefinition {
struct FEX_PACKED GDBContextDefinition {
uint64_t gregs[16];
uint64_t rip;
uint32_t eflags;
@@ -232,17 +239,17 @@ std::string GdbServer::readRegs() {
bool Found = false;
for (auto &Thread : *Threads) {
if (Thread->State.ThreadManager.GetTID() != CurrentDebuggingThread) {
if (Thread->ThreadManager.GetTID() != CurrentDebuggingThread) {
continue;
}
state = Thread->State.State;
memcpy(&state, Thread->CurrentFrame, sizeof(state));
Found = true;
break;
}
if (!Found) {
// If set to an invalid thread then just get the parent thread ID
state = CTX->GetCPUState();
memcpy(&state, CTX->ParentThread->CurrentFrame, sizeof(state));
}
// Encode the GDB context definition
@@ -272,7 +279,7 @@ std::string GdbServer::readRegs() {
return encodeHex((unsigned char *)&GDB, sizeof(GDBContextDefinition));
}
GdbServer::HandledPacketType GdbServer::readReg(std::string& packet) {
GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
size_t addr;
auto ss = std::istringstream(packet);
ss.get(); // Drop first letter
@@ -284,17 +291,17 @@ GdbServer::HandledPacketType GdbServer::readReg(std::string& packet) {
bool Found = false;
for (auto &Thread : *Threads) {
if (Thread->State.ThreadManager.GetTID() != CurrentDebuggingThread) {
if (Thread->ThreadManager.GetTID() != CurrentDebuggingThread) {
continue;
}
state = Thread->State.State;
memcpy(&state, Thread->CurrentFrame, sizeof(state));
Found = true;
break;
}
if (!Found) {
// If set to an invalid thread then just get the parent thread ID
state = CTX->GetCPUState();
memcpy(&state, CTX->ParentThread->CurrentFrame, sizeof(state));
}
@@ -350,7 +357,7 @@ GdbServer::HandledPacketType GdbServer::readReg(std::string& packet) {
return {encodeHex((unsigned char *)(&Empty), sizeof(uint32_t)), HandledPacketType::TYPE_ACK};
}
LogMan::Msg::E("Unknown GDB register 0x%lx", addr);
LogMan::Msg::EFmt("Unknown GDB register 0x{:x}", addr);
return {"E00", HandledPacketType::TYPE_ACK};
}
@@ -455,7 +462,7 @@ std::string buildTargetXML() {
return xml.str();
}
GdbServer::HandledPacketType GdbServer::handleXfer(std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
std::string object;
std::string rw;
std::string annex;
@@ -525,7 +532,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(std::string &packet) {
ss << "<threads>\n";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << "\t<thread id=\"" << std::hex << Thread->State.ThreadManager.GetTID() << "\" core=\"" << i << "\" name=\"" << getThreadName(Thread->State.ThreadManager.GetTID()) << "\">\n";
ss << "\t<thread id=\"" << std::hex << Thread->ThreadManager.GetTID() << "\" core=\"" << i << "\" name=\"" << getThreadName(Thread->ThreadManager.GetTID()) << "\">\n";
ss << "\t</thread>\n";
}
@@ -541,10 +548,9 @@ GdbServer::HandledPacketType GdbServer::handleXfer(std::string &packet) {
static size_t CheckMemMapping(uint64_t Address, size_t Size) {
uint64_t AddressEnd = Address + Size;
std::fstream fs;
fs.open("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
while (std::getline(fs, Line)) {
if (fs.eof()) break;
uint64_t Begin, End;
@@ -561,32 +567,29 @@ static size_t CheckMemMapping(uint64_t Address, size_t Size) {
}
}
fs.close();
return 0;
}
GdbServer::HandledPacketType GdbServer::handleProgramOffsets() {
std::fstream fs;
fs.open("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
std::string const &RuntimeExecutable = Filename();
while (std::getline(fs, Line)) {
uint64_t Begin, End;
char Filename[255];
if (sscanf(Line.c_str(), "%lx-%lx %*c%*c%*c%*c %*x %*x:%*x %*d%s", &Begin, &End, Filename) == 3) {
if (RuntimeExecutable == Filename) {
std::ostringstream ss;
ss << "Text=" << std::hex << Begin << ";Data=" << std::hex << Begin << ";Bss=" << std::hex << Begin;
ss << std::flush;
return {ss.str(), HandledPacketType::TYPE_ACK};
auto str = fmt::format("Text={:x};Data={:x};Bss={:x}", Begin, Begin, Begin);
return {std::move(str), HandledPacketType::TYPE_ACK};
}
}
}
fs.close();
return {"Text=0;Data=0;Bss=0", HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::handleMemory(std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleMemory(const std::string &packet) {
bool write;
size_t addr;
size_t length;
@@ -627,8 +630,8 @@ GdbServer::HandledPacketType GdbServer::handleMemory(std::string &packet) {
}
GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
const auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
if (match("qSupported")) {
return {"PacketSize=5000;xmlRegisters=i386;qXfer:exec-file:read+;qXfer:features:read+;", HandledPacketType::TYPE_ACK};
@@ -653,7 +656,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
ss << "m";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << std::hex << Thread->State.ThreadManager.TID << ",";
ss << std::hex << Thread->ThreadManager.TID << ",";
}
return {ss.str(), HandledPacketType::TYPE_ACK};
}
@@ -672,7 +675,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
if (match("qC")) {
// Returns the current Thread ID
std::ostringstream ss;
ss << "m" << std::hex << CTX->ParentThread->State.ThreadManager.TID;
ss << "m" << std::hex << CTX->ParentThread->ThreadManager.TID;
return {ss.str(), HandledPacketType::TYPE_ACK};
}
if (match("QStartNoAckMode")) {
@@ -686,8 +689,8 @@ GdbServer::HandledPacketType GdbServer::handleQuery(std::string &packet) {
return {"", HandledPacketType::TYPE_UNKNOWN};
}
GdbServer::HandledPacketType GdbServer::handleV(std::string& packet) {
auto match = [&](std::string str) -> std::optional<std::istringstream> {
GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
const auto match = [&](const std::string& str) -> std::optional<std::istringstream> {
if (packet.rfind(str, 0) == 0) {
auto ss = std::istringstream(packet);
ss.seekg(str.size());
@@ -696,18 +699,11 @@ GdbServer::HandledPacketType GdbServer::handleV(std::string& packet) {
return std::nullopt;
};
auto F = [](int result) {
std::ostringstream ss;
ss << "F" << std::hex << result;
return ss.str(); };
auto F_error = [&]() {
std::ostringstream ss;
ss << "F-1," << std::hex << errno;
return ss.str(); };
auto F_data = [&](int result, std::string data) {
std::ostringstream ss;
ss << "F" << std::hex << result << ";" << data;
return ss.str(); };
const auto F = [](int result) { return fmt::format("F{:x}", result); };
const auto F_error = [] { return fmt::format("F-1,{:x}", errno); };
const auto F_data = [](int result, const std::string& data) {
return fmt::format("F{:x};{}", result, data);
};
std::optional<std::istringstream> ss;
if((ss = match("vFile:open:"))) {
@@ -729,11 +725,11 @@ GdbServer::HandledPacketType GdbServer::handleV(std::string& packet) {
return {F(pid == 0 ? 0 : -1), HandledPacketType::TYPE_ACK}; // Only support the common filesystem
}
if((ss = match("vFile:close:"))) {
int fd;
*ss >> std::hex >> fd;
close(fd);
return {F(0), HandledPacketType::TYPE_ACK};
}
int fd;
*ss >> std::hex >> fd;
close(fd);
return {F(0), HandledPacketType::TYPE_ACK};
}
if((ss = match("vFile:pread:"))) {
int fd, count, offset;
@@ -770,7 +766,7 @@ GdbServer::HandledPacketType GdbServer::handleV(std::string& packet) {
}
if (ss->fail()) {
return {"E00", HandledPacketType::TYPE_ACK};
return {"E00", HandledPacketType::TYPE_ACK};
}
switch (action) {
@@ -780,27 +776,25 @@ GdbServer::HandledPacketType GdbServer::handleV(std::string& packet) {
}
case 's': {
CTX->Step();
SendPacketPair({"OK", HandledPacketType::TYPE_ACK});
std::ostringstream ss;
ss << "T05thread:" << std::setfill('0') << std::setw(2) << std::hex << getpid() << ";core:2c;";
SendPacketPair({ss.str(), HandledPacketType::TYPE_ACK});
SendPacketPair({"OK", HandledPacketType::TYPE_ACK});
auto str = fmt::format("T05thread:{:02x};core:2c;", getpid());
SendPacketPair({std::move(str), HandledPacketType::TYPE_ACK});
return {"OK", HandledPacketType::TYPE_ACK};
}
case 't':
// This thread isn't part of the thread pool
CTX->Stop(false /* Ignore current thread */);
return {"OK", HandledPacketType::TYPE_ACK};
return {"OK", HandledPacketType::TYPE_ACK};
default:
return {"E00", HandledPacketType::TYPE_ACK};
return {"E00", HandledPacketType::TYPE_ACK};
}
}
return {"", HandledPacketType::TYPE_ACK};
return {"", HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::handleThreadOp(std::string &packet) {
auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
GdbServer::HandledPacketType GdbServer::handleThreadOp(const std::string &packet) {
const auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
if (match("Hc")) {
// Sets thread to this ID for stepping
@@ -816,7 +810,7 @@ GdbServer::HandledPacketType GdbServer::handleThreadOp(std::string &packet) {
if (match("Hg")) {
// Sets thread for "other" operations
auto ss = std::istringstream(packet);
ss.seekg(std::string("Hg").size());
ss.seekg(std::string_view("Hg").size());
ss >> std::hex >> CurrentDebuggingThread;
// This must return quick otherwise IDA complains
@@ -827,7 +821,7 @@ GdbServer::HandledPacketType GdbServer::handleThreadOp(std::string &packet) {
return {"", HandledPacketType::TYPE_UNKNOWN};
}
GdbServer::HandledPacketType GdbServer::handleBreakpoint(std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleBreakpoint(const std::string &packet) {
auto ss = std::istringstream(packet);
bool Set{};
@@ -843,17 +837,15 @@ GdbServer::HandledPacketType GdbServer::handleBreakpoint(std::string &packet) {
return {"OK", HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::ProcessPacket(std::string &packet) {
GdbServer::HandledPacketType GdbServer::ProcessPacket(const std::string &packet) {
switch (packet[0]) {
case '?': {
// Indicates the reason that the thread has stopped
// Behaviour changes if the target is in non-stop mode
// Binja doesn't support S response here
//return {"S00", HandledPacketType::TYPE_ACK};
std::ostringstream ss;
ss << "T00thread:" << std::setfill('0') << std::setw(2) << std::hex << getpid() << ";core:2c;";
return {ss.str(), HandledPacketType::TYPE_ACK};
auto str = fmt::format("T00thread:{:02x};core:2c;", getpid());
return {std::move(str), HandledPacketType::TYPE_ACK};
}
case 'g':
return {readRegs(), HandledPacketType::TYPE_ACK};
@@ -883,14 +875,14 @@ GdbServer::HandledPacketType GdbServer::ProcessPacket(std::string &packet) {
}
}
void GdbServer::SendPacketPair(HandledPacketType response) {
void GdbServer::SendPacketPair(const HandledPacketType& response) {
std::lock_guard lk(sendMutex);
if (response.TypeResponse == HandledPacketType::TYPE_ACK ||
response.TypeResponse == HandledPacketType::TYPE_ONLYACK) {
SendACK(*CommsStream, false);
}
else if (response.TypeResponse == HandledPacketType::TYPE_NACK ||
response.TypeResponse == HandledPacketType::TYPE_ONLYNACK) {
response.TypeResponse == HandledPacketType::TYPE_ONLYNACK) {
SendACK(*CommsStream, true);
}
@@ -898,8 +890,8 @@ void GdbServer::SendPacketPair(HandledPacketType response) {
SendPacket(*CommsStream, "");
}
else if (response.TypeResponse != HandledPacketType::TYPE_ONLYNACK &&
response.TypeResponse != HandledPacketType::TYPE_ONLYACK &&
response.TypeResponse != HandledPacketType::TYPE_NONE) {
response.TypeResponse != HandledPacketType::TYPE_ONLYACK &&
response.TypeResponse != HandledPacketType::TYPE_NONE) {
SendPacket(*CommsStream, response.Response);
}
}
@@ -920,7 +912,7 @@ void GdbServer::GdbServerLoop() {
response = ProcessPacket(packet);
SendPacketPair(response);
if (response.TypeResponse == HandledPacketType::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown packet %s", packet.c_str());
LogMan::Msg::DFmt("Unknown packet {}", packet);
}
break;
}
@@ -936,13 +928,12 @@ void GdbServer::GdbServerLoop() {
break;
case '\x03': { // ASCII EOT
CTX->Pause();
std::ostringstream ss;
ss << "T02thread:" << std::setfill('0') << std::setw(2) << std::hex << getpid() << ";core:2c;";
SendPacketPair({ss.str(), HandledPacketType::TYPE_ACK});
auto str = fmt::format("T02thread:{:02x};core:2c;", getpid());
SendPacketPair({std::move(str), HandledPacketType::TYPE_ACK});
break;
}
default:
LogMan::Msg::D("GdbServer: Unexpected byte %c (%02x)", c, c);
LogMan::Msg::DFmt("GdbServer: Unexpected byte {} ({:02x})", static_cast<char>(c), c);
}
}
@@ -952,9 +943,14 @@ void GdbServer::GdbServerLoop() {
}
}
}
static void* ThreadHandler(void *Arg) {
FEXCore::GdbServer *This = reinterpret_cast<FEXCore::GdbServer*>(Arg);
This->GdbServerLoop();
return nullptr;
}
void GdbServer::StartThread() {
gdbServerThread = std::thread(&GdbServer::GdbServerLoop, this);
gdbServerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
}
std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
@@ -991,7 +987,7 @@ std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
// Block until a connection arrives
LogMan::Msg::I("GdbServer, waiting for connection on localhost:8086");
LogMan::Msg::IFmt("GdbServer, waiting for connection on localhost:8086");
listen(sockfd, 1);
new_fd = accept(sockfd, (struct sockaddr *)&their_addr, &addr_size);
+25 -14
View File
@@ -1,25 +1,36 @@
/*
$info$
tags: glue|gdbserver
$end_info$
*/
#pragma once
#include <mutex>
#include <thread>
#include "Interface/Context/Context.h"
#include "Common/NetStream.h"
#include <FEXCore/Utils/Threads.h>
#include <mutex>
namespace FEXCore {
class GdbServer {
public:
GdbServer(FEXCore::Context::Context *ctx);
// Public for threading
void GdbServerLoop();
private:
void Break(int signal);
std::unique_ptr<std::iostream> OpenSocket();
void StartThread();
void GdbServerLoop();
std::string ReadPacket(std::iostream &stream);
void SendPacket(std::ostream &stream, std::string packet);
void SendPacket(std::ostream &stream, const std::string& packet);
void SendACK(std::ostream &stream, bool NACK);
@@ -36,28 +47,28 @@ private:
ResponseType TypeResponse{};
};
void SendPacketPair(HandledPacketType packetPair);
HandledPacketType ProcessPacket(std::string &packet);
HandledPacketType handleQuery(std::string &packet);
HandledPacketType handleXfer(std::string &packet);
HandledPacketType handleMemory(std::string &packet);
HandledPacketType handleV(std::string& packet);
HandledPacketType handleThreadOp(std::string &packet);
HandledPacketType handleBreakpoint(std::string &packet);
void SendPacketPair(const HandledPacketType& packetPair);
HandledPacketType ProcessPacket(const std::string &packet);
HandledPacketType handleQuery(const std::string &packet);
HandledPacketType handleXfer(const std::string &packet);
HandledPacketType handleMemory(const std::string &packet);
HandledPacketType handleV(const std::string& packet);
HandledPacketType handleThreadOp(const std::string &packet);
HandledPacketType handleBreakpoint(const std::string &packet);
HandledPacketType handleProgramOffsets();
std::string readRegs();
HandledPacketType readReg(std::string& packet);
HandledPacketType readReg(const std::string& packet);
FEXCore::Context::Context *CTX;
std::thread gdbServerThread;
std::unique_ptr<FEXCore::Threads::Thread> gdbServerThread;
std::unique_ptr<std::iostream> CommsStream;
std::mutex sendMutex;
bool SettingNoAckMode{false};
bool NoAckMode{false};
std::string ThreadString{};
uint32_t CurrentDebuggingThread{};
FEXCore::Config::Value<std::string> Filename{FEXCore::Config::CONFIG_APP_FILENAME, ""};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
};
}
@@ -1,563 +0,0 @@
#include "Common/MathUtils.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
namespace FEXCore::CPU {
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->State.RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
using namespace vixl;
using namespace vixl::aarch64;
#define STATE x28
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
class DispatchGenerator : public vixl::aarch64::Assembler {
public:
DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
uint64_t ThreadStopHandlerAddress;
uint64_t AbsoluteLoopTopAddress;
uint64_t ThreadPauseHandlerAddress;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
private:
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
std::stack<uint64_t> SignalFrames;
};
void DispatchGenerator::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
{x25, x26},
{x27, x28},
{x29, x30},
}};
for (auto &RegPair : CalleeSaved) {
stp(RegPair.first, RegPair.second, PairOffset);
}
// Additionally we need to store the lower 64bits of v8-v15
// Here's a fun thing, we can use two ST4 instructions to store everything
// We just need a single sub to sp before that
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v8, v9, v10, v11},
{v12, v13, v14, v15},
}};
uint32_t VectorSaveSize = sizeof(uint64_t) * 8;
sub(sp, sp, VectorSaveSize);
// SP supporting move
// We just saved x19 so it is safe
add(x19, sp, 0);
MemOperand QuadOffset(x19, 32, PostIndex);
for (auto &RegQuad : FPRs) {
st4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
}
void DispatchGenerator::PopCalleeSavedRegisters() {
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v12, v13, v14, v15},
{v8, v9, v10, v11},
}};
MemOperand QuadOffset(sp, 32, PostIndex);
for (auto &RegQuad : FPRs) {
ld4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
0,
QuadOffset);
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
{x23, x24},
{x21, x22},
{x19, x20},
}};
for (auto &RegPair : CalleeSaved) {
ldp(RegPair.first, RegPair.second, PairOffset);
}
}
void DispatchGenerator::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
}
}
}
DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: vixl::aarch64::Assembler(MAX_DISPATCHER_CODE_SIZE, vixl::aarch64::PositionDependentCode)
, CTX {ctx}
, State {Thread} {
SetAllowAssembler(true);
auto Buffer = GetBuffer();
DispatchPtr = Buffer->GetOffsetAddress<CPUBackend::AsmDispatch>(GetCursorOffset());
// while (!Thread->State.RunningEvents.ShouldStop.load()) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Push all the register we need to save
PushCalleeSavedRegisters();
// Push our memory base to the correct register
// Move our thread pointer to the correct register
// This is passed in to parameter 0 (x0)
mov(STATE, x0);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)));
Label Exit;
Label LoopTop;
Label NoBlock;
Label ThreadPauseHandler;
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, State.rip)));
auto RipReg = x2;
// Mask the address by the virtual address size so we can check for aliases
LoadConstant(x3, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, x3);
{
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
LoadConstant(x0, Thread->LookupCache->GetPagePointer());
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
// Load the pointer from the offset
ldr(x0, MemOperand(x0, x1, Shift::LSL, 3));
// If page pointer is zero then we have no block
cbz(x0, &NoBlock);
// Steal the page offset
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x1, &NoBlock);
// If we've made it here then we have a real compiled block
{
mov(x0, STATE);
blr(x1);
}
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, CTX)));
ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
// If the value == 0 then branch to the top
cbz(x0, &LoopTop);
// Else we need to pause now
b(&ThreadPauseHandler);
}
else {
// Unconditionally loop to the top
// We will only stop on error when compiling a block or signal
b(&LoopTop);
}
}
{
bind(&Exit);
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
PopCalleeSavedRegisters();
// Return from the function
// LR is set to the correct return location now
ret();
}
// Need to create the block
{
bind(&NoBlock);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, CTX)));
mov(x1, STATE);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
LoadConstant(x3, Ptr.Data);
// X2 contains our guest RIP
blr(x3); // { CTX, ThreadState, RIP}
b(&LoopTop);
}
{
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
// Call our sleep handler
LoadConstant(x0, reinterpret_cast<uintptr_t>(CTX));
mov(x1, STATE);
LoadConstant(x2, reinterpret_cast<uint64_t>(SleepThread));
blr(x2);
// XXX: Unsupported atm
//PauseReturnInstruction = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
//// Fault to start running again
//hlt(0);
}
{
CallbackPtr = Buffer->GetOffsetAddress<CPUBackend::JITCallback>(GetCursorOffset());
// We expect the thunk to have previously pushed the registers it was using
PushCalleeSavedRegisters();
// First thing we need to move the thread state pointer back in to our register
mov(STATE, x0);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
sub(x2, x2, 16);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
// Store RIP to the context state
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.rip)));
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
FinalizeCode();
uint64_t CodeEnd = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), CodeEnd - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
}
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
constexpr uint32_t FPR_MAGIC = 0x46508001U;
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
struct ContextBackup {
// Host State
uint64_t GPRs[31];
uint64_t PrevSP;
uint64_t PrevPC;
uint64_t PState;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
};
void DispatchGenerator::StoreThreadState(int Signal, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = _mcontext->sp;
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ContextBackup);
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
memcpy(&Context->GPRs[0], &_mcontext->regs[0], 31 * sizeof(uint64_t));
Context->PrevSP = _mcontext->sp;
Context->PrevPC = _mcontext->pc;
Context->PState = _mcontext->pstate;
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
Context->FPSR = HostState->FPSR;
Context->FPCR = HostState->FPCR;
memcpy(&Context->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, &State->State, sizeof(FEXCore::Core::CPUState));
// Set the new SP
_mcontext->sp = NewSP;
}
void DispatchGenerator::RestoreThreadState(void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint64_t OldSP = _mcontext->sp;
uintptr_t NewSP = OldSP;
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
// First thing, reset the guest state
memcpy(&State->State, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LogMan::Throw::A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Context->FPRs[0], 32 * sizeof(__uint128_t));
Context->FPCR = HostState->FPCR;
Context->FPSR = HostState->FPSR;
// Restore GPRs and other state
_mcontext->pstate = Context->PState;
_mcontext->pc = Context->PrevPC;
_mcontext->sp = Context->PrevSP;
memcpy(&_mcontext->regs[0], &Context->GPRs[0], 31 * sizeof(uint64_t));
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool DispatchGenerator::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->pc = AbsoluteLoopTopAddress;
// Set x28 (which is our state register) to point to our guest thread data
_mcontext->regs[28 /* STATE */] = reinterpret_cast<uint64_t>(State);
State->State.State.gregs[X86State::REG_RDI] = Signal;
uint64_t OldGuestSP = State->State.State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
if (GuestAction->sa_flags & SA_SIGINFO) {
// XXX: siginfo_t(RSI), ucontext (RDX)
State->State.State.gregs[X86State::REG_RSI] = 0;
State->State.State.gregs[X86State::REG_RDX] = 0;
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
State->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
return true;
}
bool DispatchGenerator::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = State->SignalReason.load();
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->pc = ThreadPauseHandlerAddress;
// Set our state register to point to our guest thread data
_mcontext->regs[28 /* STATE */] = reinterpret_cast<uint64_t>(State);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the JIT and get out safely
_mcontext->sp = State->State.ReturningStackLocation;
// Set the new PC
_mcontext->pc = ThreadStopHandlerAddress;
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
// TODO: Implement this. It is missing from the dispatcher
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = nullptr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
DispatchGenerator *Gen = Generator;
return Gen->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
}
bool InterpreterCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
DispatchGenerator *Gen = Generator;
return Gen->HandleSignalPause(Signal, info, ucontext);
}
void InterpreterCore::DeleteAsmDispatch() {
delete Generator;
}
}
@@ -2,13 +2,15 @@
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::CPU {
class DispatchGenerator;
class X86DispatchGenerator;
class Arm64DispatchGenerator;
#define DESTMAP_AS_MAP 0
#if DESTMAP_AS_MAP
@@ -20,16 +22,14 @@ using DestMapType = std::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
~InterpreterCore() override;
std::string GetName() override { return "Interpreter"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
bool NeedsOpDispatch() override { return true; }
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
void DeleteAsmDispatch();
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
@@ -38,8 +38,6 @@ private:
FEXCore::Core::InternalThreadState *State;
uint32_t AllocateTmpSpace(size_t Size);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
template<typename Res>
Res GetDest(void* SSAData, IR::OrderedNodeWrapper Op);
@@ -47,7 +45,7 @@ private:
template<typename Res>
Res GetSrc(void* SSAData, IR::OrderedNodeWrapper Src);
DispatchGenerator *Generator{};
std::unique_ptr<Dispatcher> Dispatcher{};
};
}
@@ -2,9 +2,8 @@
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#ifdef _M_ARM_64
#include "Interface/Core/ArchHelpers/Arm64.h"
#endif
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/DebugData.h"
#include "Interface/Core/InternalThreadState.h"
@@ -22,57 +21,64 @@
#include <cmath>
#include <limits>
#include <vector>
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
#include "InterpreterOps.h"
namespace FEXCore::CPU {
static void InterpreterExecution(FEXCore::Core::InternalThreadState *Thread) {
auto IR = Thread->IRLists.find(Thread->State.State.rip)->second.get();
FEXCore::Core::DebugData *DebugData = nullptr;
// DebugData is only used in debug builds
#ifndef NDEBUG
DebugData = Thread->DebugData.find(Thread->State.State.rip)->second.get();
#endif
static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
InterpreterOps::InterpretIR(Thread, IR, DebugData);
auto LocalEntry = Thread->LocalIRCache.find(Thread->CurrentFrame->State.rip);
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
bool InterpreterCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
#ifdef _M_ARM_64
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint32_t *PC = (uint32_t*)_mcontext->pc;
uint32_t Instr = PC[0];
if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
constexpr bool is_arm64 = true;
#else
constexpr bool is_arm64 = false;
#endif
if constexpr (is_arm64) {
uint32_t *PC = reinterpret_cast<uint32_t*>(ArchHelpers::Context::GetPc(ucontext));
uint32_t Instr = PC[0];
if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x: PC: %p Instruction: 0x%08x\n", Op, PC, PC[0]);
return false;
}
}
}
return false;
}
@@ -86,7 +92,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->HandleSignalPause(Signal, info, ucontext);
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
});
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
@@ -96,7 +102,7 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal < SignalDelegator::MAX_SIGNALS; ++Signal) {
@@ -105,18 +111,12 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
}
}
InterpreterCore::~InterpreterCore() {
DeleteAsmDispatch();
}
void *InterpreterCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
return reinterpret_cast<void*>(InterpreterExecution);
}
FEXCore::CPU::CPUBackend *CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return new InterpreterCore(ctx, Thread, CompileThread);
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<InterpreterCore>(ctx, Thread, CompileThread);
}
}
@@ -1,5 +1,7 @@
#pragma once
#include <memory>
namespace FEXCore::Context {
struct Context;
}
@@ -11,6 +13,6 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
FEXCore::CPU::CPUBackend *CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
}
File diff suppressed because it is too large. Load diff
@@ -1,19 +1,42 @@
namespace FEXCore::Core {
struct InternalThreadState;
struct InternalThreadState;
}
namespace FEXCore::IR {
template<bool copy>
class IRListView;
}
namespace FEXCore::Core{
struct DebugData;
struct DebugData;
}
namespace FEXCore::CPU {
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEXCore::IR::IRListView<true> *CurrentIR, FEXCore::Core::DebugData *DebugData);
};
enum FallbackABI {
FABI_UNKNOWN,
FABI_VOID_U16,
FABI_F80_F32,
FABI_F80_F64,
FABI_F80_I16,
FABI_F80_I32,
FABI_F32_F80,
FABI_F64_F80,
FABI_I16_F80,
FABI_I32_F80,
FABI_I64_F80,
FABI_I64_F80_F80,
FABI_F80_F80,
FABI_F80_F80_F80,
};
struct FallbackInfo {
FallbackABI ABI;
void *fn;
};
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
};
};
@@ -1,470 +0,0 @@
#include "Common/MathUtils.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
class DispatchGenerator : public Xbyak::CodeGenerator {
public:
DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
uint64_t ThreadStopHandlerAddress;
uint64_t AbsoluteLoopTopAddress;
uint64_t ThreadPauseHandlerAddress;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
private:
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
};
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
// Go to sleep
Thread->StartRunning.Wait();
Thread->State.RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
}
DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE)
, CTX {ctx}
, State {Thread} {
using namespace Xbyak;
using namespace Xbyak::util;
DispatchPtr = getCurr<CPUBackend::AsmDispatch>();
// while (!Thread->State.RunningEvents.ShouldStop.load()) {
// Ptr = FindBlock(RIP)
// if (!Ptr)
// Ptr = CTX->CompileBlock(RIP);
//
// if (Ptr)
// Ptr();
// else
// {
// Ptr = FallbackCore->CompileBlock()
// if (Ptr)
// Ptr()
// else {
// ShouldStop = true;
// }
// }
// }
// Bunch of exit state stuff
// x86-64 ABI has the stack aligned when /call/ happens
// Which means the destination has a misaligned stack at that point
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
mov(STATE, rdi);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
mov(qword [rdi + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)], rsp);
Label LoopTop;
Label NoBlock;
Label ExitBlock;
Label ThreadPauseHandler;
L(LoopTop);
AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
mov(r13, Thread->LookupCache->GetPagePointer());
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
// Load page pointer
mov(rdi, qword [r13 + rax * 8]);
cmp(rdi, 0);
je(NoBlock);
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
cmp(rcx, rdx);
jne(NoBlock);
// Load the block pointer
mov(rax, qword [rdi + rax]);
cmp(rax, 0);
je(NoBlock);
// Real block if we made it here
mov(rdi, STATE);
call(rax);
if (CTX->GetGdbServerStatus()) {
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(CTX->Config.RunningMode) == 4, "This is expected to be size of 4");
mov(rax, qword [STATE + (offsetof(FEXCore::Core::InternalThreadState, CTX))]);
// If the value == 0 then branch to the top
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(LoopTop);
// Else we need to pause now
jmp(ThreadPauseHandler);
ud2();
}
else {
jmp(LoopTop);
}
}
{
L(ExitBlock);
ThreadStopHandlerAddress = getCurr<uint64_t>();
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
// Block creation
{
L(NoBlock);
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
mov(rax, Ptr.Data);
call(rax);
// rdx already contains RIP here
jmp(LoopTop);
}
{
// Pause handler
ThreadPauseHandlerAddress = getCurr<uint64_t>();
L(ThreadPauseHandler);
mov(rdi, reinterpret_cast<uintptr_t>(CTX));
mov(rsi, STATE);
mov(rax, reinterpret_cast<uint64_t>(SleepThread));
call(rax);
// XXX: Unsupported atm
// uint64_t PauseReturnInstruction = getCurr<uint64_t>();
// ud2();
}
{
CallbackPtr = getCurr<CPUBackend::JITCallback>();
push(rbx);
push(rbp);
push(r12);
push(r13);
push(r14);
push(r15);
sub(rsp, 8);
// First thing we need to move the thread state pointer back in to our register
mov(STATE, rdi);
// XXX: XMM?
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
mov(rax, CTX->X86CodeGen.CallbackReturn);
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
sub(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])], 16);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])]);
mov(qword [rbx], rax);
// Store RIP to the context state
mov(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.rip)], rsi);
// Back to the loop top now
jmp(LoopTop);
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
// rsi = rsp
mov(rsp, rsi);
// Now jump back to the thunk
// XXX: XMM?
add(rsp, 8);
pop(r15);
pop(r14);
pop(r13);
pop(r12);
pop(rbp);
pop(rbx);
ret();
}
ready();
}
struct ContextBackup {
uint64_t StoredCookie;
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[NGREG];
_libc_fpstate FPRState;
// Guest state
int Signal;
FEXCore::Core::CPUState GuestState;
};
void DispatchGenerator::StoreThreadState(int Signal, void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = _mcontext->gregs[REG_RSP];
uintptr_t NewSP = OldSP;
size_t StackOffset = sizeof(ContextBackup);
// We need to back up behind the host's red zone
// We do this on the guest side as well
NewSP -= 128;
NewSP -= StackOffset;
NewSP = AlignDown(NewSP, 16);
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
Context->StoredCookie = 0x4142434445464748ULL;
// Copy the GPRs
memcpy(&Context->GPRs[0], &_mcontext->gregs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(&Context->FPRState, _mcontext->fpregs, sizeof(_libc_fpstate));
// XXX: Save 256bit and 512bit AVX register state
// Retain the action pointer so we can see it when we return
Context->Signal = Signal;
// Save guest state
// We can't guarantee if registers are in context or host GPRs
// So we need to save everything
memcpy(&Context->GuestState, &State->State, sizeof(FEXCore::Core::CPUState));
// Set the new SP
_mcontext->gregs[REG_RSP] = NewSP;
SignalFrames.push(NewSP);
}
void DispatchGenerator::RestoreThreadState(void *ucontext) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uintptr_t NewSP = OldSP;
ContextBackup *Context = reinterpret_cast<ContextBackup*>(NewSP);
if (Context->StoredCookie != 0x4142434445464748ULL) {
LogMan::Msg::D("COOKIE WAS NOT CORRECT!\n");
exit(-1);
}
// First thing, reset the guest state
memcpy(&State->State, &Context->GuestState, sizeof(FEXCore::Core::CPUState));
// Now restore host state
// Copy the GPRs
memcpy(&_mcontext->gregs[0], &Context->GPRs[0], NGREG * sizeof(_mcontext->gregs[0]));
// Copy the FPRState
memcpy(_mcontext->fpregs, &Context->FPRState, sizeof(_libc_fpstate));
// Restore the previous signal state
// This allows recursive signals to properly handle signal masking as we are walking back up the list of signals
CTX->SignalDelegation->SetCurrentSignal(Context->Signal);
}
bool DispatchGenerator::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = AbsoluteLoopTopAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(State);
uint64_t OldGuestSP = State->State.State.gregs[X86State::REG_RSP];
uint64_t NewGuestSP = OldGuestSP;
if (!(GuestStack->ss_flags & SS_DISABLE)) {
// If our guest is already inside of the alternative stack
// Then that means we are hitting recursive signals and we need to walk back the stack correctly
uint64_t AltStackBase = reinterpret_cast<uint64_t>(GuestStack->ss_sp);
uint64_t AltStackEnd = AltStackBase + GuestStack->ss_size;
if (OldGuestSP >= AltStackBase &&
OldGuestSP <= AltStackEnd) {
// We are already in the alt stack, the rest of the code will handle adjusting this
}
else {
NewGuestSP = AltStackEnd;
}
}
// Back up past the redzone, which is 128bytes
// Don't need this offset if we aren't going to be putting siginfo in to it
NewGuestSP -= 128;
State->State.State.gregs[X86State::REG_RDI] = Signal;
if (GuestAction->sa_flags & SA_SIGINFO) {
// XXX: siginfo_t(RSI), ucontext (RDX)
State->State.State.gregs[X86State::REG_RSI] = 0;
State->State.State.gregs[X86State::REG_RDX] = 0;
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.sigaction);
}
else {
State->State.State.rip = reinterpret_cast<uint64_t>(GuestAction->sigaction_handler.handler);
}
// Set up the new SP for stack handling
NewGuestSP -= 8;
*(uint64_t*)NewGuestSP = CTX->X86CodeGen.SignalReturn;
State->State.State.gregs[X86State::REG_RSP] = NewGuestSP;
return true;
}
bool DispatchGenerator::HandleSignalPause(int Signal, void *info, void *ucontext) {
FEXCore::Core::SignalEvent SignalReason = State->SignalReason.load();
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_PAUSE) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Store our thread state so we can come back to this
StoreThreadState(Signal, ucontext);
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadPauseHandlerAddress;
// Set our state register to point to our guest thread data
_mcontext->gregs[REG_R14] = reinterpret_cast<uint64_t>(State);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_STOP) {
ucontext_t* _context = (ucontext_t*)ucontext;
mcontext_t* _mcontext = &_context->uc_mcontext;
// Our thread is stopping
// We don't care about anything at this point
// Set the stack to our starting location when we entered the core and get out safely
_mcontext->gregs[REG_RSP] = State->State.ReturningStackLocation;
// Set the new PC
_mcontext->gregs[REG_RIP] = ThreadStopHandlerAddress;
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
if (SignalReason == FEXCore::Core::SignalEvent::SIGNALEVENT_RETURN) {
RestoreThreadState(ucontext);
State->SignalReason.store(FEXCore::Core::SIGNALEVENT_NONE);
return true;
}
return false;
}
void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Generator->ReturnPtr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
DispatchGenerator *Gen = Generator;
return Gen->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
}
bool InterpreterCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
DispatchGenerator *Gen = Generator;
return Gen->HandleSignalPause(Signal, info, ucontext);
}
void InterpreterCore::DeleteAsmDispatch() {
delete Generator;
}
}
+91 -82
View File
@@ -1,3 +1,8 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -29,7 +34,7 @@ static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -41,7 +46,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LogMan::Msg::A("Unhandled Truncation size: %d", Op->Size); break;
default: LOGMAN_MSG_A("Unhandled Truncation size: %d", Op->Size); break;
}
}
@@ -51,10 +56,22 @@ DEF_OP(Constant) {
LoadConstant(Dst, Op->Constant);
}
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg<RA_64>(Node);
LoadConstant(Dst, Constant);
}
DEF_OP(InlineConstant) {
//nop
}
DEF_OP(InlineEntrypointOffset) {
//nop
}
DEF_OP(CycleCounter) {
#ifdef DEBUG_CYCLES
movz(GetReg<RA_64>(Node), 0);
@@ -78,7 +95,7 @@ DEF_OP(Add) {
case 8:
add(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Const);
break;
default: LogMan::Msg::A("Unsupported Add size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Add size: %d", OpSize);
}
} else {
switch (OpSize) {
@@ -88,7 +105,7 @@ DEF_OP(Add) {
case 8:
add(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unsupported Add size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Add size: %d", OpSize);
}
}
}
@@ -104,7 +121,7 @@ DEF_OP(Sub) {
case 8:
sub(GRS(Node), GRS(Op->Header.Args[0].ID()), Const);
break;
default: LogMan::Msg::A("Unsupported Sub size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Sub size: %d", OpSize);
}
} else {
switch (OpSize) {
@@ -114,7 +131,7 @@ DEF_OP(Sub) {
case 8:
sub(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unsupported Sub size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Sub size: %d", OpSize);
}
}
@@ -130,7 +147,7 @@ DEF_OP(Neg) {
case 8:
neg(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unsupported Not size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Not size: %d", OpSize);
}
}
@@ -142,12 +159,11 @@ DEF_OP(Mul) {
switch (OpSize) {
case 4:
mul(Dst.W(), GetReg<RA_32>(Op->Header.Args[0].ID()), GetReg<RA_32>(Op->Header.Args[1].ID()));
sxtw(Dst, Dst);
break;
case 8:
mul(Dst, GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -163,7 +179,7 @@ DEF_OP(UMul) {
case 8:
mul(Dst, GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -200,7 +216,7 @@ DEF_OP(Div) {
sdiv(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
}
default: LogMan::Msg::A("Unknown DIV Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown DIV Size: %d", Size); break;
}
}
@@ -227,7 +243,7 @@ DEF_OP(UDiv) {
udiv(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
}
default: LogMan::Msg::A("Unknown UDIV Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown UDIV Size: %d", Size); break;
}
}
@@ -274,7 +290,7 @@ DEF_OP(Rem) {
msub(GetReg<RA_64>(Node), TMP1, Divisor, Dividend);
break;
}
default: LogMan::Msg::A("Unknown REM Size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown REM Size: %d", OpSize); break;
}
}
@@ -316,7 +332,7 @@ DEF_OP(URem) {
msub(GetReg<RA_64>(Node), TMP1, Divisor, Dividend);
break;
}
default: LogMan::Msg::A("Unknown UREM Size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown UREM Size: %d", OpSize); break;
}
}
@@ -328,12 +344,12 @@ DEF_OP(MulH) {
sxtw(TMP1, GetReg<RA_64>(Op->Header.Args[0].ID()));
sxtw(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
mul(TMP1, TMP1, TMP2);
sbfx(GetReg<RA_64>(Node), TMP1, 32, 32);
ubfx(GetReg<RA_64>(Node), TMP1, 32, 32);
break;
case 8:
smulh(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -350,7 +366,7 @@ DEF_OP(UMulH) {
case 8:
umulh(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -386,7 +402,6 @@ DEF_OP(Xor) {
DEF_OP(Lshl) {
auto Op = IROp->C<IR::IROp_Lshl>();
uint8_t OpSize = IROp->Size;
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
lsl(GRS(Node), GRS(Op->Header.Args[0].ID()), (unsigned int)Const);
@@ -397,7 +412,6 @@ DEF_OP(Lshl) {
DEF_OP(Lshr) {
auto Op = IROp->C<IR::IROp_Lshr>();
uint8_t OpSize = IROp->Size;
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
lsr(GRS(Node), GRS(Op->Header.Args[0].ID()), (unsigned int)Const);
@@ -448,7 +462,7 @@ DEF_OP(Ror) {
break;
}
default: LogMan::Msg::A("Unhandled ROR size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled ROR size: %d", OpSize);
}
} else {
switch (OpSize) {
@@ -461,7 +475,7 @@ DEF_OP(Ror) {
break;
}
default: LogMan::Msg::A("Unhandled ROR size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled ROR size: %d", OpSize);
}
}
}
@@ -480,7 +494,7 @@ DEF_OP(Extr) {
break;
}
default: LogMan::Msg::A("Unhandled EXTR size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled EXTR size: %d", OpSize);
}
}
@@ -512,14 +526,10 @@ DEF_OP(LDiv) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LDIV);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -529,7 +539,7 @@ DEF_OP(LDiv) {
mov(GetReg<RA_64>(Node), x0);
break;
}
default: LogMan::Msg::A("Unknown LDIV Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown LDIV Size: %d", Size); break;
}
}
@@ -559,14 +569,10 @@ DEF_OP(LUDiv) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LUDIV);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LUDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -576,7 +582,7 @@ DEF_OP(LUDiv) {
mov(GetReg<RA_64>(Node), x0);
break;
}
default: LogMan::Msg::A("Unknown LUDIV Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown LUDIV Size: %d", Size); break;
}
}
@@ -616,14 +622,10 @@ DEF_OP(LRem) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LREM);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Result is now in x0
// Fix the stack and any values that were stepped on
@@ -633,7 +635,7 @@ DEF_OP(LRem) {
mov(GetReg<RA_64>(Node), x0);
break;
}
default: LogMan::Msg::A("Unknown LREM Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown LREM Size: %d", Size); break;
}
}
@@ -670,14 +672,11 @@ DEF_OP(LURem) {
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
#if _M_X86_64
CallRuntime(LUREM);
#else
LoadConstant(x3, reinterpret_cast<uint64_t>(LUREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
#endif
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
@@ -686,7 +685,7 @@ DEF_OP(LURem) {
mov(GetReg<RA_64>(Node), x0);
break;
}
default: LogMan::Msg::A("Unknown LUREM Size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown LUREM Size: %d", OpSize); break;
}
}
@@ -700,7 +699,7 @@ DEF_OP(Not) {
case 8:
mvn(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unsupported Not size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Not size: %d", OpSize);
}
}
@@ -731,7 +730,7 @@ DEF_OP(Popcount) {
// fmov has zero extended, unused bytes are zero
addv(VTMP1.B(), VTMP1.V8B());
break;
default: LogMan::Msg::A("Unsupported Popcount size: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported Popcount size: %d", OpSize);
}
auto Dst = GetReg<RA_32>(Node);
@@ -780,7 +779,7 @@ DEF_OP(FindMSB) {
clz(Dst, GetReg<RA_64>(Op->Header.Args[0].ID()));
sub(Dst, TMP1, Dst);
break;
default: LogMan::Msg::A("Unknown REV size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown REV size: %d", OpSize); break;
}
}
@@ -801,7 +800,7 @@ DEF_OP(FindTrailingZeros) {
rbit(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
clz(GetReg<RA_64>(Node), GetReg<RA_64>(Node));
break;
default: LogMan::Msg::A("Unknown size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown size: %d", OpSize); break;
}
}
@@ -820,7 +819,7 @@ DEF_OP(CountLeadingZeroes) {
case 8:
clz(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unknown size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown size: %d", OpSize); break;
}
}
@@ -838,7 +837,7 @@ DEF_OP(Rev) {
case 8:
rev(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unknown REV size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown REV size: %d", OpSize); break;
}
}
@@ -860,15 +859,15 @@ DEF_OP(Bfi) {
bfi(TMP1, GetReg<RA_64>(Op->Header.Args[1].ID()), Op->lsb, Op->Width);
mov(GetReg<RA_64>(Node), TMP1);
break;
default: LogMan::Msg::A("Unknown BFI size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown BFI size: %d", OpSize); break;
}
}
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
uint8_t OpSize = IROp->Size;
LogMan::Throw::A(OpSize <= 8, "OpSize is too large for BFE: %d", OpSize);
LogMan::Throw::A(Op->Width != 0, "Invalid BFE width of 0");
LOGMAN_THROW_A(OpSize <= 8, "OpSize is too large for BFE: %d", OpSize);
LOGMAN_THROW_A(Op->Width != 0, "Invalid BFE width of 0");
auto Dst = GetReg<RA_64>(Node);
ubfx(Dst, GetReg<RA_64>(Op->Header.Args[0].ID()), Op->lsb, Op->Width);
@@ -908,12 +907,12 @@ Condition MapSelectCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:;
case FEXCore::IR::COND_VC:;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
default:
LogMan::Msg::A("Unsupported compare type");
LOGMAN_MSG_A("Unsupported compare type");
return Condition::nv;
}
}
@@ -931,9 +930,9 @@ DEF_OP(Select) {
} else if (IsFPR(Op->Cmp1.ID())) {
fcmp(GRFCMP(Op->Cmp1.ID()), GRFCMP(Op->Cmp2.ID()));
} else {
LogMan::Msg::A("Select: Expected GPR or FPR");
LOGMAN_MSG_A("Select: Expected GPR or FPR");
}
auto cc = MapSelectCC(Op->Cond);
uint64_t const_true, const_false;
@@ -942,7 +941,7 @@ DEF_OP(Select) {
if (is_const_true || is_const_false) {
if (is_const_false != true || is_const_true != true || const_true != 1 || const_false != 0) {
LogMan::Msg::A("Select: Unsupported compare inline parameters");
LOGMAN_MSG_A("Select: Unsupported compare inline parameters");
}
cset(GRS(Node), cc);
} else {
@@ -966,38 +965,53 @@ DEF_OP(VExtractToGPR) {
case 8:
umov(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Idx);
break;
default: LogMan::Msg::A("Unhandled ExtractElementSize: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled ExtractElementSize: %d", OpSize);
}
}
DEF_OP(Float_ToGPR_ZU) {
LogMan::Msg::D("Unimplemented");
}
DEF_OP(Float_ToGPR_ZS) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_ZS>();
if (Op->Header.ElementSize == 8) {
fcvtzs(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).D());
aarch64::Register Dst{};
aarch64::VRegister Src{};
if (Op->SrcElementSize == 8) {
Src = GetSrc(Op->Header.Args[0].ID()).D();
}
else {
fcvtzs(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).S());
Src = GetSrc(Op->Header.Args[0].ID()).S();
}
}
DEF_OP(Float_ToGPR_U) {
LogMan::Msg::D("Unimplemented");
if (IROp->Size == 8) {
Dst = GetReg<RA_64>(Node);
}
else {
Dst = GetReg<RA_32>(Node);
}
fcvtzs(Dst, Src);
}
DEF_OP(Float_ToGPR_S) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_S>();
if (Op->Header.ElementSize == 8) {
aarch64::Register Dst{};
aarch64::VRegister Src{};
if (Op->SrcElementSize == 8) {
frinti(VTMP1.D(), GetSrc(Op->Header.Args[0].ID()).D());
fcvtzs(GetReg<RA_64>(Node), VTMP1.D());
Src = VTMP1.D();
}
else {
frinti(VTMP1.S(), GetSrc(Op->Header.Args[0].ID()).S());
fcvtzs(GetReg<RA_32>(Node), VTMP1.S());
Src = VTMP1.S();
}
if (IROp->Size == 8) {
Dst = GetReg<RA_64>(Node);
}
else {
Dst = GetReg<RA_32>(Node);
}
fcvtzs(Dst, Src);
}
DEF_OP(FCmp) {
@@ -1010,11 +1024,11 @@ DEF_OP(FCmp) {
fcmp(GetSrc(Op->Header.Args[0].ID()).D(), GetSrc(Op->Header.Args[1].ID()).D());
}
auto Dst = GetReg<RA_64>(Node);
bool set = false;
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ)) {
LogMan::Throw::A(IR::FCMP_FLAG_EQ == 0, "IR::FCMP_FLAG_EQ must equal 0");
LOGMAN_THROW_A(IR::FCMP_FLAG_EQ == 0, "IR::FCMP_FLAG_EQ must equal 0");
// EQ or unordered
cset(Dst, Condition::eq); // Z = 1
csinc(Dst, Dst, xzr, Condition::vc); // IF !V ? Z : 1
@@ -1043,17 +1057,15 @@ DEF_OP(FCmp) {
}
}
DEF_OP(F80Cmp) {
LogMan::Msg::D("Unimplemented");
}
#undef DEF_OP
void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
@@ -1090,12 +1102,9 @@ void JITCore::RegisterALUHandlers() {
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZU, Float_ToGPR_ZU);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_U, Float_ToGPR_U);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
REGISTER_OP(F80CMP, F80Cmp);
#undef REGISTER_OP
}
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
@@ -28,7 +34,7 @@ DEF_OP(CASPair) {
mov(Dst.first, TMP3);
mov(Dst.second, TMP4);
break;
default: LogMan::Msg::A("Unsupported: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported: %d", OpSize);
}
}
else {
@@ -83,7 +89,7 @@ DEF_OP(CASPair) {
bind(&LoopExpected);
break;
}
default: LogMan::Msg::A("Unsupported: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported: %d", OpSize);
}
}
}
@@ -109,7 +115,7 @@ DEF_OP(CAS) {
case 2: casalh(TMP2.W(), Desired.W(), MemOperand(MemSrc)); break;
case 4: casal(TMP2.W(), Desired.W(), MemOperand(MemSrc)); break;
case 8: casal(TMP2.X(), Desired.X(), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unsupported: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported: %d", OpSize);
}
mov(GetReg<RA_64>(Node), TMP2);
}
@@ -200,7 +206,7 @@ DEF_OP(CAS) {
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", OpSize);
}
}
}
@@ -216,7 +222,7 @@ DEF_OP(AtomicAdd) {
case 2: staddlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -258,7 +264,7 @@ DEF_OP(AtomicAdd) {
cbnz(TMP2, &LoopTop);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -275,7 +281,7 @@ DEF_OP(AtomicSub) {
case 2: staddlh(TMP2.W(), MemOperand(MemSrc)); break;
case 4: staddl(TMP2.W(), MemOperand(MemSrc)); break;
case 8: staddl(TMP2.X(), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -317,7 +323,7 @@ DEF_OP(AtomicSub) {
cbnz(TMP2, &LoopTop);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -334,7 +340,7 @@ DEF_OP(AtomicAnd) {
case 2: stclrlh(TMP2.W(), MemOperand(MemSrc)); break;
case 4: stclrl(TMP2.W(), MemOperand(MemSrc)); break;
case 8: stclrl(TMP2.X(), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -376,7 +382,7 @@ DEF_OP(AtomicAnd) {
cbnz(TMP2, &LoopTop);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -392,7 +398,7 @@ DEF_OP(AtomicOr) {
case 2: stsetlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -434,7 +440,7 @@ DEF_OP(AtomicOr) {
cbnz(TMP2, &LoopTop);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -450,7 +456,7 @@ DEF_OP(AtomicXor) {
case 2: steorlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -492,7 +498,7 @@ DEF_OP(AtomicXor) {
cbnz(TMP2, &LoopTop);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -509,7 +515,7 @@ DEF_OP(AtomicSwap) {
case 2: swplh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpl(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpl(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -552,7 +558,7 @@ DEF_OP(AtomicSwap) {
mov(GetReg<RA_64>(Node), TMP2.X());
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -567,7 +573,7 @@ DEF_OP(AtomicFetchAdd) {
case 2: ldaddalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -613,7 +619,7 @@ DEF_OP(AtomicFetchAdd) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -629,7 +635,7 @@ DEF_OP(AtomicFetchSub) {
case 2: ldaddalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -675,7 +681,7 @@ DEF_OP(AtomicFetchSub) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -691,7 +697,7 @@ DEF_OP(AtomicFetchAnd) {
case 2: ldclralh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldclral(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldclral(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -737,7 +743,7 @@ DEF_OP(AtomicFetchAnd) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -752,7 +758,7 @@ DEF_OP(AtomicFetchOr) {
case 2: ldsetalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -798,7 +804,7 @@ DEF_OP(AtomicFetchOr) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
@@ -813,7 +819,7 @@ DEF_OP(AtomicFetchXor) {
case 2: ldeoralh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
else {
@@ -859,14 +865,14 @@ DEF_OP(AtomicFetchXor) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LogMan::Msg::A("Unhandled Atomic size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled Atomic size: %d", Op->Size);
}
}
}
#undef DEF_OP
void JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
@@ -1,13 +1,20 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::D("Unimplemented");
}
@@ -22,9 +29,7 @@ DEF_OP(GuestReturn) {
DEF_OP(SignalReturn) {
// First we must reset the stack
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
@@ -36,11 +41,9 @@ DEF_OP(CallbackReturn) {
// spill back to CTX
SpillStaticRegs();
// First we must reset the stack
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
// We can now lower the ref counter again
LoadConstant(x0, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
@@ -49,9 +52,9 @@ DEF_OP(CallbackReturn) {
str(w2, MemOperand(x0));
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
add(x2, x2, 8);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])));
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
PopCalleeSavedRegisters();
@@ -64,15 +67,13 @@ DEF_OP(ExitFunction) {
Label FullLookup;
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
aarch64::Register RipReg;
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ExitFunctionLinkerAddress};
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchGuest{NewRIP};
ldr(x0, &l_BranchHost);
@@ -82,9 +83,9 @@ DEF_OP(ExitFunction) {
place(&l_BranchGuest);
} else {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// L1 Cache
LoadConstant(x0, State->LookupCache->GetL1Pointer());
LoadConstant(x0, ThreadState->LookupCache->GetL1Pointer());
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -95,8 +96,8 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
LoadConstant(TMP1, AbsoluteLoopTopAddress);
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, State.rip)));
LoadConstant(TMP1, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
}
@@ -136,12 +137,12 @@ Condition MapBranchCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:;
case FEXCore::IR::COND_VC:;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
default:
LogMan::Msg::A("Unsupported compare type");
LOGMAN_MSG_A("Unsupported compare type");
return Condition::nv;
}
}
@@ -168,10 +169,10 @@ DEF_OP(CondJump) {
bool isConst = IsInlineConstant(Op->Cmp2, &Const);
if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_EQ) {
LogMan::Throw::A(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
LOGMAN_THROW_A(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
cbz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
} else if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_NEQ) {
LogMan::Throw::A(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
LOGMAN_THROW_A(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
cbnz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
} else {
if (IsGPR(Op->Cmp1.ID())) {
@@ -182,12 +183,12 @@ DEF_OP(CondJump) {
} else if (IsFPR(Op->Cmp1.ID())) {
fcmp(GRFCMP(Op->Cmp1.ID()), GRFCMP(Op->Cmp2.ID()));
} else {
LogMan::Msg::A("CondJump: Expected GPR or FPR");
LOGMAN_MSG_A("CondJump: Expected GPR or FPR");
}
b(TrueTargetLabel, MapBranchCC(Op->Cond));
}
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
@@ -222,7 +223,7 @@ DEF_OP(Syscall) {
blr(x3);
add(sp, sp, SPOffset);
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs();
@@ -244,30 +245,35 @@ DEF_OP(Thunk) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
#if _M_X86_64
ERROR_AND_DIE("JIT: OP_THUNK not supported with arm simulator")
#else
LoadConstant(x2, Op->ThunkFnPtr);
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
blr(x2);
#endif
PopDynamicRegsAndLR();
FillStaticRegs(); // load from ctx after ra64 refill
}
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
uint8_t *NewCode = (uint8_t *)Op->CodePtr;
uint8_t *OldCode = (uint8_t *)&Op->CodeOriginalLow;
int len = Op->CodeLength;
int idx = 0;
LoadConstant(GetReg<RA_64>(Node), 0);
LoadConstant(x0, Op->CodePtr);
LoadConstant(x0, Entry + Op->Offset);
LoadConstant(x1, 1);
while (len >= 8)
{
ldr(x2, MemOperand(x0, idx));
LoadConstant(x3, *(uint32_t *)(OldCode + idx));
cmp(x2, x3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
len -= 8;
idx += 8;
}
while (len >= 4)
{
ldr(w2, MemOperand(x0, idx));
@@ -298,17 +304,16 @@ DEF_OP(ValidateCode) {
}
DEF_OP(RemoveCodeEntry) {
auto Op = IROp->C<IR::IROp_RemoveCodeEntry>();
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
PushDynamicRegsAndLR();
mov(x0, STATE);
LoadConstant(x1, Op->RIP);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntry));
LoadConstant(x1, Entry);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -319,15 +324,17 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR();
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
LoadConstant(x0, reinterpret_cast<uint64_t>(&CTX->CPUID));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t);
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t, uint32_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
@@ -350,8 +357,8 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -25,7 +31,7 @@ DEF_OP(VInsGPR) {
ins(GetDst(Node).V2D(), Op->Index, GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
}
default: LogMan::Msg::A("Unknown Element Size: %d", Op->Header.ElementSize); break;
default: LOGMAN_MSG_A("Unknown Element Size: %d", Op->Header.ElementSize); break;
}
}
@@ -46,14 +52,10 @@ DEF_OP(VCastFromGPR) {
case 8:
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Header.Args[0].ID()).X());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Float_FromGPR_U) {
LogMan::Msg::A("Unimplemented");
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
@@ -89,20 +91,7 @@ DEF_OP(Float_FToF) {
fcvt(GetDst(Node).S(), GetSrc(Op->Header.Args[0].ID()).D());
break;
}
default: LogMan::Msg::A("Unknown FCVT sizes: 0x%x", Conv);
}
}
DEF_OP(Vector_UToF) {
auto Op = IROp->C<IR::IROp_Vector_UToF>();
switch (Op->Header.ElementSize) {
case 4:
ucvtf(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
ucvtf(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown FCVT sizes: 0x%x", Conv);
}
}
@@ -115,20 +104,7 @@ DEF_OP(Vector_SToF) {
case 8:
scvtf(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Vector_FToZU) {
auto Op = IROp->C<IR::IROp_Vector_FToZU>();
switch (Op->Header.ElementSize) {
case 4:
fcvtzu(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
fcvtzu(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
@@ -141,22 +117,7 @@ DEF_OP(Vector_FToZS) {
case 8:
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Vector_FToU) {
auto Op = IROp->C<IR::IROp_Vector_FToU>();
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
fcvtzu(GetDst(Node).V4S(), GetDst(Node).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtzu(GetDst(Node).V2D(), GetDst(Node).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
@@ -171,7 +132,7 @@ DEF_OP(Vector_FToS) {
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetDst(Node).V2D());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
@@ -188,25 +149,78 @@ DEF_OP(Vector_FToF) {
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
default: LogMan::Msg::A("Unknown Conversion Type : 0%04x", Conv); break;
default: LOGMAN_MSG_A("Unknown Conversion Type : 0%04x", Conv); break;
}
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
case 4:
frintn(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
frintn(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintm(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
frintm(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintp(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
frintp(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
case 4:
frintz(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
frintz(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Header.Args[0].ID()).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Header.Args[0].ID()).V2D());
break;
}
break;
}
}
#undef DEF_OP
void JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_U, Float_FromGPR_U);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_UTOF, Vector_UToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZU, Vector_FToZU);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOU, Vector_FToU);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -83,8 +89,8 @@ DEF_OP(AESKeyGenAssist) {
}
#undef DEF_OP
void JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
@@ -1,18 +1,24 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
}
#undef DEF_OP
void JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
File diff suppressed because it is too large. Load diff
+50 -111
View File
@@ -1,9 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#pragma once
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
@@ -29,56 +36,18 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
// All but x29 are caller saved
const std::array<aarch64::Register, 16> SRA64 = {
x4, x5, x6, x7, x8, x9, x10, x11,
x12, x18, x17, x16, x15, x14, x13, x29
};
// All are callee saved
const std::array<aarch64::Register, 9> RA64 = {
x20, x21, x22, x23, x24, x25, x26, x27,
x19
};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA64Pair = {{
{x20, x21},
{x22, x23},
{x24, x25},
{x26, x27},
}};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA32Pair = {{
{w20, w21},
{w22, w23},
{w24, w25},
{w26, w27},
}};
// All are caller saved
const std::array<aarch64::VRegister, 16> SRAFPR = {
v16, v17, v18, v19, v20, v21, v22, v23,
v24, v25, v26, v27, v28, v29, v30, v31
};
// v8..v15 = (lower 64bits) Callee saved
const std::array<aarch64::VRegister, 12> RAFPR = {
/*v0, v1, v2, v3,*/v4, v5, v6, v7, // v0 ~ v3 are used as temps
v8, v9, v10, v11, v12, v13, v14, v15
};
class JITCore final : public CPUBackend, public vixl::aarch64::Assembler {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
struct CodeBuffer {
uint8_t *Ptr;
size_t Size;
};
explicit JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
explicit Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
~JITCore() override;
~Arm64JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -86,21 +55,22 @@ public:
void ClearCache() override;
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static CodeBuffer AllocateNewCodeBuffer(size_t Size);
CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
FEXCore::IR::IRListView<true> const *IR;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
std::map<IR::OrderedNodeWrapper::NodeOffsetType, aarch64::Label> JumpTargets;
@@ -126,32 +96,35 @@ private:
constexpr static uint8_t RA_FPR = 2;
template<uint8_t RAType>
aarch64::Register GetReg(uint32_t Node);
aarch64::Register GetReg(uint32_t Node) const;
template<>
aarch64::Register GetReg<RA_32>(uint32_t Node);
aarch64::Register GetReg<RA_32>(uint32_t Node) const;
template<>
aarch64::Register GetReg<RA_64>(uint32_t Node);
aarch64::Register GetReg<RA_64>(uint32_t Node) const;
template<uint8_t RAType>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node);
std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node) const;
template<>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node);
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node) const;
template<>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node);
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node) const;
aarch64::VRegister GetSrc(uint32_t Node);
aarch64::VRegister GetDst(uint32_t Node);
aarch64::VRegister GetSrc(uint32_t Node) const;
aarch64::VRegister GetDst(uint32_t Node) const;
FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node);
FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node) const;
bool IsFPR(uint32_t Node);
bool IsGPR(uint32_t Node);
IR::PhysicalRegister GetPhys(uint32_t Node) const;
bool IsFPR(uint32_t Node) const;
bool IsGPR(uint32_t Node) const;
MemOperand GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
struct LiveRange {
uint32_t Begin;
@@ -161,9 +134,6 @@ private:
#if DEBUG
vixl::aarch64::Decoder Decoder;
#endif
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
@@ -180,8 +150,6 @@ private:
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the codebuffer that our dispatcher lives in
CodeBuffer DispatcherCodeBuffer{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
@@ -189,60 +157,24 @@ private:
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true);
#if DEBUG
vixl::aarch64::Disassembler Disasm;
#endif
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread);
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
static uint64_t ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadState *Thread, uint64_t *record);
/**
* @name Dispatch Helper functions
* @{ */
uint64_t AbsoluteLoopTopAddressFillSRA{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t ThreadStopHandlerAddressSpillSRA{};
uint64_t ThreadStopHandlerAddress{};
uint64_t PauseReturnInstruction{};
uint32_t SignalHandlerRefCounter{};
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
/** @} */
static uint64_t ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
struct CompilerSharedData {
uint64_t InterpreterFallbackHelperAddress{};
uint64_t SignalReturnInstruction{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
uint32_t SpillSlots{};
void SpillStaticRegs();
void FillStaticRegs();
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
using OpHandler = void (JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
using OpHandler = void (Arm64JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -265,7 +197,9 @@ private:
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
DEF_OP(CycleCounter);
DEF_OP(Add);
DEF_OP(Sub);
@@ -306,10 +240,8 @@ private:
DEF_OP(VExtractToGPR);
DEF_OP(Float_ToGPR_ZU);
DEF_OP(Float_ToGPR_ZS);
DEF_OP(Float_ToGPR_U);
DEF_OP(Float_ToGPR_S);
DEF_OP(FCmp);
DEF_OP(F80Cmp);
///< Atomic ops
DEF_OP(CASPair);
@@ -344,16 +276,13 @@ private:
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(Float_FromGPR_U);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_UToF);
DEF_OP(Vector_SToF);
DEF_OP(Vector_FToZU);
DEF_OP(Vector_FToZS);
DEF_OP(Vector_FToU);
DEF_OP(Vector_FToS);
DEF_OP(Vector_FToF);
DEF_OP(Vector_FToI);
///< Flag ops
DEF_OP(GetHostFlag);
@@ -373,8 +302,11 @@ private:
DEF_OP(StoreMem);
DEF_OP(LoadMemTSO);
DEF_OP(StoreMemTSO);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
///< Misc ops
DEF_OP(EndBlock);
@@ -400,6 +332,7 @@ private:
DEF_OP(SplatVector4);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
DEF_OP(VOr);
DEF_OP(VXor);
DEF_OP(VAdd);
@@ -410,8 +343,10 @@ private:
DEF_OP(VSQSub);
DEF_OP(VAddP);
DEF_OP(VAddV);
DEF_OP(VUMinV);
DEF_OP(VURAvg);
DEF_OP(VAbs);
DEF_OP(VPopcount);
DEF_OP(VFAdd);
DEF_OP(VFAddP);
DEF_OP(VFSub);
@@ -431,6 +366,8 @@ private:
DEF_OP(VSMax);
DEF_OP(VZip);
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -453,6 +390,7 @@ private:
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
@@ -475,6 +413,7 @@ private:
DEF_OP(VSMull);
DEF_OP(VUMull2);
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
///< Encryption ops
+245 -73
View File
@@ -1,10 +1,17 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -23,7 +30,7 @@ DEF_OP(LoadContext) {
case 8:
ldr(GetReg<RA_64>(Node), MemOperand(STATE, Op->Offset));
break;
default: LogMan::Msg::A("Unhandled LoadContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled LoadContext size: %d", OpSize);
}
}
else {
@@ -44,7 +51,7 @@ DEF_OP(LoadContext) {
case 16:
ldr(Dst, MemOperand(STATE, Op->Offset));
break;
default: LogMan::Msg::A("Unhandled LoadContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled LoadContext size: %d", OpSize);
}
}
}
@@ -66,7 +73,7 @@ DEF_OP(StoreContext) {
case 8:
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, Op->Offset));
break;
default: LogMan::Msg::A("Unhandled StoreContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled StoreContext size: %d", OpSize);
}
}
else {
@@ -87,7 +94,7 @@ DEF_OP(StoreContext) {
case 16:
str(Src, MemOperand(STATE, Op->Offset));
break;
default: LogMan::Msg::A("Unhandled LoadContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled LoadContext size: %d", OpSize);
}
}
}
@@ -95,61 +102,60 @@ DEF_OP(StoreContext) {
DEF_OP(LoadRegister) {
auto Op = IROp->C<IR::IROp_LoadRegister>();
uint8_t OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.gregs[0])) / 8;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.gregs[0])) / 8;
auto regOffs = Op->Offset & 7;
LogMan::Throw::A(regId < SRA64.size(), "out of range regId");
LOGMAN_THROW_A(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
switch(Op->Header.Size) {
case 1:
LogMan::Throw::A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
ubfx(GetReg<RA_64>(Node), reg, regOffs * 8, 8);
break;
case 2:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
ubfx(GetReg<RA_64>(Node), reg, 0, 16);
break;
case 4:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
if (GetReg<RA_64>(Node).GetCode() != reg.GetCode())
mov(GetReg<RA_32>(Node), reg.W());
break;
case 8:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
if (GetReg<RA_64>(Node).GetCode() != reg.GetCode())
mov(GetReg<RA_64>(Node), reg);
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regOffs = Op->Offset & 15;
LogMan::Throw::A(regId < SRAFPR.size(), "out of range regId");
LOGMAN_THROW_A(regId < SRAFPR.size(), "out of range regId");
auto guest = SRAFPR[regId];
auto host = GetSrc(Node);
switch(Op->Header.Size) {
case 1:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
mov(host.B(), guest.B());
break;
case 2:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
fmov(host.H(), guest.H());
break;
case 4:
LogMan::Throw::A((regOffs & 3) == 0, "unexpected regOffs");
LOGMAN_THROW_A((regOffs & 3) == 0, "unexpected regOffs");
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode())
fmov(host.S(), guest.S());
@@ -159,7 +165,7 @@ DEF_OP(LoadRegister) {
break;
case 8:
LogMan::Throw::A((regOffs & 7) == 0, "unexpected regOffs");
LOGMAN_THROW_A((regOffs & 7) == 0, "unexpected regOffs");
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode())
mov(host.D(), guest.D());
@@ -169,56 +175,54 @@ DEF_OP(LoadRegister) {
break;
case 16:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
if (host.GetCode() != guest.GetCode())
mov(host.Q(), guest.Q());
break;
}
} else {
LogMan::Throw::A(false, "Unhandled Op->Class %d", Op->Class);
LOGMAN_THROW_A(false, "Unhandled Op->Class %d", Op->Class);
}
}
DEF_OP(StoreRegister) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
uint8_t OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = Op->Offset / 8 - 1;
auto regOffs = Op->Offset & 7;
LogMan::Throw::A(regId < SRA64.size(), "out of range regId");
LOGMAN_THROW_A(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
switch(Op->Header.Size) {
case 1:
LogMan::Throw::A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0 || regOffs == 1, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), regOffs * 8, 8);
break;
case 2:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), 0, 16);
break;
case 4:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), 0, 32);
break;
case 8:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
if (GetReg<RA_64>(Op->Value.ID()).GetCode() != reg.GetCode())
mov(reg, GetReg<RA_64>(Op->Value.ID()));
break;
}
} else if (Op->Class == IR::FPRClass) {
auto regId = (Op->Offset - offsetof(FEXCore::Core::ThreadState, State.xmm[0][0])) / 16;
auto regId = (Op->Offset - offsetof(FEXCore::Core::CpuStateFrame, State.xmm[0][0])) / 16;
auto regOffs = Op->Offset & 15;
LogMan::Throw::A(regId < SRAFPR.size(), "regId out of range");
LOGMAN_THROW_A(regId < SRAFPR.size(), "regId out of range");
auto guest = SRAFPR[regId];
auto host = GetSrc(Op->Value.ID());
@@ -229,28 +233,28 @@ DEF_OP(StoreRegister) {
break;
case 2:
LogMan::Throw::A((regOffs & 1) == 0, "unexpected regOffs");
LOGMAN_THROW_A((regOffs & 1) == 0, "unexpected regOffs");
ins(guest.V8H(), regOffs/2, host.V8H(), 0);
break;
case 4:
LogMan::Throw::A((regOffs & 3) == 0, "unexpected regOffs");
LOGMAN_THROW_A((regOffs & 3) == 0, "unexpected regOffs");
ins(guest.V4S(), regOffs/4, host.V4S(), 0);
break;
case 8:
LogMan::Throw::A((regOffs & 7) == 0, "unexpected regOffs");
LOGMAN_THROW_A((regOffs & 7) == 0, "unexpected regOffs");
ins(guest.V2D(), regOffs / 8, host.V2D(), 0);
break;
case 16:
LogMan::Throw::A(regOffs == 0, "unexpected regOffs");
LOGMAN_THROW_A(regOffs == 0, "unexpected regOffs");
if (guest.GetCode() != host.GetCode())
mov(guest.Q(), host.Q());
break;
}
} else {
LogMan::Throw::A(false, "Unhandled Op->Class %d", Op->Class);
LOGMAN_THROW_A(false, "Unhandled Op->Class %d", Op->Class);
}
}
@@ -284,15 +288,15 @@ DEF_OP(LoadContextIndexed) {
ldr(GetReg<RA_64>(Node), MemOperand(TMP1, Op->BaseOffset));
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
case 16:
LogMan::Msg::A("Invalid Class load of size 16");
LOGMAN_MSG_A("Invalid Class load of size 16");
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
else {
@@ -329,12 +333,12 @@ DEF_OP(LoadContextIndexed) {
}
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
}
@@ -370,15 +374,15 @@ DEF_OP(StoreContextIndexed) {
str(value, MemOperand(TMP1, Op->BaseOffset));
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
case 16:
LogMan::Msg::A("Invalid Class load of size 16");
LOGMAN_MSG_A("Invalid Class load of size 16");
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
else {
@@ -417,12 +421,12 @@ DEF_OP(StoreContextIndexed) {
}
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
}
@@ -450,7 +454,7 @@ DEF_OP(SpillRegister) {
str(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
@@ -466,10 +470,10 @@ DEF_OP(SpillRegister) {
str(GetSrc(Op->Header.Args[0].ID()), MemOperand(sp, SlotOffset));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else {
LogMan::Msg::A("Unhandled SpillRegister class: %d", Op->Class.Val);
LOGMAN_MSG_A("Unhandled SpillRegister class: %d", Op->Class.Val);
}
}
@@ -496,7 +500,7 @@ DEF_OP(FillRegister) {
ldr(GetReg<RA_64>(Node), MemOperand(sp, SlotOffset));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
@@ -512,10 +516,10 @@ DEF_OP(FillRegister) {
ldr(GetDst(Node), MemOperand(sp, SlotOffset));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else {
LogMan::Msg::A("Unhandled FillRegister class: %d", Op->Class.Val);
LOGMAN_MSG_A("Unhandled FillRegister class: %d", Op->Class.Val);
}
}
@@ -530,12 +534,12 @@ DEF_OP(StoreFlag) {
strb(GetReg<RA_64>(Op->Header.Args[0].ID()), MemOperand(STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag));
}
MemOperand JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
MemOperand Arm64JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
if (Offset.IsInvalid()) {
return MemOperand(Base);
} else {
if (OffsetScale != 1 && OffsetScale != AccessSize) {
LogMan::Msg::A("Unhandled GenerateMemOperand OffsetScale: %d", OffsetScale);
LOGMAN_MSG_A("Unhandled GenerateMemOperand OffsetScale: %d", OffsetScale);
}
uint64_t Const;
if (IsInlineConstant(Offset, &Const)) {
@@ -547,10 +551,12 @@ MemOperand JITCore::GenerateMemOperand(uint8_t AccessSize, aarch64::Register Bas
case IR::MEM_OFFSET_UXTW.Val: return MemOperand(Base, RegOffset.W(), Extend::UXTW, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_SXTW.Val: return MemOperand(Base, RegOffset.W(), Extend::SXTW, (int)std::log2(OffsetScale) );
default: LogMan::Msg::A("Unhandled GenerateMemOperand OffsetType: %d", OffsetType.Val); break;
default: LOGMAN_MSG_A("Unhandled GenerateMemOperand OffsetType: %d", OffsetType.Val); break;
}
}
}
FEX_UNREACHABLE;
}
DEF_OP(LoadMem) {
@@ -574,7 +580,7 @@ DEF_OP(LoadMem) {
case 8:
ldr(Dst, MemSrc);
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
}
else {
@@ -595,7 +601,7 @@ DEF_OP(LoadMem) {
case 16:
ldr(Dst, MemSrc);
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
}
}
@@ -606,7 +612,7 @@ DEF_OP(LoadMemTSO) {
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
if (!Op->Offset.IsInvalid()) {
LogMan::Msg::A("LoadMemTSO: No offset allowed");
LOGMAN_MSG_A("LoadMemTSO: No offset allowed");
}
if (SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
@@ -629,7 +635,7 @@ DEF_OP(LoadMemTSO) {
case 8:
ldapr(Dst, MemSrc);
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
nop();
}
@@ -654,7 +660,7 @@ DEF_OP(LoadMemTSO) {
case 8:
ldar(Dst, MemSrc);
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
nop();
}
@@ -675,7 +681,7 @@ DEF_OP(LoadMemTSO) {
case 16:
ldr(Dst, MemSrc);
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
dmb(InnerShareable, BarrierAll);
}
@@ -702,7 +708,7 @@ DEF_OP(StoreMem) {
case 8:
str(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
}
else {
@@ -723,7 +729,7 @@ DEF_OP(StoreMem) {
case 16:
str(Src, MemSrc);
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
}
}
@@ -733,7 +739,7 @@ DEF_OP(StoreMemTSO) {
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
if (!Op->Offset.IsInvalid()) {
LogMan::Msg::A("StoreMemTSO: No offset allowed");
LOGMAN_MSG_A("StoreMemTSO: No offset allowed");
}
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -753,7 +759,7 @@ DEF_OP(StoreMemTSO) {
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
nop();
}
@@ -777,23 +783,182 @@ DEF_OP(StoreMemTSO) {
case 16:
str(Src, MemSrc);
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
dmb(InnerShareable, BarrierAll);
}
}
DEF_OP(ParanoidLoadMemTSO) {
auto Op = IROp->C<IR::IROp_LoadMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A("LoadMemTSO: No offset allowed");
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldarb(Dst, MemSrc);
}
else {
auto Dst = GetReg<RA_64>(Node);
nop();
switch (Op->Size) {
case 2:
ldarh(Dst, MemSrc);
break;
case 4:
ldar(Dst.W(), MemSrc);
break;
case 8:
ldar(Dst, MemSrc);
break;
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
nop();
}
}
else {
auto Dst = GetDst(Node);
switch (Op->Size) {
case 2:
nop();
ldarh(TMP1, MemSrc);
nop();
fmov(Dst, TMP1);
break;
case 4:
nop();
ldar(TMP1.W(), MemSrc);
nop();
fmov(Dst, TMP1);
break;
case 8:
nop();
ldar(TMP1, MemSrc);
nop();
fmov(Dst, TMP1);
break;
case 16:
nop();
ldaxp(TMP1, TMP2, MemSrc);
clrex();
mov(Dst.V2D(), 0, TMP1);
mov(Dst.V2D(), 1, TMP2);
break;
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
}
}
DEF_OP(ParanoidStoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A("StoreMemTSO: No offset allowed");
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
}
else {
nop();
switch (Op->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
nop();
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
if (Op->Size == 1) {
// 8bit load is always aligned to natural alignment
mov(TMP1, Src.V4S(), 0);
stlrb(TMP1, MemSrc);
}
else {
switch (Op->Size) {
case 2:
mov(TMP1, Src.V4S(), 0);
nop();
stlrh(TMP1, MemSrc);
nop();
break;
case 4:
mov(TMP1, Src.V4S(), 0);
nop();
stlr(TMP1.W(), MemSrc);
nop();
break;
case 8:
mov(TMP1, Src.V2D(), 0);
nop();
stlr(TMP1, MemSrc);
nop();
break;
case 16: {
// Move vector to GPRs
mov(TMP1, Src.V2D(), 0);
mov(TMP2, Src.V2D(), 1);
Label B;
bind(&B);
nop(); // < Overwritten with DMB
// ldaxp must not have both the destination registers be the same
ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS
nop(); // < Overwritten with DMB
stlxp(TMP3, TMP1, TMP2, MemSrc); // <- Can also hit SIGBUS
cbnz(TMP3, &B); // < Overwritten with DMB
break;
}
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
}
}
}
DEF_OP(VLoadMemElement) {
LogMan::Msg::A("Unimplemented");
LOGMAN_MSG_A("Unimplemented");
}
DEF_OP(VStoreMemElement) {
LogMan::Msg::A("Unimplemented");
LOGMAN_MSG_A("Unimplemented");
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
mov(TMP1, MemReg);
for (size_t i = 0; i < std::max(1U, DCacheLineSize / 64U); ++i) {
dc(DataCacheOp::CVAU, TMP1);
add(TMP1, TMP1, DCacheLineSize);
}
dsb(InnerShareable, BarrierAll);
}
#undef DEF_OP
void JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
@@ -806,10 +971,17 @@ void JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMemTSO);
REGISTER_OP(STOREMEMTSO, StoreMemTSO);
if (ParanoidTSO()) {
REGISTER_OP(LOADMEMTSO, ParanoidLoadMemTSO);
REGISTER_OP(STOREMEMTSO, ParanoidStoreMemTSO);
}
else {
REGISTER_OP(LOADMEMTSO, LoadMemTSO);
REGISTER_OP(STOREMEMTSO, StoreMemTSO);
}
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
#undef REGISTER_OP
}
}
+54 -13
View File
@@ -1,10 +1,23 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
LogMan::Msg::D("Value: 0x%lx", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::D("Value: 0x%016lx'%016lx", ValueUpper, Value);
}
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -18,7 +31,7 @@ DEF_OP(Fence) {
case IR::Fence_Store.Val:
dmb(FullSystem, BarrierWrites);
break;
default: LogMan::Msg::A("Unknown Fence: %d", Op->Fence); break;
default: LOGMAN_MSG_A("Unknown Fence: %d", Op->Fence); break;
}
}
@@ -29,32 +42,38 @@ DEF_OP(Break) {
case 5: // Guest ud2
hlt(4);
break;
case 1: // Int <imm8>
hlt(4);
break;
case 2: // overflow
hlt(4);
break;
case 3: // int 1
hlt(4);
break;
case 4: { // HLT
// Time to quit
// Set our stack to the starting stack location
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)));
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
LoadConstant(TMP1, ThreadStopHandlerAddressSpillSRA);
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddressSpillSRA);
br(TMP1);
break;
}
case 6: { // INT3
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
LoadConstant(TMP1, ThreadPauseHandlerAddressSpillSRA);
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
br(TMP1);
break;
}
default: LogMan::Msg::A("Unknown Break reason: %d", Op->Reason);
default: LOGMAN_MSG_A("Unknown Break reason: %d", Op->Reason);
}
}
DEF_OP(GetRoundingMode) {
auto Op = IROp->C<IR::IROp_GetRoundingMode>();
auto Dst = GetReg<RA_64>(Node);
mrs(Dst, FPCR);
lsr(Dst, Dst, 22);
@@ -112,9 +131,31 @@ DEF_OP(SetRoundingMode) {
msr(FPCR, TMP1);
}
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushDynamicRegsAndLR();
if (IsGPR(Op->Header.Args[0].ID())) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintValue));
}
else {
fmov(x0, GetSrc(Op->Header.Args[0].ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Header.Args[0].ID()).V1D(), 1);
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintVectorValue));
}
SpillStaticRegs();
blr(x3);
FillStaticRegs();
PopDynamicRegsAndLR();
}
#undef DEF_OP
void JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
@@ -124,7 +165,7 @@ void JITCore::RegisterMiscHandlers() {
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Unhandled);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
@@ -1,10 +1,16 @@
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -20,7 +26,7 @@ DEF_OP(ExtractElementPair) {
mov (GetReg<RA_64>(Node), Regs[Op->Element]);
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
default: LOGMAN_MSG_A("Unknown Size"); break;
}
}
@@ -29,21 +35,24 @@ DEF_OP(CreateElementPair) {
std::pair<aarch64::Register, aarch64::Register> Dst;
aarch64::Register RegFirst;
aarch64::Register RegSecond;
aarch64::Register RegTmp;
switch (Op->Header.Size) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_32>(Op->Header.Args[1].ID());
RegTmp = w0;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetReg<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_64>(Op->Header.Args[1].ID());
RegTmp = x0;
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
default: LOGMAN_MSG_A("Unknown Size"); break;
}
if (Dst.first.GetCode() != RegSecond.GetCode()) {
@@ -53,7 +62,9 @@ DEF_OP(CreateElementPair) {
mov(Dst.second, RegSecond);
mov(Dst.first, RegFirst);
} else {
LogMan::Msg::A("Unhandled CreateElementPair");
mov(RegTmp, RegFirst);
mov(Dst.second, RegSecond);
mov(Dst.first, RegTmp);
}
}
@@ -63,8 +74,8 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void Arm64JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
File diff suppressed because it is too large. Load diff
+4 -1
View File
@@ -1,5 +1,7 @@
#pragma once
#include <memory>
namespace FEXCore::Context {
struct Context;
}
@@ -11,5 +13,6 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
FEXCore::CPU::CPUBackend *CreateJITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
}
+121 -83
View File
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -14,7 +20,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LogMan::Msg::A("Unhandled Truncation size: %d", Op->Size); break;
default: LOGMAN_MSG_A("Unhandled Truncation size: %d", Op->Size); break;
}
}
@@ -23,17 +29,28 @@ DEF_OP(Constant) {
mov(GetDst<RA_64>(Node), Op->Constant);
}
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
mov(GetDst<RA_64>(Node), Constant);
}
DEF_OP(InlineConstant) {
//nop
}
DEF_OP(InlineEntrypointOffset) {
//nop
}
DEF_OP(CycleCounter) {
#ifdef DEBUG_CYCLES
mov (GetDst<RA_64>(Node), 0);
#else
rdtsc();
shl(rdx, 32);
or(rax, rdx);
or_(rax, rdx);
mov (GetDst<RA_64>(Node), rax);
#endif
}
@@ -53,7 +70,7 @@ DEF_OP(Add) {
case 8:
add(rax, Const);
break;
default: LogMan::Msg::A("Unhandled Add size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Add size: %d", OpSize);
break;
}
} else {
@@ -64,7 +81,7 @@ DEF_OP(Add) {
case 8:
add(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled Add size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Add size: %d", OpSize);
break;
}
}
@@ -86,7 +103,7 @@ DEF_OP(Sub) {
case 8:
sub(rax, Const);
break;
default: LogMan::Msg::A("Unhandled Sub size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Sub size: %d", OpSize);
break;
}
} else {
@@ -97,7 +114,7 @@ DEF_OP(Sub) {
case 8:
sub(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled Sub size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Sub size: %d", OpSize);
break;
}
}
@@ -119,7 +136,7 @@ DEF_OP(Neg) {
Src = GetSrc<RA_64>(Op->Header.Args[0].ID());
Dst = GetDst<RA_64>(Node);
break;
default: LogMan::Msg::A("Unhandled Neg size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled Neg size: %d", OpSize);
break;
}
mov(Dst, Src);
@@ -143,7 +160,7 @@ DEF_OP(Mul) {
imul(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(Dst, rax);
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -162,7 +179,7 @@ DEF_OP(UMul) {
mul(GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), rax);
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -201,7 +218,7 @@ DEF_OP(Div) {
mov(GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unknown UDIV Size: %d", Size); break;
default: LOGMAN_MSG_A("Unknown UDIV Size: %d", Size); break;
}
}
@@ -244,7 +261,7 @@ DEF_OP(UDiv) {
mov(GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unknown UDIV OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown UDIV OpSize: %d", OpSize); break;
}
}
@@ -281,7 +298,7 @@ DEF_OP(Rem) {
mov(GetDst<RA_64>(Node), rdx);
break;
}
default: LogMan::Msg::A("Unknown UDIV Size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown UDIV Size: %d", OpSize); break;
}
}
@@ -324,7 +341,7 @@ DEF_OP(URem) {
mov(GetDst<RA_64>(Node), rdx);
break;
}
default: LogMan::Msg::A("Unknown UDIV OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown UDIV OpSize: %d", OpSize); break;
}
}
@@ -343,7 +360,7 @@ DEF_OP(MulH) {
imul(GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), rdx);
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -362,7 +379,7 @@ DEF_OP(UMulH) {
mul(GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), rdx);
break;
default: LogMan::Msg::A("Unknown Sext size: %d", OpSize);
default: LOGMAN_MSG_A("Unknown Sext size: %d", OpSize);
}
}
@@ -373,9 +390,9 @@ DEF_OP(Or) {
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
or (rax, Const);
or_(rax, Const);
} else {
or (rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -386,9 +403,9 @@ DEF_OP(And) {
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
and (rax, Const);
and_(rax, Const);
} else {
and(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -399,9 +416,9 @@ DEF_OP(Xor) {
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
xor(rax, Const);
xor_(rax, Const);
} else {
xor(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -424,11 +441,11 @@ DEF_OP(Lshl) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
shl(GetDst<RA_64>(Node), Const);
break;
default: LogMan::Msg::A("Unknown LSHL Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown LSHL Size: %d\n", OpSize); break;
};
} else {
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 4:
@@ -439,7 +456,7 @@ DEF_OP(Lshl) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
shl(GetDst<RA_64>(Node), cl);
break;
default: LogMan::Msg::A("Unknown LSHL Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown LSHL Size: %d\n", OpSize); break;
};
}
}
@@ -471,12 +488,12 @@ DEF_OP(Lshr) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
shr(GetDst<RA_64>(Node), Const);
break;
default: LogMan::Msg::A("Unknown Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown Size: %d\n", OpSize); break;
};
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 1:
@@ -495,7 +512,7 @@ DEF_OP(Lshr) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
shr(GetDst<RA_64>(Node), cl);
break;
default: LogMan::Msg::A("Unknown Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown Size: %d\n", OpSize); break;
};
}
}
@@ -529,12 +546,12 @@ DEF_OP(Ashr) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
sar(GetDst<RA_64>(Node), Const);
break;
default: LogMan::Msg::A("Unknown ASHR Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown ASHR Size: %d\n", OpSize); break;
};
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 1:
movsx(rax, GetSrc<RA_8>(Op->Header.Args[0].ID()));
@@ -554,7 +571,7 @@ DEF_OP(Ashr) {
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
sar(GetDst<RA_64>(Node), cl);
break;
default: LogMan::Msg::A("Unknown ASHR Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown ASHR Size: %d\n", OpSize); break;
};
}
}
@@ -579,11 +596,11 @@ DEF_OP(Ror) {
ror(rax, Const);
break;
}
default: LogMan::Msg::A("Unknown ROR Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown ROR Size: %d\n", OpSize); break;
}
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 4: {
mov(eax, GetSrc<RA_32>(Op->Header.Args[0].ID()));
@@ -595,7 +612,7 @@ DEF_OP(Ror) {
ror(rax, cl);
break;
}
default: LogMan::Msg::A("Unknown ROR Size: %d\n", OpSize); break;
default: LOGMAN_MSG_A("Unknown ROR Size: %d\n", OpSize); break;
}
}
mov(GetDst<RA_64>(Node), rax);
@@ -651,7 +668,7 @@ DEF_OP(LDiv) {
mov(GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unknown LDIV OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown LDIV OpSize: %d", OpSize); break;
}
}
@@ -683,7 +700,7 @@ DEF_OP(LUDiv) {
mov(GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unknown LUDIV OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown LUDIV OpSize: %d", OpSize); break;
}
}
@@ -715,7 +732,7 @@ DEF_OP(LRem) {
mov(GetDst<RA_64>(Node), rdx);
break;
}
default: LogMan::Msg::A("Unknown LREM OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown LREM OpSize: %d", OpSize); break;
}
}
@@ -747,7 +764,7 @@ DEF_OP(LURem) {
mov(GetDst<RA_64>(Node), rdx);
break;
}
default: LogMan::Msg::A("Unknown LUDIV OpSize: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown LUDIV OpSize: %d", OpSize); break;
}
}
@@ -790,10 +807,10 @@ DEF_OP(FindLSB) {
bsf(rcx, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, 0x40);
cmovz(rcx, rax);
xor(rax, rax);
xor_(rax, rax);
cmp(GetSrc<RA_64>(Op->Header.Args[0].ID()), 1);
sbb(rax, rax);
or(rax, rcx);
or_(rax, rcx);
mov (GetDst<RA_64>(Node), rax);
}
@@ -812,7 +829,7 @@ DEF_OP(FindMSB) {
case 8:
bsr(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unknown OpSize: %d", OpSize);
default: LOGMAN_MSG_A("Unknown OpSize: %d", OpSize);
}
}
@@ -836,7 +853,7 @@ DEF_OP(FindTrailingZeros) {
mov(rax, 0x40);
cmovz(GetDst<RA_64>(Node), rax);
break;
default: LogMan::Msg::A("Unknown size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown size: %d", OpSize); break;
}
}
@@ -859,7 +876,7 @@ DEF_OP(CountLeadingZeroes) {
lzcnt(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
}
default: LogMan::Msg::A("Unknown size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown size: %d", OpSize); break;
}
}
else {
@@ -870,7 +887,7 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(ax, GetSrc<RA_16>(Op->Header.Args[0].ID()));
xor(ax, 0xF);
xor_(ax, 0xF);
movzx(eax, ax);
L(Skip);
mov(GetDst<RA_32>(Node), eax);
@@ -882,7 +899,7 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(eax, GetSrc<RA_32>(Op->Header.Args[0].ID()));
xor(eax, 0x1F);
xor_(eax, 0x1F);
L(Skip);
mov(GetDst<RA_32>(Node), eax);
break;
@@ -893,12 +910,12 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
xor(rax, 0x3F);
xor_(rax, 0x3F);
L(Skip);
mov(GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unknown size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown size: %d", OpSize); break;
}
}
}
@@ -920,7 +937,7 @@ DEF_OP(Rev) {
mov (GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
bswap(GetDst<RA_64>(Node).cvt64());
break;
default: LogMan::Msg::A("Unknown REV size: %d", OpSize); break;
default: LOGMAN_MSG_A("Unknown REV size: %d", OpSize); break;
}
}
@@ -939,15 +956,15 @@ DEF_OP(Bfi) {
mov(Dst, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(TMP2, DestMask);
and(Dst, TMP2);
and_(Dst, TMP2);
mov(TMP2, SourceMask);
and(TMP1, TMP2);
and_(TMP1, TMP2);
shl(TMP1, Op->lsb);
or_(Dst, TMP1);
if (OpSize != 8) {
mov(rcx, uint64_t((1ULL << (OpSize * 8)) - 1));
and(Dst, rcx);
and_(Dst, rcx);
}
}
@@ -955,7 +972,7 @@ DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
uint8_t OpSize = IROp->Size;
LogMan::Throw::A(OpSize <= 8, "OpSize is too large for BFE: %d", OpSize);
LOGMAN_THROW_A(OpSize <= 8, "OpSize is too large for BFE: %d", OpSize);
auto Dst = GetDst<RA_64>(Node);
@@ -987,7 +1004,7 @@ DEF_OP(Bfe) {
if (Op->Width != 64) {
mov(rcx, uint64_t((1ULL << Op->Width) - 1));
and(Dst, rcx);
and_(Dst, rcx);
}
}
@@ -1056,7 +1073,7 @@ DEF_OP(Select) {
if (is_const_true || is_const_false) {
if (is_const_false != true || is_const_true != true || const_true != 1 || const_false != 0) {
LogMan::Msg::A("Select: Unsupported compare inline parameters");
LOGMAN_MSG_A("Select: Unsupported compare inline parameters");
}
(this->*SetCC)(al);
movzx(Dst, al);
@@ -1087,46 +1104,67 @@ DEF_OP(VExtractToGPR) {
pextrq(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
break;
}
default: LogMan::Msg::A("Unknown Element Size: %d", Op->Header.ElementSize); break;
default: LOGMAN_MSG_A("Unknown Element Size: %d", Op->Header.ElementSize); break;
}
}
DEF_OP(Float_ToGPR_ZU) {
LogMan::Msg::D("Unimplemented");
}
DEF_OP(Float_ToGPR_ZS) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_ZS>();
if (Op->Header.ElementSize == 8) {
cvttsd2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
}
else {
cvttss2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
}
}
DEF_OP(Float_ToGPR_U) {
LogMan::Msg::D("Unimplemented");
uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: // int64_t <- float
cvttss2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0808: // int64_t <- double
cvttsd2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0404: // int32_t <- float
cvttss2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0408: // int32_t <- double
cvttsd2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
}
}
DEF_OP(Float_ToGPR_S) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_S>();
if (Op->Header.ElementSize == 8) {
cvtsd2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
}
else {
cvtss2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: // int64_t <- float
cvtss2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0808: // int64_t <- double
cvtsd2si(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0404: // int32_t <- float
cvtss2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
case 0x0408: // int32_t <- double
cvtsd2si(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()));
break;
}
}
DEF_OP(FCmp) {
auto Op = IROp->C<IR::IROp_FCmp>();
if (Op->ElementSize == 4) {
ucomiss(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
if (Op->Flags & (1 << IR::FCMP_FLAG_UNORDERED)) {
if (Op->ElementSize == 4) {
ucomiss(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
}
else {
ucomisd(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
}
}
else {
ucomisd(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
if (Op->ElementSize == 4) {
comiss(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
}
else {
comisd(GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
}
}
mov (rdx, 0);
@@ -1136,32 +1174,34 @@ DEF_OP(FCmp) {
mov(rcx, 0);
setb(cl);
shl(rcx, IR::FCMP_FLAG_LT);
or(rdx, rcx);
or_(rdx, rcx);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_UNORDERED)) {
sahf();
mov(rcx, 0);
setp(cl);
shl(rcx, IR::FCMP_FLAG_UNORDERED);
or(rdx, rcx);
or_(rdx, rcx);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ)) {
sahf();
mov(rcx, 0);
setz(cl);
shl(rcx, IR::FCMP_FLAG_EQ);
or(rdx, rcx);
or_(rdx, rcx);
}
mov (GetDst<RA_64>(Node), rdx);
}
#undef DEF_OP
void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
@@ -1198,9 +1238,7 @@ void JITCore::RegisterALUHandlers() {
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZU, Float_ToGPR_ZU);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_U, Float_ToGPR_U);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
#undef REGISTER_OP
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
@@ -49,7 +55,7 @@ DEF_OP(CASPair) {
mov(Dst.second, rdx);
break;
}
default: LogMan::Msg::A("Unsupported: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported: %d", OpSize);
}
}
@@ -68,7 +74,6 @@ DEF_OP(CAS) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[2].ID());
mov(rdx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
// RCX now contains pointer
@@ -76,31 +81,31 @@ DEF_OP(CAS) {
// RDX contains our desired
lock();
switch (OpSize) {
case 1: {
cmpxchg(byte [MemReg], dl);
movzx(rax, al);
cmpxchg(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), al);
break;
}
case 2: {
cmpxchg(word [MemReg], dx);
movzx(rax, ax);
cmpxchg(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), ax);
break;
}
case 4: {
cmpxchg(dword [MemReg], edx);
cmpxchg(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), eax);
break;
}
case 8: {
cmpxchg(qword [MemReg], rdx);
cmpxchg(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), rax);
break;
}
default: LogMan::Msg::A("Unsupported: %d", OpSize);
default: LOGMAN_MSG_A("Unsupported: %d", OpSize);
}
// RAX now contains the result
mov (GetDst<RA_64>(Node), rax);
}
DEF_OP(AtomicAdd) {
@@ -122,7 +127,7 @@ DEF_OP(AtomicAdd) {
case 8:
add(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -144,7 +149,7 @@ DEF_OP(AtomicSub) {
case 8:
sub(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -155,18 +160,18 @@ DEF_OP(AtomicAnd) {
lock();
switch (Op->Size) {
case 1:
and(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
and(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
and(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
and(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -177,18 +182,18 @@ DEF_OP(AtomicOr) {
lock();
switch (Op->Size) {
case 1:
or(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
or(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
or(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
or(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -199,18 +204,18 @@ DEF_OP(AtomicXor) {
lock();
switch (Op->Size) {
case 1:
xor(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
xor(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
xor(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
xor(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -222,17 +227,17 @@ DEF_OP(AtomicSwap) {
switch (Op->Size) {
case 1:
mov(GetDst<RA_8>(Node), GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Header.Args[1].ID()));
lock();
xchg(byte [MemReg], GetDst<RA_8>(Node));
break;
case 2:
mov(GetDst<RA_16>(Node), GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_16>(Op->Header.Args[1].ID()));
lock();
xchg(word [MemReg], GetDst<RA_16>(Node));
break;
case 4:
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()));
lock();
xchg(dword [MemReg], GetDst<RA_32>(Node));
break;
@@ -241,7 +246,7 @@ DEF_OP(AtomicSwap) {
lock();
xchg(qword [MemReg], GetDst<RA_64>(Node));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicAdd size: %d", Op->Size);
}
}
@@ -251,13 +256,13 @@ DEF_OP(AtomicFetchAdd) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
case 1:
mov(cl, GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_8>(Op->Header.Args[1].ID()));
lock();
xadd(byte [MemReg], cl);
movzx(GetDst<RA_32>(Node), cl);
break;
case 2:
mov(cx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
lock();
xadd(word [MemReg], cx);
movzx(GetDst<RA_32>(Node), cx);
@@ -266,7 +271,7 @@ DEF_OP(AtomicFetchAdd) {
mov(ecx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
lock();
xadd(dword [MemReg], ecx);
mov(GetDst<RA_32>(Node), ecx);
mov(GetDst<RA_64>(Node), ecx);
break;
case 8:
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
@@ -274,7 +279,7 @@ DEF_OP(AtomicFetchAdd) {
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
break;
default: LogMan::Msg::A("Unhandled AtomicFetchAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicFetchAdd size: %d", Op->Size);
}
}
@@ -311,7 +316,7 @@ DEF_OP(AtomicFetchSub) {
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
break;
default: LogMan::Msg::A("Unhandled AtomicFetchAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicFetchAdd size: %d", Op->Size);
}
}
@@ -329,7 +334,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
and(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -345,7 +350,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
and(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -362,7 +367,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
and(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -379,7 +384,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
and(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -389,7 +394,7 @@ DEF_OP(AtomicFetchAnd) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LogMan::Msg::A("Unhandled AtomicFetchAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicFetchAdd size: %d", Op->Size);
}
}
@@ -406,7 +411,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
or(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -422,7 +427,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
or(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -439,7 +444,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
or(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -456,7 +461,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
or(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -466,7 +471,7 @@ DEF_OP(AtomicFetchOr) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LogMan::Msg::A("Unhandled AtomicFetchAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicFetchAdd size: %d", Op->Size);
}
}
@@ -483,7 +488,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
xor(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -499,7 +504,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
xor(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -516,7 +521,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
xor(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -533,7 +538,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
xor(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -543,13 +548,13 @@ DEF_OP(AtomicFetchXor) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LogMan::Msg::A("Unhandled AtomicFetchAdd size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled AtomicFetchAdd size: %d", Op->Size);
}
}
#undef DEF_OP
void JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
@@ -1,11 +1,18 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::D("Unimplemented");
}
@@ -39,7 +46,7 @@ DEF_OP(CallbackReturn) {
sub(dword [rax], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.gregs[X86State::REG_RSP])], 8);
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
// Now jump back to the thunk
// XXX: XMM?
@@ -66,7 +73,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP)) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Label l_BranchHost;
Label l_BranchGuest;
@@ -74,12 +81,12 @@ DEF_OP(ExitFunction) {
jmp(qword[rax]);
L(l_BranchHost);
dq(ExitFunctionLinkerAddress);
dq(ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress);
L(l_BranchGuest);
dq(NewRIP);
} else {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, ThreadState->LookupCache->GetL1Pointer());
mov(rax, RipReg);
@@ -88,14 +95,14 @@ DEF_OP(ExitFunction) {
shl(rax, 4);
Xbyak::RegExp LookupBase = rcx + rax;
cmp(qword[LookupBase + 8], RipReg);
jne(FullLookup);
jmp(qword[LookupBase + 0]);
L(FullLookup);
mov(rax, AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::InternalThreadState, State.State.rip)], RipReg);
mov(rax, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(rax);
}
@@ -227,7 +234,9 @@ DEF_OP(Thunk) {
mov(rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, reinterpret_cast<uintptr_t>(Op->ThunkFnPtr));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
mov(rax, reinterpret_cast<uintptr_t>(thunkFn));
call(rax);
if (NumPush & 1)
@@ -244,7 +253,7 @@ DEF_OP(ValidateCode) {
int idx = 0;
xor_(GetDst<RA_64>(Node), GetDst<RA_64>(Node));
mov(rax, Op->CodePtr);
mov(rax, Entry + Op->Offset);
mov(rbx, 1);
while (len >= 4) {
cmp(dword[rax + idx], *(uint32_t*)(OldCode + idx));
@@ -268,8 +277,6 @@ DEF_OP(ValidateCode) {
}
DEF_OP(RemoveCodeEntry) {
auto Op = IROp->C<IR::IROp_RemoveCodeEntry>();
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -279,11 +286,11 @@ DEF_OP(RemoveCodeEntry) {
sub(rsp, 8); // Align
mov(rdi, STATE);
mov(rax, Op->RIP); // imm64 move
mov(rax, Entry); // imm64 move
mov(rsi, rax);
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntry));
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
call(rax);
if (NumPush & 1)
@@ -296,7 +303,7 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function);
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function, uint32_t Leaf);
union {
ClassPtrType ClassPtr;
uint64_t Raw;
@@ -313,6 +320,7 @@ DEF_OP(CPUID) {
// Result: RAX, RDX. 4xi32
mov (rsi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov (rdx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov (rdi, reinterpret_cast<uint64_t>(&CTX->CPUID));
auto NumPush = RA64.size();
@@ -338,8 +346,8 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -25,7 +31,7 @@ DEF_OP(VInsGPR) {
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[1].ID()), Op->Index);
break;
}
default: LogMan::Msg::A("Unknown Element Size: %d", Op->Header.ElementSize); break;
default: LOGMAN_MSG_A("Unknown Element Size: %d", Op->Header.ElementSize); break;
}
}
@@ -46,14 +52,10 @@ DEF_OP(VCastFromGPR) {
case 8:
vmovq(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()).cvt64());
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Float_FromGPR_U) {
LogMan::Msg::A("Unimplemented");
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
@@ -89,14 +91,10 @@ DEF_OP(Float_FToF) {
cvtsd2ss(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
}
default: LogMan::Msg::A("Unknown FCVT sizes: 0x%x", Conv);
default: LOGMAN_MSG_A("Unknown FCVT sizes: 0x%x", Conv);
}
}
DEF_OP(Vector_UToF) {
LogMan::Msg::A("Unimplemented");
}
DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
switch (Op->Header.ElementSize) {
@@ -115,14 +113,10 @@ DEF_OP(Vector_SToF) {
cvtsi2sd(xmm15, rax);
movlhps(GetDst(Node), xmm15);
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Vector_FToZU) {
LogMan::Msg::A("Unimplemented");
}
DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
switch (Op->Header.ElementSize) {
@@ -132,14 +126,10 @@ DEF_OP(Vector_FToZS) {
case 8:
cvttpd2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
DEF_OP(Vector_FToU) {
LogMan::Msg::A("Unimplemented");
}
DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
switch (Op->Header.ElementSize) {
@@ -149,7 +139,7 @@ DEF_OP(Vector_FToS) {
case 8:
cvtpd2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
default: LogMan::Msg::A("Unknown castGPR element size: %d", Op->Header.ElementSize);
default: LOGMAN_MSG_A("Unknown castGPR element size: %d", Op->Header.ElementSize);
}
}
@@ -166,25 +156,54 @@ DEF_OP(Vector_FToF) {
cvtpd2ps(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
}
default: LogMan::Msg::A("Unknown Conversion Type : 0%04x", Conv); break;
default: LOGMAN_MSG_A("Unknown Conversion Type : 0%04x", Conv); break;
}
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
uint8_t RoundMode{};
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
RoundMode = 0b0000'0'0'00;
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
RoundMode = 0b0000'0'0'01;
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
RoundMode = 0b0000'0'0'10;
break;
case FEXCore::IR::Round_Towards_Zero.Val:
RoundMode = 0b0000'0'0'11;
break;
case FEXCore::IR::Round_Host.Val:
RoundMode = 0b0000'0'1'00;
break;
}
switch (Op->Header.ElementSize) {
case 4:
roundps(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), RoundMode);
break;
case 8:
roundpd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), RoundMode);
break;
}
}
#undef DEF_OP
void JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_U, Float_FromGPR_U);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_UTOF, Vector_UToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZU, Vector_FToZU);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOU, Vector_FToU);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
@@ -1,8 +1,14 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -35,8 +41,8 @@ DEF_OP(AESKeyGenAssist) {
}
#undef DEF_OP
void JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
@@ -1,21 +1,27 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
shr(rax, Op->Flag);
and(rax, 1);
and_(rax, 1);
mov(GetDst<RA_64>(Node), rax);
}
#undef DEF_OP
void JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
File diff suppressed because it is too large. Load diff
@@ -1 +0,0 @@
+48 -51
View File
@@ -1,11 +1,18 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#pragma once
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JIT.h"
#include "Common/MathUtils.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
@@ -51,15 +58,15 @@ namespace FEXCore::CPU {
using namespace Xbyak::util;
const std::array<Xbyak::Reg, 9> RA64 = { rsi, r8, r9, r10, r11, rbp, r12, r13, r15 };
const std::array<std::pair<Xbyak::Reg, Xbyak::Reg>, 4> RA64Pair = {{ {rsi, r8}, {r9, r10}, {r11, rbp}, {r12, r13} }};
const std::array<Xbyak::Reg, 11> RAXMM = { xmm0, xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10 };
const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm0, xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10 };
const std::array<Xbyak::Reg, 11> RAXMM = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
class JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
~JITCore() override;
explicit X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
~X86JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -69,17 +76,15 @@ public:
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
bool HandleSIGILL(int Signal, void *info, void *ucontext);
bool HandleSignalPause(int Signal, void *info, void *ucontext);
bool HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView<true> const *IR;
FEXCore::IR::IRListView const *IR;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
uint64_t Entry;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, Label> JumpTargets;
Xbyak::util::Cpu Features{};
@@ -107,27 +112,27 @@ private:
constexpr static uint8_t RA_64 = 3;
constexpr static uint8_t RA_XMM = 4;
IR::PhysicalRegister GetPhys(uint32_t Node);
IR::PhysicalRegister GetPhys(uint32_t Node) const;
bool IsFPR(uint32_t Node);
bool IsGPR(uint32_t Node);
bool IsFPR(uint32_t Node) const;
bool IsGPR(uint32_t Node) const;
template<uint8_t RAType>
Xbyak::Reg GetSrc(uint32_t Node);
Xbyak::Reg GetSrc(uint32_t Node) const;
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node);
std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node) const;
template<uint8_t RAType>
Xbyak::Reg GetDst(uint32_t Node);
Xbyak::Reg GetDst(uint32_t Node) const;
Xbyak::Xmm GetSrc(uint32_t Node);
Xbyak::Xmm GetDst(uint32_t Node);
Xbyak::Xmm GetSrc(uint32_t Node) const;
Xbyak::Xmm GetDst(uint32_t Node) const;
Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
void CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread);
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
@@ -135,13 +140,11 @@ private:
bool GetSamplingData {true};
#endif
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 1;
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
static uint64_t ExitFunctionLink(JITCore* code, FEXCore::Core::InternalThreadState *Thread, uint64_t *record);
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
@@ -152,41 +155,26 @@ private:
// For code safety we can't delete code buffers until outside of all signals
std::vector<CodeBuffer> CodeBuffers{};
// This is the codebuffer that our dispatcher lives in
CodeBuffer DispatcherCodeBuffer{};
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t ThreadStopHandlerAddress{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t PauseReturnInstruction{};
uint32_t SignalHandlerRefCounter{};
struct CompilerSharedData {
void *InterpreterFallbackHelperAddress;
uint64_t SignalHandlerReturnAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
void StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
uint32_t SpillSlots{};
using SetCC = void (JITCore::*)(const Operand& op);
using CMovCC = void (JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (JITCore::*)(const Label& label, LabelType type);
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (X86JITCore::*)(const Label& label, LabelType type);
std::tuple<SetCC, CMovCC, JCC> GetCC(IR::CondClassType cond);
using OpHandler = void (JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
using OpHandler = void (X86JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -198,6 +186,9 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
void PushRegs();
void PopRegs();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
///< Unhandled handler
@@ -209,7 +200,9 @@ private:
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
DEF_OP(CycleCounter);
DEF_OP(Add);
DEF_OP(Sub);
@@ -248,9 +241,7 @@ private:
DEF_OP(Sbfe);
DEF_OP(Select);
DEF_OP(VExtractToGPR);
DEF_OP(Float_ToGPR_ZU);
DEF_OP(Float_ToGPR_ZS);
DEF_OP(Float_ToGPR_U);
DEF_OP(Float_ToGPR_S);
DEF_OP(FCmp);
DEF_OP(F80Cmp);
@@ -288,16 +279,14 @@ private:
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(Float_FromGPR_U);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_UToF);
DEF_OP(Vector_SToF);
DEF_OP(Vector_FToZU);
DEF_OP(Vector_FToZS);
DEF_OP(Vector_FToU);
DEF_OP(Vector_FToS);
DEF_OP(Vector_FToF);
DEF_OP(Vector_FToI);
///< Flag ops
DEF_OP(GetHostFlag);
@@ -315,6 +304,7 @@ private:
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
///< Misc ops
DEF_OP(EndBlock);
@@ -339,6 +329,7 @@ private:
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
DEF_OP(VOr);
DEF_OP(VXor);
DEF_OP(VAdd);
@@ -349,8 +340,10 @@ private:
DEF_OP(VSQSub);
DEF_OP(VAddP);
DEF_OP(VAddV);
DEF_OP(VUMinV);
DEF_OP(VURAvg);
DEF_OP(VAbs);
DEF_OP(VPopcount);
DEF_OP(VFAdd);
DEF_OP(VFAddP);
DEF_OP(VFSub);
@@ -370,6 +363,8 @@ private:
DEF_OP(VSMax);
DEF_OP(VZip);
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -392,6 +387,7 @@ private:
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
@@ -414,6 +410,7 @@ private:
DEF_OP(VSMull);
DEF_OP(VUMull2);
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
///< Encryption ops
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -5,7 +11,7 @@
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -30,10 +36,10 @@ DEF_OP(LoadContext) {
}
break;
case 16: {
LogMan::Msg::A("Invalid GPR load of size 16");
LOGMAN_MSG_A("Invalid GPR load of size 16");
}
break;
default: LogMan::Msg::A("Unhandled LoadContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled LoadContext size: %d", OpSize);
}
}
else {
@@ -63,7 +69,7 @@ DEF_OP(LoadContext) {
movups(GetDst(Node), xword [STATE + Op->Offset]);
}
break;
default: LogMan::Msg::A("Unhandled LoadContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled LoadContext size: %d", OpSize);
}
}
}
@@ -94,7 +100,7 @@ DEF_OP(StoreContext) {
case 16:
LogMan::Msg::D("Invalid store size of 16");
break;
default: LogMan::Msg::A("Unhandled StoreContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled StoreContext size: %d", OpSize);
}
}
else {
@@ -123,7 +129,7 @@ DEF_OP(StoreContext) {
movups(xword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
}
break;
default: LogMan::Msg::A("Unhandled StoreContext size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled StoreContext size: %d", OpSize);
}
}
}
@@ -154,15 +160,15 @@ DEF_OP(LoadContextIndexed) {
mov(GetDst<RA_64>(Node), qword [rax + index * Op->Stride]);
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
case 16:
LogMan::Msg::A("Invalid Class load of size 16");
LOGMAN_MSG_A("Invalid Class load of size 16");
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
@@ -189,7 +195,7 @@ DEF_OP(LoadContextIndexed) {
vmovq(GetDst(Node), qword [rax + index * Op->Stride]);
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
@@ -217,12 +223,12 @@ DEF_OP(LoadContextIndexed) {
movups(GetDst(Node), xword [STATE + rax]);
break;
default:
LogMan::Msg::A("Unhandled LoadContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled LoadContextIndexed size: %d", Op->Size);
}
break;
}
default:
LogMan::Msg::A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled LoadContextIndexed stride: %d", Op->Stride);
}
}
}
@@ -242,13 +248,13 @@ DEF_OP(StoreContextIndexed) {
case 4:
case 8: {
if (!(size == 1 || size == 2 || size == 4 || size == 8)) {
LogMan::Msg::A("Unhandled StoreContextIndexed size: %d", Op->Size);
LOGMAN_MSG_A("Unhandled StoreContextIndexed size: %d", Op->Size);
}
mov(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value);
break;
}
default:
LogMan::Msg::A("Unhandled StoreContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled StoreContextIndexed stride: %d", Op->Stride);
}
}
else {
@@ -273,7 +279,7 @@ DEF_OP(StoreContextIndexed) {
vmovq(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value);
break;
default:
LogMan::Msg::A("Unhandled StoreContextIndexed size: %d", size);
LOGMAN_MSG_A("Unhandled StoreContextIndexed size: %d", size);
}
break;
}
@@ -301,12 +307,12 @@ DEF_OP(StoreContextIndexed) {
movups(xword [STATE + rax], value);
break;
default:
LogMan::Msg::A("Unhandled StoreContextIndexed size: %d", size);
LOGMAN_MSG_A("Unhandled StoreContextIndexed size: %d", size);
}
break;
}
default:
LogMan::Msg::A("Unhandled StoreContextIndexed stride: %d", Op->Stride);
LOGMAN_MSG_A("Unhandled StoreContextIndexed stride: %d", Op->Stride);
}
}
}
@@ -334,7 +340,7 @@ DEF_OP(SpillRegister) {
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
@@ -350,10 +356,10 @@ DEF_OP(SpillRegister) {
movaps(xword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
break;
}
default: LogMan::Msg::A("Unhandled SpillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled SpillRegister size: %d", OpSize);
}
} else {
LogMan::Msg::A("Unhandled SpillRegister class: %d", Op->Class.Val);
LOGMAN_MSG_A("Unhandled SpillRegister class: %d", Op->Class.Val);
}
@@ -382,7 +388,7 @@ DEF_OP(FillRegister) {
mov(GetDst<RA_64>(Node), qword [rsp + SlotOffset]);
break;
}
default: LogMan::Msg::A("Unhandled FillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled FillRegister size: %d", OpSize);
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
@@ -398,10 +404,10 @@ DEF_OP(FillRegister) {
movaps(GetDst(Node), xword [rsp + SlotOffset]);
break;
}
default: LogMan::Msg::A("Unhandled FillRegister size: %d", OpSize);
default: LOGMAN_MSG_A("Unhandled FillRegister size: %d", OpSize);
}
} else {
LogMan::Msg::A("Unhandled FillRegister class: %d", Op->Class.Val);
LOGMAN_MSG_A("Unhandled FillRegister class: %d", Op->Class.Val);
}
}
@@ -419,16 +425,16 @@ DEF_OP(StoreFlag) {
mov(byte [STATE + (offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag)], al);
}
Xbyak::RegExp JITCore::GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
Xbyak::RegExp X86JITCore::GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) const {
if (Offset.IsInvalid()) {
return Base;
} else {
if (OffsetScale != 1 && OffsetScale != 2 && OffsetScale != 4 && OffsetScale != 8) {
LogMan::Msg::A("Unhandled GenerateModRM OffsetScale: %d", OffsetScale);
LOGMAN_MSG_A("Unhandled GenerateModRM OffsetScale: %d", OffsetScale);
}
if (OffsetType != IR::MEM_OFFSET_SXTX) {
LogMan::Msg::A("Unhandled GenerateModRM OffsetType: %d", OffsetType.Val);
LOGMAN_MSG_A("Unhandled GenerateModRM OffsetType: %d", OffsetType.Val);
}
uint64_t Const;
@@ -469,7 +475,7 @@ DEF_OP(LoadMem) {
mov(Dst, qword [MemPtr]);
}
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
}
else
@@ -505,7 +511,7 @@ DEF_OP(LoadMem) {
}
}
break;
default: LogMan::Msg::A("Unhandled LoadMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled LoadMem size: %d", Op->Size);
}
}
}
@@ -531,7 +537,7 @@ DEF_OP(StoreMem) {
case 8:
mov(qword [MemPtr], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
}
else {
@@ -554,22 +560,30 @@ DEF_OP(StoreMem) {
else
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled StoreMem size: %d", Op->Size);
default: LOGMAN_MSG_A("Unhandled StoreMem size: %d", Op->Size);
}
}
}
DEF_OP(VLoadMemElement) {
LogMan::Msg::A("Unimplemented");
LOGMAN_MSG_A("Unimplemented");
}
DEF_OP(VStoreMemElement) {
LogMan::Msg::A("Unimplemented");
LOGMAN_MSG_A("Unimplemented");
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
clflush(ptr [MemReg]);
}
#undef DEF_OP
void JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, Unhandled); // SRA specific, not supported on this backend
@@ -586,6 +600,7 @@ void JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
#undef REGISTER_OP
}
}
+43 -26
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -6,7 +12,11 @@ static void PrintValue(uint64_t Value) {
LogMan::Msg::D("Value: 0x%lx", Value);
}
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::D("Value: 0x%016lx'%016lx", ValueUpper, Value);
}
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -20,7 +30,7 @@ DEF_OP(Fence) {
case IR::Fence_Store.Val:
sfence();
break;
default: LogMan::Msg::A("Unknown Fence: %d", Op->Fence); break;
default: LOGMAN_MSG_A("Unknown Fence: %d", Op->Fence); break;
}
}
@@ -31,13 +41,22 @@ DEF_OP(Break) {
case 5: // Guest ud2
ud2();
break;
case 1: // Int <imm8>
ud2();
break;
case 2: // overflow
ud2();
break;
case 3: // int 1
ud2();
break;
case 4: { // HLT
// Time to quit
// Set our stack to the starting stack location
mov(rsp, qword [STATE + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)]);
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadStopHandlerAddress);
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
break;
}
@@ -48,23 +67,23 @@ DEF_OP(Break) {
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
// This jump target needs to be a constant offset here
mov(TMP1, ThreadPauseHandlerAddress);
mov(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddress);
jmp(TMP1);
}
else {
// If we don't have a gdb server attached then....crash?
// Treat this case like HLT
mov(rsp, qword [STATE + offsetof(FEXCore::Core::ThreadState, ReturningStackLocation)]);
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadStopHandlerAddress);
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
}
break;
}
default: LogMan::Msg::A("Unknown Break reason: %d", Op->Reason);
default: LOGMAN_MSG_A("Unknown Break reason: %d", Op->Reason);
}
}
@@ -89,10 +108,10 @@ DEF_OP(SetRoundingMode) {
mov(TMP1.cvt32(), dword [rsp]);
// Insert the new rounding mode
and(TMP1.cvt32(), ~(0b111 << 13));
and_(TMP1.cvt32(), ~(0b111 << 13));
mov(TMP2.cvt32(), Src);
shl(TMP2.cvt32(), 13);
or(TMP1.cvt32(), TMP2.cvt32());
or_(TMP1.cvt32(), TMP2.cvt32());
// Store it to mxcsr
// Only loads from memory
@@ -104,29 +123,27 @@ DEF_OP(SetRoundingMode) {
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
for (auto &Reg : RA64)
push(Reg);
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
mov(rax, reinterpret_cast<uintptr_t>(PrintValue));
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, reinterpret_cast<uintptr_t>(PrintValue));
mov(rax, reinterpret_cast<uintptr_t>(PrintVectorValue));
}
call(rax);
if (NumPush & 1)
add(rsp, 8); // Align
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
PopRegs();
}
#undef DEF_OP
void JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
@@ -1,9 +1,15 @@
/*
$info$
tags: backend|x86-64
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -19,7 +25,7 @@ DEF_OP(ExtractElementPair) {
mov (GetDst<RA_64>(Node), Regs[Op->Element]);
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
default: LOGMAN_MSG_A("Unknown Size"); break;
}
}
@@ -28,21 +34,24 @@ DEF_OP(CreateElementPair) {
std::pair<Xbyak::Reg, Xbyak::Reg> Dst;
Xbyak::Reg RegFirst;
Xbyak::Reg RegSecond;
Xbyak::Reg RegTmp;
switch (Op->Header.Size) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetSrc<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_32>(Op->Header.Args[1].ID());
RegTmp = eax;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetSrc<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_64>(Op->Header.Args[1].ID());
RegTmp = rax;
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
default: LOGMAN_MSG_A("Unknown Size"); break;
}
if (Dst.first != RegSecond) {
@@ -52,7 +61,9 @@ DEF_OP(CreateElementPair) {
mov(Dst.second, RegSecond);
mov(Dst.first, RegFirst);
} else {
LogMan::Msg::A("Unhandled CreateElementPair");
mov(RegTmp, RegFirst);
mov(Dst.second, RegSecond);
mov(Dst.first, RegTmp);
}
}
@@ -62,8 +73,8 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
void X86JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
File diff suppressed because it is too large. Load diff
+17 -8
View File
@@ -1,6 +1,15 @@
/*
$info$
tags: glue|block-database
desc: Stores information about blocks, and provides C++ implementations to lookup the blocks
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/LookupCache.h"
#include <FEXCore/Utils/Allocator.h>
#include <sys/mman.h>
namespace FEXCore {
@@ -19,27 +28,27 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// Allocate a region of memory that we can use to back our block pointers
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(mmap(nullptr, ctx->Config.VirtualMemSize / 4096 * 8, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, ctx->Config.VirtualMemSize / 4096 * 8, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
// Allocate our memory backing our pages
// We need 32KB per guest page (One pointer per byte)
// XXX: We can drop down to 16KB if we store 4byte offsets from the code base
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = reinterpret_cast<uintptr_t>(mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LogMan::Throw::A(PageMemory != -1ULL, "Failed to allocate page memory");
PageMemory = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = reinterpret_cast<uintptr_t>(mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LogMan::Throw::A(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
L1Pointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
LookupCache::~LookupCache() {
munmap(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8);
munmap(reinterpret_cast<void*>(PageMemory), CODE_SIZE);
munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8);
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PageMemory), CODE_SIZE);
FEXCore::Allocator::munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
}
void LookupCache::HintUsedRange(uint64_t Address, uint64_t Size) {
+19 -9
View File
@@ -35,11 +35,21 @@ public:
}
}
void AddBlockMapping(uint64_t Address, void *HostCode) {
auto InsertPoint = BlockList.emplace(Address, (uintptr_t)HostCode);
LogMan::Throw::A(InsertPoint.second == true, "Dupplicate block mapping added");
std::map<uint64_t, std::vector<uint64_t>> CodePages;
// no need to update L1 or L2, they will get updated on first lookup
void AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
auto InsertPoint = BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A(InsertPoint.second == true, "Dupplicate block mapping added");
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
CodePages[CurrentPage].push_back(Address);
}
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
}
void Erase(uint64_t Address) {
@@ -88,8 +98,8 @@ public:
void HintUsedRange(uint64_t Address, uint64_t Size);
uintptr_t GetL1Pointer() { return L1Pointer; }
uintptr_t GetPagePointer() { return PagePointer; }
uintptr_t GetL1Pointer() const { return L1Pointer; }
uintptr_t GetPagePointer() const { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
@@ -99,9 +109,8 @@ private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = L1Entry.HostCode = 0;
}
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
// Do ful map
auto FullAddress = Address;
@@ -197,6 +206,7 @@ private:
}
};
std::map<BlockLinkTag, std::function<void()>> BlockLinks;
std::map<uint64_t, uint64_t> BlockList;
File diff suppressed because it is too large. Load diff
+45 -33
View File
@@ -49,7 +49,7 @@ public:
OrderedNode* GetNewJumpBlock(uint64_t RIP) {
auto it = JumpTargets.find(RIP);
LogMan::Throw::A(it != JumpTargets.end(), "Couldn't find block generated for 0x%lx", RIP);
LOGMAN_THROW_A(it != JumpTargets.end(), "Couldn't find block generated for 0x%lx", RIP);
return it->second.BlockEntry;
}
@@ -59,7 +59,7 @@ public:
it->second.HaveEmitted = true;
if (CurrentCodeBlock->Wrapped(ListData.Begin()).ID() == it->second.BlockEntry->Wrapped(ListData.Begin()).ID()) return;
if (CurrentCodeBlock->Wrapped(DualListData.ListBegin()).ID() == it->second.BlockEntry->Wrapped(DualListData.ListBegin()).ID()) return;
// We have hit a RIP that is a jump target
// Thus we need to end up in a new block
@@ -81,14 +81,15 @@ public:
// rdi, 0x8
// cmp qword [rdi-8], 0
// jne .label
if (!BlockSetRIP) {
if (LastOp && !BlockSetRIP) {
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end() && LastOp) {
if (it == JumpTargets.end()) {
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
// If we don't have a jump target to a new block then we have to leave
// Set the RIP to the next instruction and leave
_ExitFunction(_Constant(GPRSize * 8, NextRIP));
auto RelocatedNextRIP = _EntrypointOffset(NextRIP - Entry, GPRSize);
_ExitFunction(RelocatedNextRIP);
}
else if (it != JumpTargets.end()) {
_Jump(it->second.BlockEntry);
@@ -103,7 +104,8 @@ public:
OpDispatchBuilder(FEXCore::Context::Context *ctx);
void ResetWorkingList();
bool HadDecodeFailure() { return DecodeFailure; }
void ResetDecodeFailure() { DecodeFailure = false; }
bool HadDecodeFailure() const { return DecodeFailure; }
void BeginFunction(uint64_t RIP, std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
void Finalize();
@@ -258,12 +260,6 @@ public:
template<size_t ElementSize>
void PSUBQOp(OpcodeArgs);
template<size_t ElementSize>
void PMINUOp(OpcodeArgs);
template<size_t ElementSize>
void PMAXUOp(OpcodeArgs);
void PMINSWOp(OpcodeArgs);
void PMAXSWOp(OpcodeArgs);
template<size_t ElementSize>
void MOVMSKOp(OpcodeArgs);
void MOVMSKOpOne(OpcodeArgs);
template<size_t ElementSize>
@@ -273,10 +269,6 @@ public:
void PSHUFBOp(OpcodeArgs);
template<size_t ElementSize, bool HalfSize, bool Low>
void PSHUFDOp(OpcodeArgs);
template<size_t ElementSize>
void PCMPEQOp(OpcodeArgs);
template<size_t ElementSize>
void PCMPGTOp(OpcodeArgs);
void MOVDOp(OpcodeArgs);
template<size_t ElementSize, bool Scalar, uint32_t SrcIndex>
void PSRLDOp(OpcodeArgs);
@@ -295,21 +287,21 @@ public:
template<size_t ElementSize>
void PAVGOp(OpcodeArgs);
void MOVDDUPOp(OpcodeArgs);
template<size_t DstElementSize, bool Signed>
template<size_t DstElementSize>
void CVTGPR_To_FPR(OpcodeArgs);
template<size_t SrcElementSize, bool Signed, bool HostRoundingMode>
template<size_t SrcElementSize, bool HostRoundingMode>
void CVTFPR_To_GPR(OpcodeArgs);
template<size_t SrcElementSize, bool Signed, bool Widen>
template<size_t SrcElementSize, bool Widen>
void Vector_CVT_Int_To_Float(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void Scalar_CVT_Float_To_Float(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void Vector_CVT_Float_To_Float(OpcodeArgs);
template<size_t SrcElementSize, bool Signed, bool Narrow, bool HostRoundingMode>
template<size_t SrcElementSize, bool Narrow, bool HostRoundingMode>
void Vector_CVT_Float_To_Int(OpcodeArgs);
template<size_t SrcElementSize, bool Signed, bool Widen>
void MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<size_t SrcElementSize, bool Signed, bool Narrow, bool HostRoundingMode>
template<size_t SrcElementSize, bool Narrow, bool HostRoundingMode>
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs);
void MASKMOVOp(OpcodeArgs);
void MOVBetweenGPR_FPR(OpcodeArgs);
@@ -323,17 +315,13 @@ public:
void ANDNOp(OpcodeArgs);
template<size_t ElementSize>
void PINSROp(OpcodeArgs);
void InsertPSOp(OpcodeArgs);
template<size_t ElementSize>
void PExtrOp(OpcodeArgs);
template<size_t ElementSize, bool Signed>
void PMULOp(OpcodeArgs);
template<size_t ElementSize>
void PSIGN(OpcodeArgs);
template<size_t ElementSize>
void PABS(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -349,6 +337,8 @@ public:
void FST(OpcodeArgs);
void FST(OpcodeArgs);
template<bool Truncate>
void FIST(OpcodeArgs);
enum class OpResult {
@@ -381,6 +371,7 @@ public:
void X87TAN(OpcodeArgs);
void X87ATAN(OpcodeArgs);
void X87LDENV(OpcodeArgs);
void X87FLDCW(OpcodeArgs);
void X87FNSTENV(OpcodeArgs);
void X87FSTCW(OpcodeArgs);
void X87LDSW(OpcodeArgs);
@@ -454,6 +445,8 @@ public:
template<uint8_t FenceType>
void FenceOp(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void PSADBW(OpcodeArgs);
void AESImcOp(OpcodeArgs);
@@ -463,6 +456,23 @@ public:
void AESDecLastOp(OpcodeArgs);
void AESKeyGenAssist(OpcodeArgs);
template<size_t ElementSize, size_t DstElementSize, bool Signed>
void ExtendVectorElements(OpcodeArgs);
template<size_t ElementSize, bool Scalar>
void VectorRound(OpcodeArgs);
template<size_t ElementSize>
void VectorBlend(OpcodeArgs);
template<size_t ElementSize>
void VectorVariableBlend(OpcodeArgs);
void PTestOp(OpcodeArgs);
void PHMINPOSUWOp(OpcodeArgs);
template<size_t ElementSize>
void DPPOp(OpcodeArgs);
void MPSADBWOp(OpcodeArgs);
void UnimplementedOp(OpcodeArgs);
#undef OpcodeArgs
@@ -471,22 +481,24 @@ public:
OrderedNode *GetPackedRFLAG(bool Lower8);
void SetMultiblock(bool _Multiblock) { Multiblock = _Multiblock; }
bool GetMultiblock() { return Multiblock; }
bool HandledLock = false;
private:
bool DecodeFailure{false};
FEXCore::IR::IROp_IRHeader *Current_Header{};
OrderedNode *Current_HeaderNode{};
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align);
uint8_t GetDstSize(FEXCore::X86Tables::DecodedOp Op);
uint8_t GetSrcSize(FEXCore::X86Tables::DecodedOp Op);
uint8_t GetDstSize(FEXCore::X86Tables::DecodedOp Op) const;
uint8_t GetSrcSize(FEXCore::X86Tables::DecodedOp Op) const;
template<unsigned BitOffset>
void SetRFLAG(OrderedNode *Value);
@@ -516,12 +528,12 @@ private:
OrderedNode * GetX87Top();
void SetX87Top(OrderedNode *Value);
bool DestIsLockedMem(FEXCore::X86Tables::DecodedOp Op) {
return Op->Dest.TypeNone.Type !=FEXCore::X86Tables::DecodedOperand::TYPE_GPR && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK);
bool DestIsLockedMem(FEXCore::X86Tables::DecodedOp Op) const {
return DestIsMem(Op) && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK) != 0;
}
bool DestIsMem(FEXCore::X86Tables::DecodedOp Op) {
return Op->Dest.TypeNone.Type !=FEXCore::X86Tables::DecodedOperand::TYPE_GPR;
bool DestIsMem(FEXCore::X86Tables::DecodedOp Op) const {
return !Op->Dest.IsGPR();
}
void CreateJumpBlocks(std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
+13 -5
View File
@@ -1,4 +1,12 @@
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include <FEXCore/Utils/Allocator.h>
#include <cstring>
#include <stdlib.h>
@@ -24,15 +32,15 @@ X86GeneratedCode::X86GeneratedCode() {
}
X86GeneratedCode::~X86GeneratedCode() {
munmap(CodePtr, CODE_SIZE);
FEXCore::Allocator::munmap(CodePtr, CODE_SIZE);
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
FEXCore::Config::Value<bool> Is64BitMode{FEXCore::Config::CONFIG_IS64BIT_MODE, 0};
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
return mmap(nullptr, Size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
return FEXCore::Allocator::mmap(nullptr, Size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
}
// First 64bit page
@@ -42,14 +50,14 @@ void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void *Ptr = mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
void *Ptr = FEXCore::Allocator::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED &&
reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
munmap(Ptr, Size);
FEXCore::Allocator::munmap(Ptr, Size);
continue;
}
+6
View File
@@ -1,3 +1,9 @@
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <FEXCore/Config/Config.h>
+7
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: frontend|x86-tables ~ Metadata that drives the frontend x86/64 decoding
tags: frontend|x86-tables
$end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/Context.h>
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
@@ -279,13 +285,13 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(BaseOps, BaseOpTable, sizeof(BaseOpTable) / sizeof(BaseOpTable[0]));
GenerateTable(BaseOps, BaseOpTable, std::size(BaseOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(BaseOps, BaseOpTable_64, sizeof(BaseOpTable_64) / sizeof(BaseOpTable_64[0]));
GenerateTable(BaseOps, BaseOpTable_64, std::size(BaseOpTable_64));
}
else {
GenerateTable(BaseOps, BaseOpTable_32, sizeof(BaseOpTable_32) / sizeof(BaseOpTable_32[0]));
GenerateTable(BaseOps, BaseOpTable_32, std::size(BaseOpTable_32));
}
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -38,6 +44,6 @@ void InitializeDDDTables() {
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
};
GenerateTable(DDDNowOps, DDDNowOpTable, sizeof(DDDNowOpTable) / sizeof(DDDNowOpTable[0]));
GenerateTable(DDDNowOps, DDDNowOpTable, std::size(DDDNowOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -20,6 +26,6 @@ void InitializeEVEXTables() {
{0xE7, 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
};
GenerateTable(EVEXTableOps, EVEXTable, sizeof(EVEXTable) / sizeof(EVEXTable[0]));
GenerateTable(EVEXTableOps, EVEXTable, std::size(EVEXTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -35,10 +41,10 @@ void InitializeH0F38Tables() {
{OPD(PF_38_NONE, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x10), 1, X86InstInfo{"PBLENDVB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x14), 1, X86InstInfo{"BLENDVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x15), 1, X86InstInfo{"BLENDVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x17), 1, X86InstInfo{"PTEST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x10), 1, X86InstInfo{"PBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x14), 1, X86InstInfo{"BLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x15), 1, X86InstInfo{"BLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x17), 1, X86InstInfo{"PTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x1D), 1, X86InstInfo{"PABSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -46,34 +52,34 @@ void InitializeH0F38Tables() {
{OPD(PF_38_NONE, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x20), 1, X86InstInfo{"PMOVSXBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x21), 1, X86InstInfo{"PMOVSXBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x22), 1, X86InstInfo{"PMOVSXBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x23), 1, X86InstInfo{"PMOVSXWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x24), 1, X86InstInfo{"PMOVSXWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x25), 1, X86InstInfo{"PMOVSXDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x28), 1, X86InstInfo{"PMULDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x29), 1, X86InstInfo{"PCMPEQQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x2B), 1, X86InstInfo{"PACKUSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x20), 1, X86InstInfo{"PMOVSXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x21), 1, X86InstInfo{"PMOVSXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x22), 1, X86InstInfo{"PMOVSXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x23), 1, X86InstInfo{"PMOVSXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x24), 1, X86InstInfo{"PMOVSXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x25), 1, X86InstInfo{"PMOVSXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x28), 1, X86InstInfo{"PMULDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x29), 1, X86InstInfo{"PCMPEQQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x2B), 1, X86InstInfo{"PACKUSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x30), 1, X86InstInfo{"PMOVZXBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x31), 1, X86InstInfo{"PMOVZXBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x32), 1, X86InstInfo{"PMOVZXBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x3B), 1, X86InstInfo{"PMINUD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3C), 1, X86InstInfo{"PMAXSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x3D), 1, X86InstInfo{"PMAXSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x3E), 1, X86InstInfo{"PMAXUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x3F), 1, X86InstInfo{"PMAXUD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x30), 1, X86InstInfo{"PMOVZXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x31), 1, X86InstInfo{"PMOVZXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x32), 1, X86InstInfo{"PMOVZXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3B), 1, X86InstInfo{"PMINUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3C), 1, X86InstInfo{"PMAXSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3D), 1, X86InstInfo{"PMAXSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3E), 1, X86InstInfo{"PMAXUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3F), 1, X86InstInfo{"PMAXUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDB), 1, X86InstInfo{"AESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDC), 1, X86InstInfo{"AESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -92,6 +98,6 @@ void InitializeH0F38Tables() {
};
#undef OPD
GenerateTable(H0F38TableOps, H0F38Table, sizeof(H0F38Table) / sizeof(H0F38Table[0]));
GenerateTable(H0F38TableOps, H0F38Table, std::size(H0F38Table));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -10,26 +16,26 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
const U16U8InfoStruct H0F3ATable[] = {
{OPD(0, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_8BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_8BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(0, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(0, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(0, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -48,10 +54,10 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(H0F3ATableOps, H0F3ATable, sizeof(H0F3ATable) / sizeof(H0F3ATable[0]));
GenerateTable(H0F3ATableOps, H0F3ATable, std::size(H0F3ATable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(H0F3ATableOps, H0F3ATable_64, sizeof(H0F3ATable_64) / sizeof(H0F3ATable_64[0]));
GenerateTable(H0F3ATableOps, H0F3ATable_64, std::size(H0F3ATable_64));
}
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -152,12 +158,12 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable, sizeof(PrimaryGroupOpTable) / sizeof(PrimaryGroupOpTable[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_64, sizeof(PrimaryGroupOpTable_64) / sizeof(PrimaryGroupOpTable_64[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
}
else {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_32, sizeof(PrimaryGroupOpTable_32) / sizeof(PrimaryGroupOpTable_32[0]));
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_32, std::size(PrimaryGroupOpTable_32));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -477,7 +483,7 @@ void InitializeSecondaryGroupTables() {
};
#undef OPD
GenerateTable(SecondInstGroupOps, SecondaryExtensionOpTable, sizeof(SecondaryExtensionOpTable) / sizeof(SecondaryExtensionOpTable[0]));
GenerateTable(SecondInstGroupOps, SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -46,6 +52,6 @@ void InitializeSecondaryModRMTables() {
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondModRMTableOps, SecondaryModRMExtensionOpTable, sizeof(SecondaryModRMExtensionOpTable) / sizeof(SecondaryModRMExtensionOpTable[0]));
GenerateTable(SecondModRMTableOps, SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -571,18 +577,18 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondBaseOps, TwoByteOpTable, sizeof(TwoByteOpTable) / sizeof(TwoByteOpTable[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable, std::size(TwoByteOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(SecondBaseOps, TwoByteOpTable_64, sizeof(TwoByteOpTable_64) / sizeof(TwoByteOpTable_64[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable_64, std::size(TwoByteOpTable_64));
}
else {
GenerateTable(SecondBaseOps, TwoByteOpTable_32, sizeof(TwoByteOpTable_32) / sizeof(TwoByteOpTable_32[0]));
GenerateTable(SecondBaseOps, TwoByteOpTable_32, std::size(TwoByteOpTable_32));
}
GenerateTableWithCopy(RepModOps, RepModOpTable, sizeof(RepModOpTable) / sizeof(RepModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(RepNEModOps, RepNEModOpTable, sizeof(RepNEModOpTable) / sizeof(RepNEModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(OpSizeModOps, OpSizeModOpTable, sizeof(OpSizeModOpTable) / sizeof(OpSizeModOpTable[0]), SecondBaseOps);
GenerateTableWithCopy(RepModOps, RepModOpTable, std::size(RepModOpTable), SecondBaseOps);
GenerateTableWithCopy(RepNEModOps, RepNEModOpTable, std::size(RepNEModOpTable), SecondBaseOps);
GenerateTableWithCopy(OpSizeModOps, OpSizeModOpTable, std::size(OpSizeModOpTable), SecondBaseOps);
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -500,7 +506,7 @@ void InitializeVEXTables() {
};
#undef OPD
GenerateTable(VEXTableOps, VEXTable, sizeof(VEXTable) / sizeof(VEXTable[0]));
GenerateTable(VEXTableGroupOps, VEXGroupTable, sizeof(VEXGroupTable) / sizeof(VEXGroupTable[0]));
GenerateTable(VEXTableOps, VEXTable, std::size(VEXTable));
GenerateTable(VEXTableGroupOps, VEXGroupTable, std::size(VEXGroupTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#pragma once
#include <FEXCore/Debug/X86Tables.h>
@@ -27,7 +33,7 @@ static inline void GenerateTable(X86InstInfo *FinalTable, U8U8InfoStruct const *
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LogMan::Throw::A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
@@ -44,7 +50,7 @@ static inline void GenerateTable(X86InstInfo *FinalTable, U16U8InfoStruct const
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LogMan::Throw::A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
@@ -61,7 +67,7 @@ static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, U8U8InfoStruct
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LogMan::Throw::A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
if (Info.Type == TYPE_COPY_OTHER) {
FinalTable[OpNum + i] = OtherLocal[OpNum + i];
}
@@ -83,7 +89,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, U16U8InfoStruct con
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LogMan::Throw::A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
if ((OpNum & 0b11'000'000) == 0b11'000'000) {
// If the mod field is 0b11 then it is a regular op
FinalTable[OpNum + i] = Info;
@@ -91,7 +97,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, U16U8InfoStruct con
else {
// If the mod field is !0b11 then this instruction is duplicated through the whole mod [0b00, 0b10] range
// and the modrm.rm space because that is used part of the instruction encoding
LogMan::Throw::A((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
LOGMAN_THROW_A((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
for (uint16_t mod = 0b00'000'000; mod < 0b11'000'000; mod += 0b01'000'000) {
for (uint16_t rm = 0b000; rm < 0b1'000; ++rm) {
FinalTable[(OpNum | mod | rm) + i] = Info;
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -254,6 +260,6 @@ void InitializeX87Tables() {
#undef OPD
#undef OPDReg
GenerateX87Table(X87Ops, X87OpTable, sizeof(X87OpTable) / sizeof(X87OpTable[0]));
GenerateX87Table(X87Ops, X87OpTable, std::size(X87OpTable));
}
}
@@ -1,3 +1,9 @@
/*
$info$
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
namespace FEXCore::X86Tables {
@@ -119,7 +125,7 @@ void InitializeXOPTables() {
};
#undef OPD
GenerateTable(XOPTableOps, XOPTable, sizeof(XOPTable) / sizeof(XOPTable[0]));
GenerateTable(XOPTableGroupOps, XOPGroupTable, sizeof(XOPGroupTable) / sizeof(XOPGroupTable[0]));
GenerateTable(XOPTableOps, XOPTable, std::size(XOPTable));
GenerateTable(XOPTableGroupOps, XOPGroupTable, std::size(XOPGroupTable));
}
}
+25 -14
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: glue|thunks ~ FEXCore side of thunks: Registration, Lookup
tags: glue|thunks
$end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include "Thunks.h"
@@ -7,6 +14,7 @@
#include <string>
#include <map>
#include <array>
#include <Interface/Context/Context.h>
#include "Interface/Core/InternalThreadState.h"
#include "FEXCore/Core/X86Enums.h"
@@ -22,23 +30,26 @@ static thread_local FEXCore::Core::InternalThreadState *Thread;
namespace FEXCore {
struct ExportEntry { const char* Name; ThunkedFunction* Fn; };
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
class ThunkHandler_impl final: public ThunkHandler {
std::shared_mutex ThunksMutex;
std::map<std::string, ThunkedFunction*> Thunks = {
{ "fex:loadlib", &LoadLib}
std::map<IR::SHA256Sum, ThunkedFunction*> Thunks = {
{
// sha256(fex:loadlib)
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80},
&LoadLib
}
};
/*
Set arg0/1 to arg regs, use CTX::HandleCallback to handle the callback
*/
static void CallCallback(void *callback, void *arg0, void* arg1) {
Thread->State.State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->State.State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
Thread->CTX->HandleCallback((uintptr_t)callback);
}
@@ -52,10 +63,10 @@ namespace FEXCore {
auto Name = Args->Name;
auto CallbackThunks = Args->CallbackThunks;
auto SOName = CTX->Config.ThunkLibsPath + "/" + (const char*)Name + "-host.so";
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
LogMan::Msg::D("Load lib: %s -> %s", Name, SOName.c_str());
auto Handle = dlopen(SOName.c_str(), RTLD_LOCAL | RTLD_NOW);
if (!Handle) {
@@ -73,7 +84,7 @@ namespace FEXCore {
LogMan::Msg::E("Load lib: failed to find export %s", InitSym.c_str());
return;
}
auto Exports = InitFN((void*)&CallCallback, CallbackThunks);
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
@@ -82,8 +93,8 @@ namespace FEXCore {
std::unique_lock lk(That->ThunksMutex);
int i;
for (i = 0; Exports[i].Name; i++) {
That->Thunks[Exports[i].Name] = Exports[i].Fn;
for (i = 0; Exports[i].sha256; i++) {
That->Thunks[*reinterpret_cast<IR::SHA256Sum*>(Exports[i].sha256)] = Exports[i].Fn;
}
LogMan::Msg::D("Loaded %d syms", i);
@@ -92,11 +103,11 @@ namespace FEXCore {
public:
ThunkedFunction* LookupThunk(const char *Name) {
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) {
std::shared_lock lk(ThunksMutex);
auto it = Thunks.find(Name);
auto it = Thunks.find(sha256);
if (it != Thunks.end()) {
return it->second;
+8 -1
View File
@@ -1,4 +1,11 @@
/*
$info$
tags: glue|thunks
$end_info$
*/
#pragma once
#include <FEXCore/IR/IR.h>
namespace FEXCore::Core {
struct InternalThreadState;
@@ -9,7 +16,7 @@ namespace FEXCore {
class ThunkHandler {
public:
virtual ThunkedFunction* LookupThunk(const char *name) = 0;
virtual ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) = 0;
virtual void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) = 0;
virtual ~ThunkHandler() { }
+226 -111
View File
@@ -60,6 +60,12 @@
"constexpr static uint8_t ROUND_MODE_TOWARDS_ZERO = 3",
"constexpr static uint8_t ROUND_MODE_FLUSH_TO_ZERO = 1 << 2",
"static constexpr FEXCore::IR::RoundType Round_Nearest {ROUND_MODE_NEAREST}",
"static constexpr FEXCore::IR::RoundType Round_Negative_Infinity {ROUND_MODE_NEGATIVE_INFINITY}",
"static constexpr FEXCore::IR::RoundType Round_Positive_Infinity {ROUND_MODE_POSITIVE_INFINITY}",
"static constexpr FEXCore::IR::RoundType Round_Towards_Zero {ROUND_MODE_TOWARDS_ZERO} /* Truncate */",
"static constexpr FEXCore::IR::RoundType Round_Host {ROUND_MODE_TOWARDS_ZERO + 1}",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_SXTX {0};",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_UXTW {1};",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_SXTW {2};"
@@ -79,9 +85,7 @@
"Blocks"
],
"Args": [
"uint64_t", "Entry",
"uint32_t", "BlockCount",
"bool", "ShouldInterpret"
"uint32_t", "BlockCount"
]
},
"CodeBlock": {
@@ -133,17 +137,14 @@
"Args": [
"uint64_t", "CodeOriginalLow",
"uint64_t", "CodeOriginalHigh",
"uint64_t", "CodePtr",
"int64_t", "Offset",
"uint8_t", "CodeLength"
]
},
"RemoveCodeEntry": {
"HasSideEffects": true,
"OpClass": "Misc",
"Args": [
"uint64_t", "RIP"
]
"OpClass": "Misc"
},
"GuestCallDirect": {
@@ -284,6 +285,32 @@
]
},
"EntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "RegisterSize",
"Args": [
"int64_t", "Offset"
],
"HelperArgs": [
"uint8_t", "RegisterSize"
]
},
"InlineEntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"OpClass": "ALU",
"DestSize": "RegisterSize",
"Args": [
"int64_t", "Offset"
],
"HelperArgs": [
"uint8_t", "RegisterSize"
]
},
"Constant": {
"Desc": ["Generates a 64bit constant inside of a GPR",
"Unsupported to create a constant in FPR"
@@ -644,8 +671,7 @@
"ArgPtr"
],
"Args":[
"const char*", "ThunkName",
"uintptr_t", "ThunkFnPtr"
"SHA256Sum", "ThunkNameHash"
]
},
@@ -772,6 +798,17 @@
]
},
"CacheLineClear": {
"Desc": ["Does a 64 byte cacheline clear at the address specified"
],
"HasSideEffects": true,
"OpClass": "Memory",
"SSAArgs": "1",
"SSANames": [
"Addr"
]
},
"Add": {
"Desc": [ "Integer Add",
"Will truncate to 64 or 32bits"
@@ -1117,7 +1154,11 @@
"DestClass": "GPRPair",
"FixedDestSize": "8",
"NumElements": "2",
"SSAArgs": "1"
"SSAArgs": "2",
"SSANames": [
"Function",
"Leaf"
]
},
"Bfi": {
@@ -1442,24 +1483,6 @@
]
},
"Float_ToGPR_U": {
"Desc": ["Moves the scalar element to a GPR with conversion",
"Converts the 32bit or 64bit float to an unsigned integer",
"Rounding mode determined by host flag's rounding mode"
],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "ElementSize",
"SSAArgs": "1",
"SSANames": [
"Scalar"
],
"Args": [
"uint8_t", "ElementSize"
]
},
"Float_ToGPR_S": {
"Desc": ["Moves the scalar element to a GPR with conversion",
"Converts the 32bit or 64bit float to an signed integer",
@@ -1468,30 +1491,16 @@
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "ElementSize",
"DestSize": "DestElementSize",
"SSAArgs": "1",
"SSANames": [
"Scalar"
],
"Args": [
"uint8_t", "ElementSize"
]
},
"Float_ToGPR_ZU": {
"Desc": ["Moves the scalar element to a GPR with conversion",
"Converts the 32bit or 64bit float to an unsigned integer rounding towards zero (Truncating)"
],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "ElementSize",
"SSAArgs": "1",
"SSANames": [
"Scalar"
"HelperArgs": [
"uint8_t", "DestElementSize"
],
"Args": [
"uint8_t", "ElementSize"
"uint8_t", "SrcElementSize"
]
},
@@ -1502,13 +1511,16 @@
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "ElementSize",
"DestSize": "DestElementSize",
"SSAArgs": "1",
"SSANames": [
"Scalar"
],
"HelperArgs": [
"uint8_t", "DestElementSize"
],
"Args": [
"uint8_t", "ElementSize"
"uint8_t", "SrcElementSize"
]
},
@@ -1538,7 +1550,6 @@
"Depending on backend, may only support GPR printing"
],
"OpClass": "Misc",
"DestSize": "GetOpSize(ssa0)",
"SSAArgs": "1",
"SSANames": [
"Value"
@@ -1630,6 +1641,23 @@
]
},
"VBic": {
"OpClass": "Vector",
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "2",
"SSANames": [
"Vector1",
"Vector2"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VOr": {
"OpClass": "Vector",
"HasDest": true,
@@ -1775,8 +1803,8 @@
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "2",
"SSANames": [
"Vector1",
"Vector2"
"VectorLower",
"VectorUpper"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
@@ -1803,6 +1831,25 @@
]
},
"VUMinV": {
"OpClass": "Vector",
"Desc": ["Does a horizontal vector unsigned minimum of elements across the source vector",
"Result is a zero extended scalar"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VURAvg": {
"OpClass": "Vector",
"Desc": ["Does an unsigned rounded average", "dst_elem = (src1_elem + src2_elem + 1) >> 1"],
@@ -1839,6 +1886,24 @@
]
},
"VPopcount": {
"OpClass": "Vector",
"Desc": ["Does a popcount for each element of the register"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VFAdd": {
"OpClass": "Vector",
"HasDest": true,
@@ -1865,8 +1930,8 @@
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "2",
"SSANames": [
"Vector1",
"Vector2"
"VectorLow",
"VectorHigh"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
@@ -2158,6 +2223,40 @@
]
},
"VUnZip": {
"OpClass": "Vector",
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "2",
"SSANames": [
"Lower",
"Upper"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VUnZip2": {
"OpClass": "Vector",
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "2",
"SSANames": [
"Lower",
"Upper"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VBSL": {
"Desc": ["Does a vector bitwise select.",
"If the bit in the field is 1 then the corresponding bit is pulled from VectorTrue",
@@ -2549,6 +2648,26 @@
]
},
"VDupElement": {
"Desc": ["Duplicates one element from the source register across the whole register"],
"OpClass": "Vector",
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
],
"Args": [
"uint8_t", "Index"
]
},
"VExtr": {
"Desc": ["Concats two vector registers together and extracts a full width register from the element index",
"Index is an element index. So it is offset by ElementSize argument",
@@ -2901,27 +3020,6 @@
]
},
"Float_FromGPR_U": {
"OpClass": "Conv",
"Desc": ["Scalar op: Converts unsigned GPR to Scalar float",
"Zeroes the upper bits of the vector register"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "DstElementSize",
"NumElements": "1",
"SSAArgs": "1",
"SSANames": [
"GPR"
],
"HelperArgs": [
"uint8_t", "DstElementSize"
],
"Args": [
"uint8_t", "SrcElementSize"
]
},
"Float_FromGPR_S": {
"OpClass": "Conv",
"Desc": ["Scalar op: Converts signed GPR to Scalar float",
@@ -2998,25 +3096,6 @@
]
},
"Vector_FToU": {
"OpClass": "Conv",
"Desc": ["Vector op: Converts float to unsigned integer",
"Rounding mode determined by host rounding mode"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"Vector_FToS": {
"OpClass": "Conv",
"Desc": ["Vector op: Converts float to signed integer, rounding towards zero",
@@ -3036,23 +3115,6 @@
]
},
"Vector_FToZU": {
"OpClass": "Conv",
"Desc": "Vector op: Converts float to unsigned integer, rounding towards zero",
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"Vector_FToZS": {
"OpClass": "Conv",
"Desc": "Vector op: Converts float to signed integer, rounding towards zero",
@@ -3090,6 +3152,28 @@
]
},
"Vector_FToI": {
"OpClass": "Conv",
"Desc": ["Vector op: Rounds float to integral",
"Rounding mode determined by argument"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"SSAArgs": "1",
"SSANames": [
"Vector"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
],
"Args":[
"FEXCore::IR::RoundType", "Round"
]
},
"VUMul": {
"OpClass": "Vector",
"HasDest": true,
@@ -3197,6 +3281,25 @@
]
},
"VUABDL": {
"OpClass": "Vector",
"Desc": ["Unsigned Absolute Difference Long"
],
"HasDest": true,
"DestClass": "FPR",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)",
"SSAArgs": "2",
"SSANames": [
"Vector1",
"Vector2"
],
"HelperArgs": [
"uint8_t", "RegisterSize",
"uint8_t", "ElementSize"
]
},
"VTBL1": {
"Desc": ["Does a vector table lookup from one register in to the destination",
"Lookup is byte sized per byte element.",
@@ -3310,6 +3413,15 @@
]
},
"F80LoadFCW": {
"OpClass": "Vector",
"HasSideEffects": true,
"SSAArgs": "1",
"SSANames": [
"X80FCW"
]
},
"F80Add": {
"OpClass": "Vector",
"HasDest": true,
@@ -3439,6 +3551,9 @@
"SSANames": [
"X80Src"
],
"Args": [
"bool", "Truncate"
],
"HelperArgs": [
"uint8_t", "Size"
]
+38 -12
View File
@@ -1,8 +1,17 @@
/*
$info$
meta: ir|dumper ~ IR -> Text
tags: ir|dumper
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <iomanip>
namespace FEXCore::IR {
#define IROP_GETNAME_IMPL
#define IROP_GETRAARGS_IMPL
@@ -12,16 +21,23 @@ namespace FEXCore::IR {
#include <FEXCore/IR/IRDefines.inc>
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, uint64_t Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, const SHA256Sum &Arg) {
*out << "sha256:";
for(auto byte: Arg.data)
*out << std::hex << std::setfill('0') << std::setw(2) << (unsigned int)byte;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, uint64_t Arg) {
*out << "#0x" << std::hex << Arg;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, const char* Arg) {
[[maybe_unused]]
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, const char* Arg) {
*out << Arg;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, CondClassType Arg) {
std::array<std::string, 22> CondNames = {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, CondClassType Arg) {
static constexpr std::array<std::string_view, 22> CondNames = {
"EQ",
"NEQ",
"UGE",
@@ -49,8 +65,8 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
*out << CondNames[Arg];
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, MemOffsetType Arg) {
std::array<std::string, 3> Names = {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, MemOffsetType Arg) {
static constexpr std::array<std::string_view, 3> Names = {
"SXTX",
"UXTW",
"SXTW",
@@ -59,7 +75,7 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
*out << Names[Arg];
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, RegisterClassType Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, RegisterClassType Arg) {
if (Arg == GPRClass.Val)
*out << "GPR";
else if (Arg == GPRFixedClass.Val)
@@ -74,7 +90,7 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
*out << "Unknown Registerclass " << Arg;
}
static void PrintArg(std::stringstream *out, IRListView<false> const* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationData *RAData) {
static void PrintArg(std::stringstream *out, IRListView const* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationData *RAData) {
auto [CodeNode, IROp] = IR->at(Arg)();
if (Arg.ID() == 0) {
@@ -83,7 +99,7 @@ static void PrintArg(std::stringstream *out, IRListView<false> const* IR, Ordere
*out << "%ssa" << std::to_string(Arg.ID());
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(Arg.ID());
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(GPRFixed"; break;
@@ -123,7 +139,7 @@ static void PrintArg(std::stringstream *out, IRListView<false> const* IR, Ordere
}
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, FEXCore::IR::FenceType Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, FEXCore::IR::FenceType Arg) {
if (Arg == IR::Fence_Load) {
*out << "Loads";
}
@@ -138,7 +154,18 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
}
}
void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAllocationData *RAData) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, FEXCore::IR::RoundType Arg) {
switch (Arg) {
case FEXCore::IR::Round_Nearest: *out << "Nearest"; break;
case FEXCore::IR::Round_Negative_Infinity: *out << "-Inf"; break;
case FEXCore::IR::Round_Positive_Infinity: *out << "+Inf"; break;
case FEXCore::IR::Round_Towards_Zero: *out << "Towards Zero"; break;
case FEXCore::IR::Round_Host: *out << "Host"; break;
default: *out << "<Unknown Round Type>"; break;
}
}
void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationData *RAData) {
auto HeaderOp = IR->GetHeader();
int8_t CurrentIndent = 0;
@@ -151,7 +178,6 @@ void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAlloc
++CurrentIndent;
AddIndent();
*out << "(%ssa0) " << "IRHeader ";
*out << "#0x" << std::hex << HeaderOp->Entry << ", ";
*out << "%ssa" << HeaderOp->Blocks.ID() << ", ";
*out << "#" << std::dec << HeaderOp->BlockCount << std::endl;
+23 -17
View File
@@ -1,19 +1,25 @@
/*
$info$
meta: ir|emitter ~ C++ Functions to generate IR. See IR.json for spec.
tags: ir|emitter
$end_info$
*/
#include <FEXCore/IR/IREmitter.h>
namespace FEXCore::IR {
void IREmitter::ResetWorkingList() {
Data.Reset();
ListData.Reset();
DualListData.Reset();
CodeBlocks.clear();
CurrentWriteCursor = nullptr;
// This is necessary since we do "null" pointer checks
InvalidNode = reinterpret_cast<OrderedNode*>(ListData.Allocate(sizeof(OrderedNode)));
InvalidNode = reinterpret_cast<OrderedNode*>(DualListData.ListAllocate(sizeof(OrderedNode)));
memset(InvalidNode, 0, sizeof(OrderedNode));
CurrentCodeBlock = nullptr;
}
void IREmitter::ReplaceAllUsesWithRange(OrderedNode *Node, OrderedNode *NewNode, AllNodesIterator After, AllNodesIterator End) {
uintptr_t ListBegin = ListData.Begin();
uintptr_t ListBegin = DualListData.ListBegin();
auto NodeId = Node->Wrapped(ListBegin).ID();
while (After != End) {
@@ -38,8 +44,8 @@ void IREmitter::ReplaceAllUsesWithRange(OrderedNode *Node, OrderedNode *NewNode,
}
void IREmitter::ReplaceNodeArgument(OrderedNode *Node, uint8_t Arg, OrderedNode *NewArg) {
uintptr_t ListBegin = ListData.Begin();
uintptr_t DataBegin = Data.Begin();
uintptr_t ListBegin = DualListData.ListBegin();
uintptr_t DataBegin = DualListData.DataBegin();
FEXCore::IR::IROp_Header *IROp = Node->Op(DataBegin);
OrderedNodeWrapper OldArgWrapper = IROp->Args[Arg];
@@ -50,8 +56,8 @@ void IREmitter::ReplaceNodeArgument(OrderedNode *Node, uint8_t Arg, OrderedNode
}
void IREmitter::RemoveArgUses(OrderedNode *Node) {
uintptr_t ListBegin = ListData.Begin();
uintptr_t DataBegin = Data.Begin();
uintptr_t ListBegin = DualListData.ListBegin();
uintptr_t DataBegin = DualListData.DataBegin();
FEXCore::IR::IROp_Header *IROp = Node->Op(DataBegin);
@@ -65,7 +71,7 @@ void IREmitter::RemoveArgUses(OrderedNode *Node) {
void IREmitter::Remove(OrderedNode *Node) {
RemoveArgUses(Node);
Node->Unlink(ListData.Begin());
Node->Unlink(DualListData.ListBegin());
}
IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode* insertAfter) {
@@ -76,14 +82,14 @@ IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode
if (insertAfter) {
LinkCodeBlocks(insertAfter, CodeNode);
} else {
LogMan::Throw::A(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
LOGMAN_THROW_A(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
// Find last block
auto LastBlock = CurrentCodeBlock;
while (LastBlock->Header.Next.GetNode(ListData.Begin()) != InvalidNode)
LastBlock = LastBlock->Header.Next.GetNode(ListData.Begin());
while (LastBlock->Header.Next.GetNode(DualListData.ListBegin()) != InvalidNode)
LastBlock = LastBlock->Header.Next.GetNode(DualListData.ListBegin());
// Append it after the last block
LinkCodeBlocks(LastBlock, CodeNode);
}
@@ -95,12 +101,12 @@ IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode
void IREmitter::SetCurrentCodeBlock(OrderedNode *Node) {
CurrentCodeBlock = Node;
LogMan::Throw::A(Node->Op(Data.Begin())->Op == OP_CODEBLOCK, "Node wasn't codeblock. It was '%s'", std::string(IR::GetName(Node->Op(Data.Begin())->Op)).c_str());
SetWriteCursor(Node->Op(Data.Begin())->CW<IROp_CodeBlock>()->Begin.GetNode(ListData.Begin()));
LOGMAN_THROW_A(Node->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Node wasn't codeblock. It was '%s'", std::string(IR::GetName(Node->Op(DualListData.DataBegin())->Op)).c_str());
SetWriteCursor(Node->Op(DualListData.DataBegin())->CW<IROp_CodeBlock>()->Begin.GetNode(DualListData.ListBegin()));
}
void IREmitter::ReplaceWithConstant(OrderedNode *Node, uint64_t Value) {
auto Header = Node->Op(Data.Begin());
auto Header = Node->Op(DualListData.DataBegin());
if (IRSizes[Header->Op] >= sizeof(IROp_Constant)) {
// Unlink any arguments the node currently has
+166 -114
View File
@@ -1,3 +1,10 @@
/*
$info$
meta: ir|parser ~ Text -> IR
tags: ir|parser
$end_info$
*/
#include <string>
#include <vector>
#include <istream>
@@ -59,7 +66,8 @@ std::string DecodeErrorToString(DecodeFailure Failure) {
case DecodeFailure::DECODE_INVALID_CONDFLAG: return "Invalid Conditional name";
case DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE: return "Invalid Memory Offset Type";
case DecodeFailure::DECODE_INVALID_FENCETYPE: return "Invalid Fence Type";
};
}
return "Unknown Error";
}
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
@@ -67,22 +75,31 @@ std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
class IRParser: public FEXCore::IR::IREmitter {
public:
template<typename Type>
std::pair<DecodeFailure, Type> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, Type> DecodeValue(const std::string &Arg) {
return {DecodeFailure::DECODE_UNKNOWN_TYPE, {}};
}
template<>
std::pair<DecodeFailure, uint8_t> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, uint8_t> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
}
template<>
std::pair<DecodeFailure, uint16_t> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, bool> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE || Result > 1) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result != 0};
}
template<>
std::pair<DecodeFailure, uint16_t> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint16_t Result = strtoul(&Arg.at(1), nullptr, 0);
@@ -91,7 +108,7 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, uint32_t> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, uint32_t> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint32_t Result = strtoul(&Arg.at(1), nullptr, 0);
@@ -100,7 +117,7 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, uint64_t> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, uint64_t> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint64_t Result = strtoull(&Arg.at(1), nullptr, 0);
@@ -109,7 +126,47 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, int64_t> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
int64_t Result = (int64_t)strtoull(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, IR::SHA256Sum> DecodeValue(const std::string &Arg) {
IR::SHA256Sum Result;
if (Arg.at(0) != 's' || Arg.at(1) != 'h' || Arg.at(2) != 'a' || Arg.at(3) != '2' || Arg.at(4) != '5' || Arg.at(5) != '6' || Arg.at(6) != ':')
return {DecodeFailure::DECODE_INVALIDCHAR, Result};
auto GetDigit = [](const std::string &Arg, int pos, uint8_t *val) {
auto chr = Arg.at(pos);
if (chr >= '0' && chr <= '9') {
*val = chr - '0';
return true;
} else if (chr >= 'a' && chr <= 'f') {
*val = 10 + chr - 'a';
return true;
} else {
return false;
}
};
for (size_t i = 0; i < sizeof(Result.data); i++) {
uint8_t high, low;
if (!GetDigit(Arg, 7 + 2 * i + 0, &high) || !GetDigit(Arg, 7 + 2 * i + 1, &low)) {
return {DecodeFailure::DECODE_INVALIDRANGE, Result};
}
Result.data[i] = high * 16 + low;
}
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType> DecodeValue(const std::string &Arg) {
if (Arg == "GPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRClass};
}
@@ -127,7 +184,7 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition> DecodeValue(const std::string &Arg) {
uint8_t Size{}, Elements{1};
int NumArgs = sscanf(Arg.c_str(), "i%hhdv%hhd", &Size, &Elements);
@@ -139,8 +196,8 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, FEXCore::IR::CondClassType> DecodeValue(std::string &Arg) {
std::array<std::string, 22> CondNames = {
std::pair<DecodeFailure, FEXCore::IR::CondClassType> DecodeValue(const std::string &Arg) {
static constexpr std::array<std::string_view, 22> CondNames = {
"EQ",
"NEQ",
"UGE",
@@ -174,8 +231,8 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType> DecodeValue(std::string &Arg) {
std::array<std::string, 3> Names = {
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType> DecodeValue(const std::string &Arg) {
static constexpr std::array<std::string_view, 3> Names = {
"SXTX",
"UXTW",
"SXTW",
@@ -190,8 +247,8 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, FEXCore::IR::FenceType> DecodeValue(std::string &Arg) {
std::array<std::string, 3> Names = {
std::pair<DecodeFailure, FEXCore::IR::FenceType> DecodeValue(const std::string &Arg) {
static constexpr std::array<std::string_view, 3> Names = {
"Loads",
"Stores",
"LoadStores",
@@ -206,23 +263,22 @@ class IRParser: public FEXCore::IR::IREmitter {
}
template<>
std::pair<DecodeFailure, OrderedNode*> DecodeValue(std::string &Arg) {
std::pair<DecodeFailure, OrderedNode*> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
// Strip off the type qualifier from the ssa value
size_t ArgEnd = std::string::npos;
std::string SSAName = trim(Arg);
ArgEnd = SSAName.find_first_of(" ");
const size_t ArgEnd = SSAName.find_first_of(' ');
if (ArgEnd != std::string::npos) {
SSAName = SSAName.substr(0, ArgEnd);
}
SSAName = SSAName.substr(0, ArgEnd);
}
// Forward declarations may make this not succed
// Forward declarations may make this not succed
auto Op = SSANameMapper.find(SSAName);
if (Op == SSANameMapper.end()) {
if (Op == SSANameMapper.end()) {
return {DecodeFailure::DECODE_UNKNOWN_SSA, nullptr};
}
}
return {DecodeFailure::DECODE_OKAY, Op->second};
}
@@ -246,21 +302,21 @@ class IRParser: public FEXCore::IR::IREmitter {
IRParser(std::istream *text) {
InitializeStaticTables();
std::string TmpLine;
while (!text->eof()) {
std::getline(*text, TmpLine);
if (text->eof()) {
break;
}
if (text->eof()) {
break;
}
if (text->fail()) {
LogMan::Msg::E("Failed to getline on line: %ld", Lines.size());
LogMan::Msg::EFmt("Failed to getline on line: {}", Lines.size());
return;
}
Lines.emplace_back(TmpLine);
}
ResetWorkingList();
ResetWorkingList();
Loaded = Parse();
}
@@ -271,11 +327,11 @@ class IRParser: public FEXCore::IR::IREmitter {
bool Parse() {
auto CheckPrintError = [&](LineDefinition &Def, DecodeFailure Failure) -> bool {
const auto CheckPrintError = [&](const LineDefinition &Def, DecodeFailure Failure) -> bool {
if (Failure != DecodeFailure::DECODE_OKAY) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Value Couldn't be decoded due to %s", DecodeErrorToString(Failure).c_str());
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("Value Couldn't be decoded due to {}", DecodeErrorToString(Failure));
return false;
}
@@ -283,13 +339,13 @@ class IRParser: public FEXCore::IR::IREmitter {
};
// String parse every line for our definitions
for (size_t i = 0; i < Lines.size(); ++i) {
std::string Line = Lines[i];
for (size_t i = 0; i < Lines.size(); ++i) {
std::string Line = Lines[i];
LineDefinition Def{};
CurrentDef = &Def;
CurrentDef = &Def;
Def.LineNumber = i;
Line = trim(Line);
Line = trim(Line);
// Skip empty lines
if (Line.empty()) {
@@ -303,35 +359,37 @@ class IRParser: public FEXCore::IR::IREmitter {
}
size_t CurrentPos{};
// Let's see if this node is assigning something first
if (Line[0] == '%') {
// Let's see if this node is assigning something first
if (Line[0] == '%') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of("=", CurrentPos)) != std::string::npos) {
if ((DefinitionEnd = Line.find_first_of('=', CurrentPos)) != std::string::npos) {
Def.Definition = Line.substr(0, DefinitionEnd);
Def.Definition = trim(Def.Definition);
Def.HasDefinition = true;
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA declaration without assignment");
LogMan::Msg::EFmt("Error on Line: {}", i);
LogMan::Msg::EFmt("{}", Lines[i]);
LogMan::Msg::EFmt("SSA declaration without assignment");
return false;
}
}
}
// Check if we are pulling in some IR from the IR Printer
// Prints (%ssa%d) at the start of lines without a definition
if (Line[0] == '(') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of(")", CurrentPos)) != std::string::npos) {
if ((DefinitionEnd = Line.find_first_of(')', CurrentPos)) != std::string::npos) {
size_t SSAEnd = std::string::npos;
if ((SSAEnd = Line.find_last_of(" ", DefinitionEnd)) != std::string::npos) {
if ((SSAEnd = Line.find_last_of(' ', DefinitionEnd)) != std::string::npos) {
std::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
Type = trim(Type);
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
if (!CheckPrintError(Def, DefinitionSize.first)) {
return false;
}
Def.Size = DefinitionSize.second;
}
@@ -340,9 +398,9 @@ class IRParser: public FEXCore::IR::IREmitter {
CurrentPos = DefinitionEnd + 1;
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA value with numbered SSA provided but no closing parentheses");
LogMan::Msg::EFmt("Error on Line: {}", i);
LogMan::Msg::EFmt("{}", Lines[i]);
LogMan::Msg::EFmt("SSA value with numbered SSA provided but no closing parentheses");
return false;
}
}
@@ -350,7 +408,7 @@ class IRParser: public FEXCore::IR::IREmitter {
if (Def.HasDefinition) {
// Let's check if we have a size declared with this variable
size_t NameEnd = std::string::npos;
if ((NameEnd = Def.Definition.find_first_of(" ")) != std::string::npos) {
if ((NameEnd = Def.Definition.find_first_of(' ')) != std::string::npos) {
std::string Type = Def.Definition.substr(NameEnd + 1);
Type = trim(Type);
Def.Definition = trim(Def.Definition.substr(0, NameEnd));
@@ -361,9 +419,9 @@ class IRParser: public FEXCore::IR::IREmitter {
}
if (Def.Definition == "%Invalid") {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Definition tried to define reserved %Invalid ssa node");
LogMan::Msg::EFmt("Error on Line: {}", i);
LogMan::Msg::EFmt("{}", Lines[i]);
LogMan::Msg::EFmt("Definition tried to define reserved %Invalid ssa node");
return false;
}
}
@@ -380,9 +438,9 @@ class IRParser: public FEXCore::IR::IREmitter {
}
else {
if (RemainingLine.empty()) {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Line without an IROp?");
LogMan::Msg::EFmt("Error on Line: {}", i);
LogMan::Msg::EFmt("{}", Lines[i]);
LogMan::Msg::EFmt("Line without an IROp?");
return false;
}
@@ -399,12 +457,10 @@ class IRParser: public FEXCore::IR::IREmitter {
}
else {
while (!RemainingLine.empty()) {
size_t ArgEnd = std::string::npos;
ArgEnd = RemainingLine.find_first_of(",");
const size_t ArgEnd = RemainingLine.find(',');
std::string Arg = trim(RemainingLine.substr(0, ArgEnd));
std::string Arg = RemainingLine.substr(0, ArgEnd);
Arg = trim(Arg);
Def.Args.emplace_back(Arg);
Def.Args.emplace_back(std::move(Arg));
RemainingLine.erase(0, ArgEnd+1); // +1 to ensure we go past the ','
if (ArgEnd == std::string::npos)
@@ -413,17 +469,17 @@ class IRParser: public FEXCore::IR::IREmitter {
}
}
Defs.emplace_back(Def);
}
CurrentDef = &Defs.emplace_back(std::move(Def));
}
// Ensure all of the ops are real ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
auto Op = NameToOpMap.find(Def.IROp);
if (Op == NameToOpMap.end()) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("IROp '%s' doesn't exist", Def.IROp.c_str());
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("IROp '{}' doesn't exist", Def.IROp);
return false;
}
Def.OpEnum = Op->second;
@@ -433,40 +489,38 @@ class IRParser: public FEXCore::IR::IREmitter {
IRPair<IROp_IRHeader> IRHeader;
{
auto &Def = Defs[0];
CurrentDef = &Def;
CurrentDef = &Def;
if (Def.OpEnum != FEXCore::IR::IROps::OP_IRHEADER) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("First op needs to be IRHeader. Was '%s'", Def.IROp.c_str());
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("First op needs to be IRHeader. Was '{}'", Def.IROp);
return false;
}
auto Entry = DecodeValue<uint64_t>(Def.Args[0]);
auto CodeBlockCount = DecodeValue<uint64_t>(Def.Args[2]);
auto CodeBlockCount = DecodeValue<uint64_t>(Def.Args[1]);
if (!CheckPrintError(Def, Entry.first)) return false;
if (!CheckPrintError(Def, CodeBlockCount.first)) return false;
IRHeader = _IRHeader(InvalidNode, Entry.second, CodeBlockCount.second, false);
IRHeader = _IRHeader(InvalidNode, CodeBlockCount.second);
}
SetWriteCursor(nullptr); // isolate the header from everything following
// Initialize SSANameMapper with Invalid value
SSANameMapper["%Invalid"] = Invalid();
SSANameMapper.insert_or_assign("%Invalid", Invalid());
// Spin through the blocks and generate basic block ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
if (Def.OpEnum == FEXCore::IR::IROps::OP_CODEBLOCK) {
auto CodeBlock = _CodeBlock(InvalidNode, InvalidNode);
SSANameMapper[Def.Definition] = CodeBlock.Node;
SSANameMapper.insert_or_assign(Def.Definition, CodeBlock.Node);
Def.Node = CodeBlock.Node;
if (i == 1) {
// First code block is the entry block
// Link the header to the first block
IRHeader.first->Blocks = CodeBlock.Node->Wrapped(ListData.Begin());
IRHeader.first->Blocks = CodeBlock.Node->Wrapped(DualListData.ListBegin());
}
CodeBlocks.emplace_back(CodeBlock.Node);
}
@@ -478,63 +532,64 @@ class IRParser: public FEXCore::IR::IREmitter {
FEXCore::IR::IROp_CodeBlock *CurrentBlockOp{};
for(size_t i = 1; i < Defs.size(); ++i) {
auto &Def = Defs[i];
CurrentDef = &Def;
CurrentDef = &Def;
switch (Def.OpEnum) {
// Special handled
case FEXCore::IR::IROps::OP_IRHEADER:
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("IRHEADER used in the middle of the block!");
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("IRHEADER used in the middle of the block!");
return false; // only one OP_IRHEADER allowed per block
case FEXCore::IR::IROps::OP_CODEBLOCK: {
SetWriteCursor(nullptr); // isolate from previous block
if (CurrentBlock != nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("CodeBlock being used inside of already existing codeblock!");
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("CodeBlock being used inside of already existing codeblock!");
return false;
}
CurrentBlock = Def.Node;
CurrentBlockOp = CurrentBlock->Op(Data.Begin())->CW<FEXCore::IR::IROp_CodeBlock>();
CurrentBlockOp = CurrentBlock->Op(DualListData.DataBegin())->CW<FEXCore::IR::IROp_CodeBlock>();
break;
}
case FEXCore::IR::IROps::OP_BEGINBLOCK: {
if (CurrentBlock == nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("EndBlock being used outside of a block!");
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("EndBlock being used outside of a block!");
return false;
}
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
if (!CheckPrintError(Def, Adjust.first)) return false;
if (!CheckPrintError(Def, Adjust.first)) {
return false;
}
Def.Node = _BeginBlock(Adjust.second);
CurrentBlockOp->Begin = Def.Node->Wrapped(ListData.Begin());
CurrentBlockOp->Begin = Def.Node->Wrapped(DualListData.ListBegin());
break;
}
case FEXCore::IR::IROps::OP_ENDBLOCK: {
if (CurrentBlock == nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("EndBlock being used outside of a block!");
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("EndBlock being used outside of a block!");
return false;
}
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
if (!CheckPrintError(Def, Adjust.first)) return false;
if (!CheckPrintError(Def, Adjust.first)) {
return false;
}
Def.Node = _EndBlock(Adjust.second);
CurrentBlockOp->Last = Def.Node->Wrapped(ListData.Begin());
CurrentBlockOp->Last = Def.Node->Wrapped(DualListData.ListBegin());
CurrentBlock = nullptr;
CurrentBlockOp = nullptr;
@@ -543,25 +598,23 @@ class IRParser: public FEXCore::IR::IREmitter {
}
case FEXCore::IR::IROps::OP_DUMMY: {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Dummy op must not be used");
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("Dummy op must not be used");
break;
}
#define IROP_PARSER_SWITCH_HELPERS
#include <FEXCore/IR/IRDefines.inc>
default: {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Unhandled Op enum '%s' in parser", Def.IROp.c_str());
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
LogMan::Msg::EFmt("Unhandled Op enum '{}' in parser", Def.IROp);
return false;
break;
}
}
if (Def.HasDefinition) {
auto IROp = Def.Node->Op(Data.Begin());
auto IROp = Def.Node->Op(DualListData.DataBegin());
if (Def.Size.Elements()) {
IROp->Size = Def.Size.Bytes() * Def.Size.Elements();
IROp->ElementSize = Def.Size.Bytes();
@@ -570,7 +623,7 @@ class IRParser: public FEXCore::IR::IREmitter {
IROp->Size = Def.Size.Bytes();
IROp->ElementSize = 0;
}
SSANameMapper[Def.Definition] = Def.Node;
SSANameMapper.insert_or_assign(Def.Definition, Def.Node);
}
}
@@ -578,11 +631,11 @@ class IRParser: public FEXCore::IR::IREmitter {
}
void InitializeStaticTables() {
if (NameToOpMap.size() == 0) {
if (NameToOpMap.empty()) {
for (FEXCore::IR::IROps Op = FEXCore::IR::IROps::OP_DUMMY;
Op <= FEXCore::IR::IROps::OP_LAST;
Op = static_cast<FEXCore::IR::IROps>(static_cast<uint32_t>(Op) + 1)) {
NameToOpMap[FEXCore::IR::GetName(Op)] = Op;
NameToOpMap.insert_or_assign(FEXCore::IR::GetName(Op), Op);
}
}
}
@@ -590,13 +643,12 @@ class IRParser: public FEXCore::IR::IREmitter {
} // anon namespace
IREmitter* Parse(std::istream *in) {
auto parser = new IRParser(in);
std::unique_ptr<IREmitter> Parse(std::istream *in) {
auto parser = std::make_unique<IRParser>(in);
if (parser->Loaded) {
return parser;
} else {
delete parser;
return nullptr;
}
}
Loaded 100 of 463 files, more files were not shown because too many files have changed in this diff. Show more