Compare commits

...
91 Commits
Author SHA1 Message Date
Ryan Houdek 597d524f9e Docs: Update for release FEX-2111 2021-11-06 21:50:00 -07:00
Ryan Houdek b4a71a2144 Merge pull request #1353 from lioncash/bmi
OpcodeDispatcher: Implement BLSR/BLSMSK
2021-11-06 20:56:23 -07:00
Ryan Houdek 09ee6d3bf4 Merge pull request #1352 from Sonicadvance1/more_symlink
Linux/FM: Follow more symlinks in emulation
2021-11-06 20:53:45 -07:00
lioncash 7d2b3d0846 CPUID: Signify support for BMI1
Now that all of BMI1's instructions are implemented, we can signify that
we support it in CPUID.
2021-11-06 23:37:54 -04:00
lioncash e0973e19fc OpcodeDispatcher: Implement handling for BLSMSK 2021-11-06 23:36:52 -04:00
lioncash ff9190204c OpcodeDispatcher: Implement handling for BLSR 2021-11-06 23:28:23 -04:00
Ryan Houdek decd8bec31 Linux/FM: Follow more symlinks in emulation
Depending on how wine is launching it may do a PATH scan.
So we need to follow symlinks in a few more syscalls
2021-11-06 18:49:52 -07:00
Ryan Houdek a393d6609f Merge pull request #1351 from Sonicadvance1/fix_execve_softlinks
Linux: Fixes execve on softlinks in rootfs
2021-11-06 17:24:53 -07:00
Ryan Houdek 8aebbbd0ca Merge pull request #1350 from Sonicadvance1/FEXConfig_fix_timeout
FEXConfig: Fixes timeout in select causing 100% CPU load
2021-11-06 17:24:47 -07:00
Ryan Houdek 8b64546579 Merge pull request #1349 from Sonicadvance1/fix_paranoid
Arm64: Fixes paranoid TSO mode
2021-11-06 17:24:41 -07:00
Ryan Houdek 95457bc78c Merge pull request #1348 from Sonicadvance1/sigchld_drop
SignalDelegator: No longer do magic on SIGCHLD
2021-11-06 17:24:35 -07:00
Ryan Houdek a9d31227bf Merge pull request #1347 from Sonicadvance1/cpuid_hybrid_flag
CPUID: Adds support for hybrid flag
2021-11-06 17:24:29 -07:00
Ryan Houdek cae2f8cac4 Merge pull request #1346 from Sonicadvance1/hide_48bit_va
Allocator: Reserve upper 128TB of VA on 64-bit process
2021-11-06 17:24:00 -07:00
Ryan Houdek d7764d37db Linux: Fixes execve on softlinks in rootfs
Ubuntu soft links a bunch of binaries in /usr/bin to softlinks that live
in /etc/alternatives/

When hitting any of these alternative softlinks execve would fail if the
host also didn't have the same softlink paths.

Allows us to correctly follow the symlinks on execve as well which fixes
launching wine directly from the wine symlink.
Alternatively you could have launched /usr/bin/wine-stable directly.

Also fixes FEX strace again.
2021-11-06 16:26:23 -07:00
Ryan Houdek bf5042bdf0 FEXConfig: Fixes timeout in select causing 100% CPU load
glibc 2.34 changed the select interface to update the timeout on return
to more closely match the kernel interface.
glibc 2.33 always made a copy instead of updating.
Make sure to set the timeout on each iteration of select otherwise we
will end up having a timeout of zero. Thus burning a CPU core.
2021-11-06 14:33:44 -07:00
Ryan Houdek 61a0508ff6 Arm64: Fixes paranoid TSO mode
Vector loadstores were crashing. Now we emulate on load and backpatch on
store.

Store can't effectively emulate so it's better to backpatch.
2021-11-06 04:03:37 -07:00
Ryan Houdek 4ebbca45be CPUID: Adds support for hybrid flag
CPUID lets the application know if it is running on a CPU with hybrid
CPU clusters.
This matches big.little fairly easily. Walk the affinity mask and
check if we are running on a big.little system and report it to the
guest.

For x86-64 host just pass through the flag.
2021-11-06 03:01:39 -07:00
Ryan Houdek 8244d55276 SignalDelegator: No longer do magic on SIGCHLD
We have been setting the host sa_flags to handle this for a while now.
So just pass the signals to the guest as expected
2021-11-06 03:00:00 -07:00
Ryan Houdek 4b47e66135 FEXCore/Utils: Adds File loading helper
This will be used in multiple locations now.
2021-11-06 02:56:31 -07:00
Ryan Houdek df2f1ad074 Allocator: Reserve upper 128TB of VA on 64-bit process
Only a partial fix for #1330, still needs preemption disabled to work.

On x86-64 hosts the Linux kernel resides in the top bit of VA which
isn't mapped in to userspace.
This means that userspace will never receive pointers living with that
top bit set unless you're running a 57bit VA host.

This results in userspace pointers never needing to do the sign
extending pointer canonicalization. But additionally some applications
actually don't understand the pointer canonicalization.
This results in bugs like: https://github.com/golang/go/issues/49405
Now if you're running on a 57bit VA host, this will end up behaving like
FEX but it seems like no one in golang land has really messed with 57bit
VA yet.

In AArch64, when configured with a 48bit VA, the userspace gets the full
48bit VA space and on EL mode switch has the full address range change
to the kernel's 48bit VA.
This means that we will /very/ likely allocate pointers in the high
48bit space since Linux currently allocates top-down.

So behave more like x86-64, hide the top 128TB of memory space from the
guest before boot.

Testing: Took the M1Max 15ms to 21ms allocate the top 128TB.
2021-11-06 01:10:00 -07:00
Ryan Houdek 8f170d4aa0 Merge pull request #1345 from Sonicadvance1/EnvironmentLoader_Parse
Fixes environment loader not hooked up to ArgumentLoader
2021-11-05 17:31:21 -07:00
Ryan Houdek 285ef38717 Merge pull request #1344 from lioncash/bmi
OpcodeDispatcher: Implement handling for BLSI
2021-11-05 17:31:10 -07:00
Ryan Houdek be12059e8f Fixes environment loader not hooked up to ArgumentLoader
Fixes #1334

Fixes the issue of `FEX_CORE=irjit` not working.
2021-11-04 23:41:22 -07:00
lioncash b47cb20619 OpcodeDispatcher: Implement handling for BLSI
Now all that remains is handling for BLSMSK and BLSR
2021-11-04 18:49:15 -04:00
lioncash 166c96320c Frontend: Handle VEX-encoded destination operands
BLSI, BLSMSK, and BLSR make use of these, for example.
2021-11-04 17:52:57 -04:00
lioncash ff24fe872d X86Tables: Relocate size descriptors at the end of uint64_t
This leaves the remaining bits available for use without needing to work
around the size fields.
2021-11-04 16:10:23 -04:00
lioncash e317424b86 X86Tables: Increase InstFlags to uint64_t
We've run out of the range of 32 bits already and will need to use
another flag in upcoming changes, so we need to expand our flags to be
64-bit.

While we're at it, we can use a dedicated type alias for the instruction
flags to make the interface changeable from one spot in the future.
2021-11-04 16:10:20 -04:00
Ryan Houdek babb81a240 Merge pull request #1343 from Sonicadvance1/sigbus_share
Arm64: Consolidate HandleSIGBUS
2021-11-03 01:46:36 -07:00
Ryan Houdek 235367b67a Merge pull request #1342 from Sonicadvance1/tear_telemetry
Telemetry: Adds telemetry for when an application tears
2021-11-03 01:46:27 -07:00
Ryan Houdek 56e5e78b25 Merge pull request #1341 from Sonicadvance1/store_op_size
IR: Fixes memory ops having a duplicate size field
2021-11-03 01:46:17 -07:00
Ryan Houdek 9e8af23456 Merge pull request #1340 from Sonicadvance1/syscall_nanosleep
Syscall: Fix 32-bit nanosleep always passing valid remainder
2021-11-03 01:46:04 -07:00
Ryan Houdek 5b9da4f2be Arm64: Consolidate HandleSIGBUS
We can share this between the interpreter and the JIT. Necessary to
support the TSO-correct interpreter path.

With this change the interpreter is TSO-correct for GPRs. Just not FPRs
yet.
2021-11-02 23:54:25 -07:00
Ryan Houdek 1f4a10ef1f unittests: Update tests for new IR operand ordering 2021-11-02 22:51:38 -07:00
Ryan Houdek 4c712ca111 Telemetry: Adds telemetry for when an application tears
This can be used as an early indicator of an application doing nefarious
things.
2021-11-02 22:46:50 -07:00
Ryan Houdek 3bcc8ca695 IR: Fixes memory ops having a duplicate size field
There's zero need for these to have an independent size field and it was
just confusing.
For stores it was always set to zero and for loads it was just
duplicated.

In addition this allows introspection of the store op without casting
the op, which can be useful in edge cases
2021-11-02 22:38:59 -07:00
Ryan Houdek 83073a880a Syscall: Fix 32-bit nanosleep always passing valid remainder
This doesn't really change behaviour but makes sure we are consistent
2021-11-02 21:54:40 -07:00
Ryan Houdek 34b2f93ddf Merge pull request #1338 from lioncash/bic
IR: Add handling for ANDN operations
2021-11-02 19:29:29 -07:00
lioncash 49dae08b3d OpcodeDispatcher: Make use of the new Andn IR op where applicable
Now that we have the handling in place, we can make use of it to
simplify some operations and resolve some lingering TODO comments.
2021-11-02 21:59:56 -04:00
lioncash 0b700de7d9 IR: Add handling for ANDN operations
This is a pretty straightforward operation that can be nicely modeled
by the BIC instruction on ARMv8, which is nice since we can get rid of
the need to manually perform the And and Not operations.
2021-11-02 21:59:53 -04:00
Ryan Houdek 43454abc63 Merge pull request #1339 from lioncash/nodiscard
Core: Mark relevant Interpreter/JIT functions as [[nodiscard]]
2021-11-02 18:29:51 -07:00
Ryan Houdek 76538be0e0 Merge pull request #1337 from lioncash/bmi-bextr
OpcodeDispatcher: Handle BMI1 BEXTR
2021-11-02 18:26:52 -07:00
Ryan Houdek aa1c47cd75 Merge pull request #1336 from lioncash/fmt
ALUOps: Fix left-over printf specifier in fmt log
2021-11-02 18:19:26 -07:00
lioncash 002867bc2a Core: Mark relevant Interpreter/JIT functions as [[nodiscard]]
Lets the compiler warn loudly when the result from any of these
functions are left unused (indicating a bug).
2021-11-02 18:31:41 -04:00
lioncash 79d6bf2840 OpcodeDispatcher: Handle BMI BEXTR 2021-11-02 16:07:37 -04:00
Lioncash 31030e6f85 ALUOps: Fix left-over printf specifier in fmt log 2021-11-02 14:26:44 -04:00
lioncash 26d493a66e Frontend: Handle VEX on second source operands 2021-11-01 14:55:32 -04:00
Ryan Houdek e0343647c9 Merge pull request #1333 from Sonicadvance1/virtio_ioctls
Linux: Implements virtio ioctls for 32-bit
2021-10-28 11:09:03 -07:00
Ryan Houdek 33151e16a2 Linux: Implements virtio ioctls for 32-bit
This makes running Steam under parallels more sane
2021-10-27 13:01:28 -07:00
Ryan Houdek 7e9201cf0d Merge pull request #1325 from lioncash/bmi
Frontend: Handle VEX source operands
2021-10-22 08:38:14 -07:00
Lioncash 877db85428 OpcodeDecoder: Handle ANDN 2021-10-22 11:18:46 -04:00
Ryan Houdek f9078f8ded Merge pull request #1329 from Sonicadvance1/fix_fexloader_argument_passing
Linux: Fixes FEXLoader argument passing
2021-10-21 23:42:47 -07:00
Ryan Houdek e547f0cad6 Merge pull request #1328 from Sonicadvance1/static_pie_error
Cmake: Change static-pie message to indicate compiled without it
2021-10-21 23:42:37 -07:00
Ryan Houdek a3b39afef2 Merge pull request #1327 from Sonicadvance1/less_native
Arm64: Don't fall back to native
2021-10-21 23:42:24 -07:00
Ryan Houdek 43431edd45 Linux: Fixes FEXLoader argument passing
In the case of binfmt_misc being installed, but the user was still using
FEXLoader to pass in arguments then we wouldn't pass the arguments
forward to applications passed through execve.

This resolves an issue where Wine would fail to know where the rootfs
is since Wine launches a bunch of processes.

eg: `FEXLoader -R Ubuntu_21_04 wine winecfg` would fail before

Fixes #1323
2021-10-21 21:21:40 -07:00
Ryan Houdek c59efaef7a Cmake: Change static-pie message to indicate compiled without it
If glibc is compiled without static-pie then we can't detect that. We
will just get a compile failure.
Looks like ALARM is compiling glibc without --enable-static-pie for
whatever reason.

Fixes #1326 as much as we can. We need to ask the ALARM maintainers to
change their configuration.
2021-10-21 20:40:53 -07:00
Ryan Houdek 0bfc1bbe70 Arm64: Don't fall back to native
In the case of Arm64, make sure not to fallback to native if we hit an
unsupported CPU.
Can cause issues depending on system configuration.
2021-10-21 20:39:49 -07:00
Ryan Houdek e9937d9a85 Merge pull request #1307 from Sonicadvance1/InterpreterDispatcher
Interpreter: Splits ops in to separate files
2021-10-21 16:22:45 -07:00
Lioncash f088f0a236 Frontend: Handle VEX source operands
This will allow us to begin implementing BMI instructions.
2021-10-21 10:50:13 -04:00
Ryan Houdek a40a0cbb12 Interpreter: Splits ops in to separate files
I need this for something else so I'm doing this now
2021-10-20 00:15:33 -07:00
Ryan Houdek 28d084bf78 Merge pull request #1321 from Sonicadvance1/fix_arm_asserts
JIT: Fixes asserts added to the JIT
2021-10-19 11:13:33 -07:00
Ryan Houdek 435137e1a2 JIT: Fixes asserts added to the JIT
Fixes #1319
2021-10-19 10:42:55 -07:00
Ryan Houdek ff74e0a0ad Merge pull request #1317 from Sonicadvance1/JITSymbols_by_library
JITSymbols: Allow grouping JIT symbols by guest named regions
2021-10-16 22:13:59 -07:00
Ryan Houdek d847f6e1b3 JITSymbols: Allow grouping JIT symbols by guest named regions
This lets us have JITsymbols grouped by library.
Useful for determining where to thunk.

Sadly perf doesn't have an option to deduplicate regions by name, so
some external tooling is necessary to make it look nice.
2021-10-16 21:10:57 -07:00
Ryan Houdek 64aa4f00ca Merge pull request #1316 from neobrain/fix_attribute_warnings
Thunks/vulkan: Suppress compiler warnings about unknown attributes
2021-10-15 20:40:30 -07:00
Tony Wasserka 50c165d291 Thunks/vulkan: Suppress compiler warnings about unknown attributes 2021-10-15 10:47:16 +02:00
Ryan Houdek eb8a8bf929 Merge pull request #1315 from lioncash/test
TestHarnessRunner: Make argument check more strict
2021-10-14 21:16:55 -07:00
Lioncash 17fd5f7f79 TestHarnessRunner: Make argument check more strict
Overlooked that more than one argument was being when replacing the
throw macro.
2021-10-15 00:06:30 -04:00
Ryan Houdek c9c352627f Merge pull request #1314 from lioncash/test
TestHarnessRunner: Convert LOGMAN_THROW_A into error log and exit
2021-10-14 18:58:46 -07:00
Lioncash e670f8f0e6 TestHarnessRunner: Convert logging calls over to fmt
Given we're in the same area, we may as well move things over to the
other logging system.
2021-10-14 21:45:03 -04:00
Lioncash c431cdebcc TestHarnessRunner: Convert LOGMAN_THROW_A into error log and exit
In release builds LOGMAN_THROW_A doesn't do anything, so running the
program without arguments would lead to a segfault.
2021-10-14 21:43:53 -04:00
Ryan Houdek 8b3c46154d Merge pull request #1312 from Sonicadvance1/JITSymbolsConfig
JITSymbols: Change over to runtime enablement of symbols
2021-10-13 18:09:32 -07:00
Ryan Houdek 1d9b66044a JITSymbols: Change over to runtime enablement of symbols
Adds a new option for just describing all JIT state as a single symbol.
Useful for simple profiling of total time spent in the JIT
2021-10-13 17:48:34 -07:00
Ryan Houdek 031fa8a7d6 Merge pull request #1311 from lioncash/op
OpcodeDispatcher: Deduplicate OpToIndex definition
2021-10-13 15:00:46 -07:00
Lioncash c9621da51c OpcodeDispatcher: Deduplicate OpToIndex definition
We can just make the one defined in X86Tables visible instead to keep
everything in one spot.
2021-10-13 16:18:39 -04:00
Ryan Houdek b1ab252c68 Merge pull request #1310 from lioncash/printf
DeadContextStoreElimination: Fix missing printf specifier entry
2021-10-13 10:51:49 -07:00
Ryan Houdek d09706aa1c Merge pull request #1309 from lioncash/tables
X86Tables: Make flag helper functions constexpr
2021-10-13 10:51:33 -07:00
Lioncash eb8ca16402 DeadContextStoreElimination: Fix missing printf specifier entry
Previously the offset mismatch error was expecting two arguments, but
only one was provided.

While we're in the area we can convert the logging type over to the
fmt-capable one which can catch these.
2021-10-13 12:43:03 -04:00
Lioncash 8df16460d1 X86Tables: Mark initialization instruction tables as static constexpr
While the previous change eliminated much of the codegen caused by
constructing everything individually on the stack, it didn't eliminate a
memcpy of all the elements onto the stack.

This eliminates the memcpys by allowing the compiler to place all the
data into RO and just reference that data.
2021-10-13 12:30:23 -04:00
Lioncash 5758c65983 X86Tables: Make flag helper functions constexpr
These only perform bit arithmetic, so we can allow them to be used in
constexpr contexts.

This allows clang to collapse quite a bit of code for the table
initializing functions. For example, in InitializeVEXTables(),
with these as inline (but not constexpr) functions, clang will
individually put all of the table entries onto the stack.

With these as constexpr functions, clang will be able to deduce that it
can construct the tables at compile time and reduces the amount of
generated code quite a bit.
2021-10-13 11:51:14 -04:00
Ryan Houdek fa1648c6d5 Merge pull request #1308 from Sonicadvance1/fix_missing_drm_include_path
Thunks: Fix missing libdrm include path
2021-10-11 23:17:03 -07:00
Ryan Houdek 98ba0bfa82 Merge pull request #1306 from Sonicadvance1/spill_fprs
Arm64: Make sure to spill static FPRs on guest signal
2021-10-11 23:16:51 -07:00
Ryan Houdek 366122338e Merge pull request #1305 from Sonicadvance1/spill_slot_debug
RAPass: Add debug compile option to disable spill slot reuse
2021-10-11 23:16:35 -07:00
Ryan Houdek bdc66a33ef Merge pull request #1304 from Sonicadvance1/explicit_x87_abi
Arm64: Be more explicit about x87 ABI usage
2021-10-11 23:16:07 -07:00
Ryan Houdek 48955da5f3 Thunks: Fix missing libdrm include path 2021-10-11 19:43:36 -07:00
Ryan Houdek 6cd73a6724 Merge pull request #1303 from Sonicadvance1/destdir_thunks
Thunks: Respect DESTDIR environment variable
2021-10-11 06:18:58 -07:00
Ryan Houdek 8dfe305aab Merge pull request #1302 from Sonicadvance1/missing_header_xcb
Thunks: XCB Add missing header file
2021-10-11 06:18:50 -07:00
Ryan Houdek 6fb0b3d85c Arm64: Make sure to spill FPRs on guest signal
This wasn't ever wired up
2021-10-10 22:17:41 -07:00
Ryan Houdek 69b27d7715 RAPass: Add debug compile option to disable spill slot reuse
Useful for debugging if spill slots are bugged
2021-10-10 20:45:21 -07:00
Ryan Houdek cffd10d0f7 Arm64: Be more explicit about x87 ABI usage
Just using zero extending moves to ensure that we don't fill any
register's upper bits with garbage
2021-10-10 20:43:27 -07:00
Ryan Houdek f2ef58630c Thunks: Respect DESTDIR environment variable
This allows local install to actually work
2021-10-08 20:57:23 -07:00
Ryan Houdek de8d8d8751 Thunks: XCB Add missing header file 2021-10-08 20:49:19 -07:00
108 changed files with 8722 additions and 6670 deletions

No files matched your search

+6 -6
View File
@@ -208,7 +208,7 @@ if (ENABLE_STATIC_PIE)
message (FATAL_ERROR "Application has __rela_iplt_{start,end} symbols. Which means static-pie can't be enabled")
endif()
else()
message (FATAL_ERROR "Couldn't compile static-pie test. Static-pie can't be enabled!")
message (FATAL_ERROR "Couldn't compile static-pie test. Static-pie can't be enabled! Is your glibc compiled without static-pie?")
endif()
endif()
@@ -300,11 +300,6 @@ if(ENUM_ENUM_WARNING)
add_compile_options(-Wno-deprecated-enum-enum-conversion)
endif()
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
@@ -336,6 +331,11 @@ if(_M_ARM_64)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
endif()
endif()
else()
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
endif()
endif()
if (ENABLE_IWYU)
-1
View File
@@ -16,7 +16,6 @@ endif()
set(ENABLE_JIT_X86_64 ${_M_X86_64} CACHE BOOL "Enable the x86_64 JIT")
set(ENABLE_JIT_ARM64 ${_M_ARM_64} CACHE BOOL "Enable the ARM64 JIT")
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_JITSYMBOLS "Enable visibility of JITSymbols in profiling tools" FALSE)
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
cmake_policy(SET CMP0083 NEW) # Follow new PIE policy
+20 -1
View File
@@ -374,7 +374,7 @@ def print_parse_argloader_options(options):
conversion_func = "std::to_string"
if ("ArgumentHandler" in op_vals):
NeedsString = True
conversion_func = "FEX::Handler::{0}".format(op_vals["ArgumentHandler"])
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
if (value_type == "str"):
NeedsString = True
conversion_func = ""
@@ -396,6 +396,21 @@ def print_parse_argloader_options(options):
output_argloader.write("#endif\n")
def print_parse_envloader_options(options):
output_argloader.write("#ifdef ENVLOADER\n")
output_argloader.write("#undef ENVLOADER\n")
output_argloader.write("if (false) {}\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
if ("ArgumentHandler" in op_vals):
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = {0}(Value);\n".format(conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def check_for_duplicate_options(options):
short_map = []
long_map = []
@@ -470,4 +485,8 @@ output_man.close()
output_argloader = open(output_argumentloader_filename, "w")
print_argloader_options(options);
print_parse_argloader_options(options);
# Generate environment loader code
print_parse_envloader_options(options);
output_argloader.close()
+12 -4
View File
@@ -97,6 +97,17 @@ set (SRCS
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -130,6 +141,7 @@ set (SRCS
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/FileLoading.cpp
Utils/LogManager.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
@@ -182,10 +194,6 @@ if (ENABLE_JIT_ARM64)
Interface/Core/JIT/Arm64/VectorOps.cpp)
endif()
if (ENABLE_JITSYMBOLS)
list(APPEND DEFINES -DENABLE_JITSYMBOLS=1)
endif()
set (LIBS vixl dl fmt::fmt xxhash tiny-json)
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
+21
View File
@@ -41,5 +41,26 @@ namespace FEXCore {
String << std::hex << HostAddr << " " << CodeSize << " " << Name << "_" << HostAddr << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
}
void JITSymbols::RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << Name << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
}
void JITSymbols::RegisterJITSpace(void *HostAddr, uint32_t CodeSize) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " FEXJIT" << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
}
}
+2
View File
@@ -10,6 +10,8 @@ public:
~JITSymbols();
void Register(void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterJITSpace(void *HostAddr, uint32_t CodeSize);
private:
FILE* fp{};
+12 -46
View File
@@ -1,5 +1,6 @@
#include "Common/StringConv.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
@@ -40,45 +41,6 @@ namespace DefaultValues {
#include <FEXCore/Config/ConfigValues.inl>
}
static bool LoadConfigFile(std::vector<char> &Data, const std::string &Config) {
std::fstream ConfigFile;
ConfigFile.open(Config, std::ios::in);
if (!ConfigFile.is_open()) {
return false;
}
if (!ConfigFile.seekg(0, std::fstream::end)) {
LogMan::Msg::D("Couldn't load configuration file: Seek end");
return false;
}
auto FileSize = ConfigFile.tellg();
if (ConfigFile.fail()) {
LogMan::Msg::D("Couldn't load configuration file: tellg");
return false;
}
if (!ConfigFile.seekg(0, std::fstream::beg)) {
LogMan::Msg::D("Couldn't load configuration file: Seek beginning");
return false;
}
if (FileSize > 0) {
Data.resize(FileSize);
if (!ConfigFile.read(&Data.at(0), FileSize)) {
// Probably means permissions aren't set. Just early exit
return false;
}
ConfigFile.close();
}
else {
return false;
}
return true;
}
namespace JSON {
struct JsonAllocator {
jsonPool_t PoolObject;
@@ -99,7 +61,7 @@ namespace JSON {
static void LoadJSonConfig(const std::string &Config, std::function<void(const char *Name, const char *ConfigSring)> Func) {
std::vector<char> Data;
if (!LoadConfigFile(Data, Config)) {
if (!FEXCore::FileLoading::LoadFile(Data, Config)) {
return;
}
@@ -436,7 +398,7 @@ namespace JSON {
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (LoadConfigFile(Manager, ContainerManager)) {
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
@@ -539,7 +501,7 @@ namespace JSON {
return Meta->Get(Option);
}
void Set(ConfigOption Option, std::string Data) {
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -547,7 +509,7 @@ namespace JSON {
Meta->Erase(Option);
}
void EraseSet(ConfigOption Option, std::string Data) {
void EraseSet(ConfigOption Option, std::string_view Data) {
Meta->EraseSet(Option, Data);
}
@@ -716,9 +678,13 @@ namespace JSON {
if (std::string::npos == pos)
continue;
std::string_view Ident = Var.substr(0,pos);
std::string_view Value = Var.substr(pos+1);
EnvMap[Ident]=Value;
std::string_view Key = Var.substr(0,pos);
std::string_view Value {Var.substr(pos+1)};
#define ENVLOADER
#include <FEXCore/Config/ConfigOptions.inl>
EnvMap[Key]=Value;
}
std::function GetVar = [=](const std::string_view id) -> std::optional<std::string_view> {
+27 -1
View File
@@ -163,8 +163,34 @@
"Potentially useful for debugging memory problems",
"32-bit allocator is always used if your host kernel is older than 4.17"
]
},
"GlobalJITNaming": {
"Type": "bool",
"Default": "false",
"Desc": [
"Uses JITSymbols to name all JIT state as one symbol",
"Useful for querying how much time is spent inside of the JIT",
"Profiling tools will show JIT time as FEXJIT"
]
},
"LibraryJITNaming": {
"Type": "bool",
"Default": "false",
"Desc": [
"Uses JITSymbols to name JIT symbols grouped by library",
"Useful for querying how much time is spent in each guest library",
"Can be used to help guide thunk generation"
]
},
"BlockJITNaming": {
"Type": "bool",
"Default": "false",
"Desc": [
"Uses JITSymbols to name JIT symbols",
"Useful for determining hot blocks of code",
"Has some file writing overhead per JIT block"
]
}
},
"Logging": {
"SilentLog": {
+8 -6
View File
@@ -1,5 +1,6 @@
#pragma once
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
@@ -12,9 +13,6 @@
#include <FEXCore/Utils/Event.h>
#include <stdint.h>
#ifdef ENABLE_JITSYMBOLS
#include <Common/JITSymbols.h>
#endif
#include <atomic>
#include <condition_variable>
@@ -127,6 +125,10 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(DumpIR, DUMPIR);
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
} Config;
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
@@ -175,9 +177,11 @@ namespace FEXCore::Context {
bool ContainsCode;
};
std::map<uint64_t, AddrToFileEntry> AddrToFile;
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
std::map<std::string, std::string> FilesWithCode;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
@@ -328,9 +332,7 @@ namespace FEXCore::Context {
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
#if ENABLE_JITSYMBOLS
FEXCore::JITSymbols Symbols;
#endif
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
+499 -273
View File
@@ -8,10 +8,15 @@
#include <stdint.h>
#include <signal.h>
#include "aarch64/cpu-aarch64.h"
namespace FEXCore::ArchHelpers::Arm64 {
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock, TYPE_HAS_SPLIT_LOCKS);
FEXCORE_TELEMETRY_STATIC_INIT(SplitLock16B, TYPE_16BYTE_SPLIT);
FEXCORE_TELEMETRY_STATIC_INIT(Cas16Tear, TYPE_CAS_16BIT_TEAR);
FEXCORE_TELEMETRY_STATIC_INIT(Cas32Tear, TYPE_CAS_32BIT_TEAR);
FEXCORE_TELEMETRY_STATIC_INIT(Cas64Tear, TYPE_CAS_64BIT_TEAR);
FEXCORE_TELEMETRY_STATIC_INIT(Cas128Tear, TYPE_CAS_128BIT_TEAR);
static __uint128_t LoadAcquire128(uint64_t Addr) {
__uint128_t Result{};
@@ -64,272 +69,6 @@ static bool StoreCAS8(uint8_t &Expected, uint8_t Val, uint64_t Addr) {
return Atom->compare_exchange_strong(Expected, Val);
}
static bool RunCASPAL(void *_ucontext, void *_info, uint32_t Size, uint32_t DesiredReg1, uint32_t DesiredReg2, uint32_t ExpectedReg1, uint32_t ExpectedReg2, uint32_t AddressReg) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
//Bus_ADRALN check happens in HandleCASPAL and HandleCASPAL_ARMv8
if (Size == 0) {
// 32bit
uint64_t Addr = mcontext->regs[AddressReg];
uint32_t DesiredLower = mcontext->regs[DesiredReg1];
uint32_t DesiredUpper = mcontext->regs[DesiredReg2];
uint32_t ExpectedLower = mcontext->regs[ExpectedReg1];
uint32_t ExpectedUpper = mcontext->regs[ExpectedReg2];
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
// Intel will do a "split lock" which locks the full bus
// AMD will tear instead
// Both cross-cacheline and cross 16byte both need dual CAS loops that can tear
// ARMv8.4 LSE2 solves all atomic issues except cross-cacheline
// Check for Split lock across a cacheline
if ((Addr & 63) > 56) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) > 8) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
uint64_t Alignment = Addr & 0b111;
Addr &= ~0b111ULL;
uint64_t AddrUpper = Addr + 8;
// Crosses a 16byte boundary
// Need to do 256bit atomic, but since that doesn't exist we need to do a dual CAS loop
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
__uint128_t NegMask = ~Mask;
__uint128_t TmpExpected{};
__uint128_t TmpDesired{};
__uint128_t Desired = DesiredUpper;
Desired <<= 32;
Desired |= DesiredLower;
Desired <<= Alignment * 8;
__uint128_t Expected = ExpectedUpper;
Expected <<= 32;
Expected |= ExpectedLower;
Expected <<= Alignment * 8;
while (1) {
__uint128_t LoadOrderUpper = LoadAcquire64(AddrUpper);
LoadOrderUpper <<= 64;
__uint128_t TmpActual = LoadOrderUpper | LoadAcquire64(Addr);
// Set up expected
TmpExpected = TmpActual;
TmpExpected &= NegMask;
TmpExpected |= Expected;
// Set up desired
TmpDesired = TmpExpected;
TmpDesired &= NegMask;
TmpDesired |= Desired;
uint64_t TmpExpectedLower = TmpExpected;
uint64_t TmpExpectedUpper = TmpExpected >> 64;
uint64_t TmpDesiredLower = TmpDesired;
uint64_t TmpDesiredUpper = TmpDesired >> 64;
if (TmpExpected == TmpActual) {
if (StoreCAS64(TmpExpectedUpper, TmpDesiredUpper, AddrUpper)) {
if (StoreCAS64(TmpExpectedLower, TmpDesiredLower, Addr)) {
// Stored successfully
return true;
}
else {
// CAS managed to tear, we can't really solve this
// Continue down the path to let the guest know values weren't expected
}
}
TmpExpected = TmpExpectedUpper;
TmpExpected <<= 64;
TmpExpected |= TmpExpectedLower;
}
else {
// Mismatch up front
TmpExpected = TmpActual;
}
// Not successful
// Now we need to check the results to see if we need to try again
__uint128_t FailedResultOurBits = TmpExpected & Mask;
__uint128_t FailedResultNotOurBits = TmpExpected & NegMask;
__uint128_t FailedDesiredOurBits = TmpDesired & Mask;
__uint128_t FailedDesiredNotOurBits = TmpDesired & NegMask;
if ((FailedResultNotOurBits ^ FailedDesiredNotOurBits) != 0) {
// If the bits changed that weren't part of our regular CAS then we need to try again
continue;
}
if ((FailedResultOurBits ^ FailedDesiredOurBits) != 0) {
// If the bits changed that we were wanting to change then we have failed and can return
// We need to extract the bits and return them in EXPECTED
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
// This happens in the case that between Load and CAS that something has store our desired in to the memory location
// This means our CAS fails because what we wanted to store was already stored
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
}
else {
// Fits within a 16byte region
uint64_t Alignment = Addr & 0b1111;
Addr &= ~0b1111ULL;
std::atomic<__uint128_t> *Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
__uint128_t NegMask = ~Mask;
__uint128_t TmpExpected{};
__uint128_t TmpDesired{};
__uint128_t Desired = (uint64_t)DesiredUpper << 32 | DesiredLower;
Desired <<= Alignment * 8;
__uint128_t Expected = (uint64_t)ExpectedUpper << 32 | ExpectedLower;
Expected <<= Alignment * 8;
while (1) {
TmpExpected = Atomic128->load();
// Set up expected
TmpExpected &= NegMask;
TmpExpected |= Expected;
// Set up desired
TmpDesired = TmpExpected;
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return true;
}
else {
// Not successful
// Now we need to check the results to see if we need to try again
__uint128_t FailedResultOurBits = TmpExpected & Mask;
__uint128_t FailedResultNotOurBits = TmpExpected & NegMask;
__uint128_t FailedDesiredNotOurBits = TmpDesired & NegMask;
if ((FailedResultNotOurBits ^ FailedDesiredNotOurBits) != 0) {
// If the bits changed that weren't part of our regular CAS then we need to try again
continue;
}
// This happens in the case that between Load and CAS that something has store our desired in to the memory location
// This means our CAS fails because what we wanted to store was already stored
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
}
}
}
return false;
}
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = (Instr >> 30) & 1;
uint32_t DesiredReg1 = Instr & 0b11111;
uint32_t DesiredReg2 = DesiredReg1 + 1;
uint32_t ExpectedReg1 = (Instr >> 16) & 0b11111;
uint32_t ExpectedReg2 = ExpectedReg1 + 1;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
return RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, ExpectedReg1, ExpectedReg2, AddressReg);
}
uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return 0;
}
// caspair
// [1] ldaxp(TMP2.W(), TMP3.W(), MemOperand(MemSrc)); <-- DataReg & AddrReg
// [2] cmp(TMP2.W(), Expected.first.W()); <-- ExpectedReg1
// [3] ccmp(TMP3.W(), Expected.second.W(), NoFlag, Condition::eq); <-- ExpectedREg2
// [4] b(&LoopNotExpected, Condition::ne);
// [5] stlxp(TMP2.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc)); <-- DesiredReg
// [6] cbnz(TMP2.W(), &LoopTop);
// [7] mov(Dst.first.W(), Expected.first.W());
// [8] mov(Dst.second.W(), Expected.second.W());
// [9] b(&LoopExpected);
// [10] mov(Dst.first.W(), TMP2.W());
// [11] mov(Dst.second.W(), TMP3.W());
// [12] clrex();
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Size = (Instr >> 30) & 1;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
uint32_t ExpectedReg1{};
uint32_t ExpectedReg2{};
uint32_t DesiredReg1{};
uint32_t DesiredReg2{};
if(Size != 0) { //Only 32-bit pairs
return 0;
}
for(int i = 1; i < 10; i++) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
ExpectedReg1 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::CCMP_MASK) == FEXCore::ArchHelpers::Arm64::CCMP_INST) {
ExpectedReg2 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) {
DesiredReg1 = (NextInstr & 0x1F);
DesiredReg2 = (NextInstr >> 10) & 0x1F;
}
}
//mov expected into the temp registers used by JIT
mcontext->regs[DataReg] = mcontext->regs[ExpectedReg1];
mcontext->regs[DataReg2] = mcontext->regs[ExpectedReg2];
if(RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, DataReg, DataReg2, AddrReg)) {
return 9 * sizeof(uint32_t); // skip to mov + clrex
} else {
return 0;
}
}
uint16_t DoLoad16(uint64_t Addr) {
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) == 15) {
@@ -499,6 +238,347 @@ std::pair<uint64_t, uint64_t> DoLoad128(uint64_t Addr) {
return {ResultLower, ResultUpper};
}
static bool RunCASPAL(void *_ucontext, void *_info, uint32_t Size, uint32_t DesiredReg1, uint32_t DesiredReg2, uint32_t ExpectedReg1, uint32_t ExpectedReg2, uint32_t AddressReg) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
//Bus_ADRALN check happens in HandleCASPAL and HandleCASPAL_ARMv8
if (Size == 0) {
// 32bit
uint64_t Addr = mcontext->regs[AddressReg];
uint32_t DesiredLower = mcontext->regs[DesiredReg1];
uint32_t DesiredUpper = mcontext->regs[DesiredReg2];
uint32_t ExpectedLower = mcontext->regs[ExpectedReg1];
uint32_t ExpectedUpper = mcontext->regs[ExpectedReg2];
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
// Intel will do a "split lock" which locks the full bus
// AMD will tear instead
// Both cross-cacheline and cross 16byte both need dual CAS loops that can tear
// ARMv8.4 LSE2 solves all atomic issues except cross-cacheline
// Check for Split lock across a cacheline
if ((Addr & 63) > 56) {
FEXCORE_TELEMETRY_SET(SplitLock, 1);
}
uint64_t AlignmentMask = 0b1111;
if ((Addr & AlignmentMask) > 8) {
FEXCORE_TELEMETRY_SET(SplitLock16B, 1);
uint64_t Alignment = Addr & 0b111;
Addr &= ~0b111ULL;
uint64_t AddrUpper = Addr + 8;
// Crosses a 16byte boundary
// Need to do 256bit atomic, but since that doesn't exist we need to do a dual CAS loop
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
__uint128_t NegMask = ~Mask;
__uint128_t TmpExpected{};
__uint128_t TmpDesired{};
__uint128_t Desired = DesiredUpper;
Desired <<= 32;
Desired |= DesiredLower;
Desired <<= Alignment * 8;
__uint128_t Expected = ExpectedUpper;
Expected <<= 32;
Expected |= ExpectedLower;
Expected <<= Alignment * 8;
while (1) {
__uint128_t LoadOrderUpper = LoadAcquire64(AddrUpper);
LoadOrderUpper <<= 64;
__uint128_t TmpActual = LoadOrderUpper | LoadAcquire64(Addr);
// Set up expected
TmpExpected = TmpActual;
TmpExpected &= NegMask;
TmpExpected |= Expected;
// Set up desired
TmpDesired = TmpExpected;
TmpDesired &= NegMask;
TmpDesired |= Desired;
uint64_t TmpExpectedLower = TmpExpected;
uint64_t TmpExpectedUpper = TmpExpected >> 64;
uint64_t TmpDesiredLower = TmpDesired;
uint64_t TmpDesiredUpper = TmpDesired >> 64;
if (TmpExpected == TmpActual) {
if (StoreCAS64(TmpExpectedUpper, TmpDesiredUpper, AddrUpper)) {
if (StoreCAS64(TmpExpectedLower, TmpDesiredLower, Addr)) {
// Stored successfully
return true;
}
else {
// CAS managed to tear, we can't really solve this
// Continue down the path to let the guest know values weren't expected
FEXCORE_TELEMETRY_SET(Cas128Tear, 1);
}
}
TmpExpected = TmpExpectedUpper;
TmpExpected <<= 64;
TmpExpected |= TmpExpectedLower;
}
else {
// Mismatch up front
TmpExpected = TmpActual;
}
// Not successful
// Now we need to check the results to see if we need to try again
__uint128_t FailedResultOurBits = TmpExpected & Mask;
__uint128_t FailedResultNotOurBits = TmpExpected & NegMask;
__uint128_t FailedDesiredOurBits = TmpDesired & Mask;
__uint128_t FailedDesiredNotOurBits = TmpDesired & NegMask;
if ((FailedResultNotOurBits ^ FailedDesiredNotOurBits) != 0) {
// If the bits changed that weren't part of our regular CAS then we need to try again
continue;
}
if ((FailedResultOurBits ^ FailedDesiredOurBits) != 0) {
// If the bits changed that we were wanting to change then we have failed and can return
// We need to extract the bits and return them in EXPECTED
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
// This happens in the case that between Load and CAS that something has store our desired in to the memory location
// This means our CAS fails because what we wanted to store was already stored
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
}
else {
// Fits within a 16byte region
uint64_t Alignment = Addr & 0b1111;
Addr &= ~0b1111ULL;
std::atomic<__uint128_t> *Atomic128 = reinterpret_cast<std::atomic<__uint128_t>*>(Addr);
__uint128_t Mask = ~0ULL;
Mask <<= Alignment * 8;
__uint128_t NegMask = ~Mask;
__uint128_t TmpExpected{};
__uint128_t TmpDesired{};
__uint128_t Desired = (uint64_t)DesiredUpper << 32 | DesiredLower;
Desired <<= Alignment * 8;
__uint128_t Expected = (uint64_t)ExpectedUpper << 32 | ExpectedLower;
Expected <<= Alignment * 8;
while (1) {
TmpExpected = Atomic128->load();
// Set up expected
TmpExpected &= NegMask;
TmpExpected |= Expected;
// Set up desired
TmpDesired = TmpExpected;
TmpDesired &= NegMask;
TmpDesired |= Desired;
bool CASResult = Atomic128->compare_exchange_strong(TmpExpected, TmpDesired);
if (CASResult) {
// Successful, so we are done
return true;
}
else {
// Not successful
// Now we need to check the results to see if we need to try again
__uint128_t FailedResultOurBits = TmpExpected & Mask;
__uint128_t FailedResultNotOurBits = TmpExpected & NegMask;
__uint128_t FailedDesiredNotOurBits = TmpDesired & NegMask;
if ((FailedResultNotOurBits ^ FailedDesiredNotOurBits) != 0) {
// If the bits changed that weren't part of our regular CAS then we need to try again
continue;
}
// This happens in the case that between Load and CAS that something has store our desired in to the memory location
// This means our CAS fails because what we wanted to store was already stored
uint64_t FailedResult = FailedResultOurBits >> (Alignment * 8);
mcontext->regs[ExpectedReg1] = FailedResult & ~0U;
mcontext->regs[ExpectedReg2] = FailedResult >> 32;
return true;
}
}
}
}
return false;
}
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return false;
}
uint32_t Size = (Instr >> 30) & 1;
uint32_t DesiredReg1 = Instr & 0b11111;
uint32_t DesiredReg2 = DesiredReg1 + 1;
uint32_t ExpectedReg1 = (Instr >> 16) & 0b11111;
uint32_t ExpectedReg2 = ExpectedReg1 + 1;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
return RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, ExpectedReg1, ExpectedReg2, AddressReg);
}
uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return 0;
}
// caspair
// [1] ldaxp(TMP2.W(), TMP3.W(), MemOperand(MemSrc)); <-- DataReg & AddrReg
// [2] cmp(TMP2.W(), Expected.first.W()); <-- ExpectedReg1
// [3] ccmp(TMP3.W(), Expected.second.W(), NoFlag, Condition::eq); <-- ExpectedREg2
// [4] b(&LoopNotExpected, Condition::ne);
// [5] stlxp(TMP2.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc)); <-- DesiredReg
// [6] cbnz(TMP2.W(), &LoopTop);
// [7] mov(Dst.first.W(), Expected.first.W());
// [8] mov(Dst.second.W(), Expected.second.W());
// [9] b(&LoopExpected);
// [10] mov(Dst.first.W(), TMP2.W());
// [11] mov(Dst.second.W(), TMP3.W());
// [12] clrex();
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Size = (Instr >> 30) & 1;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
uint32_t ExpectedReg1{};
uint32_t ExpectedReg2{};
uint32_t DesiredReg1{};
uint32_t DesiredReg2{};
if(Size == 1) {
// 64-bit pair happens on paranoid vector loads
// [1] ldaxp(TMP1, TMP2, MemSrc);
// [2] clrex();
//
// 64-bit pair happens on paranoid vector stores
// [1] ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS
// [2] stlxp(TMP3, TMP1, TMP2, MemSrc); // <- Can also hit SIGBUS
// [3] cbnz(TMP3, &B); // < Overwritten with DMB
if (DataReg == 31) {
}
else {
uint32_t NextInstr = PC[1];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::CLREX_MASK) == FEXCore::ArchHelpers::Arm64::CLREX_INST) {
uint64_t Addr = mcontext->regs[AddrReg];
auto Res = DoLoad128(Addr);
// We set the result register if it isn't a zero register
if (DataReg != 31) {
mcontext->regs[DataReg] = std::get<0>(Res);
}
if (DataReg2 != 31) {
mcontext->regs[DataReg2] = std::get<1>(Res);
}
// Skip ldaxp and clrex
return 2 * sizeof(uint32_t);
}
}
return 0;
}
//Only 32-bit pairs
for(int i = 1; i < 10; i++) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
ExpectedReg1 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::CCMP_MASK) == FEXCore::ArchHelpers::Arm64::CCMP_INST) {
ExpectedReg2 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) {
DesiredReg1 = (NextInstr & 0x1F);
DesiredReg2 = (NextInstr >> 10) & 0x1F;
}
}
//mov expected into the temp registers used by JIT
mcontext->regs[DataReg] = mcontext->regs[ExpectedReg1];
mcontext->regs[DataReg2] = mcontext->regs[ExpectedReg2];
if(RunCASPAL(_ucontext, _info, Size, DesiredReg1, DesiredReg2, DataReg, DataReg2, AddrReg)) {
return 9 * sizeof(uint32_t); // skip to mov + clrex
} else {
return 0;
}
}
bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
// This only handles alignment problems
return 0;
}
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Size = (Instr >> 30) & 1;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
if(Size == 1) {
// 64-bit pair happens on paranoid vector stores
// [0] ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS. Overwritten with DMB
// [1] stlxp(TMP3, TMP1, TMP2, MemSrc); // <- Can also hit SIGBUS
// [2] cbnz(TMP3, &B); // < Overwritten with DMB
if (DataReg == 31) {
uint32_t NextInstr = PC[1];
AddrReg = (NextInstr >> 5) & 0x1F;
DataReg = NextInstr & 0x1F;
DataReg2 = (NextInstr >> 10) & 0x1F;
uint32_t STP =
(0b10 << 30) |
(0b101001000000000 << 15) |
(DataReg2 << 10) |
(AddrReg << 5) |
DataReg;
PC[0] = DMB;
PC[1] = STP;
PC[2] = DMB;
// Back up one instruction and have another go
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[0], 16);
return true;
}
}
return false;
}
template <typename T>
using CASExpectedFn = T (*)(T Src, T Expected);
template <typename T>
@@ -557,6 +637,7 @@ uint16_t DoCAS16(
// CAS managed to tear, we can't really solve this
// Continue down the path to let the guest know values weren't expected
Tear = true;
FEXCORE_TELEMETRY_SET(Cas16Tear, 1);
}
}
@@ -850,6 +931,7 @@ uint32_t DoCAS32(
// CAS managed to tear, we can't really solve this
// Continue down the path to let the guest know values weren't expected
Tear = true;
FEXCORE_TELEMETRY_SET(Cas32Tear, 1);
}
}
@@ -1089,6 +1171,7 @@ uint64_t DoCAS64(
// CAS managed to tear, we can't really solve this
// Continue down the path to let the guest know values weren't expected
Tear = true;
FEXCORE_TELEMETRY_SET(Cas64Tear, 1);
}
}
@@ -1196,7 +1279,6 @@ uint64_t DoCAS64(
static bool RunCASAL(void *_ucontext, void *_info, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
uint64_t Addr = mcontext->regs[AddressReg];
@@ -1278,7 +1360,6 @@ static bool RunCASAL(void *_ucontext, void *_info, uint32_t Size, uint32_t Desir
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -1639,8 +1720,7 @@ bool HandleAtomicLoad128(void *_ucontext, void *_info, uint32_t Instr) {
static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
{
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
// ARMv8.0 CAS
// [1] ldaxrb(TMP2.W(), MemOperand(MemSrc))
// [2] cmp (TMP2.W(), Expected.W())
@@ -1651,12 +1731,12 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
// [7] b
// [8] mov (.., TMP2.W());
// [9] clrex
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(_ucontext);
uint32_t Instr = PC[0];
uint32_t Size = 1 << (Instr >> 30);
uint32_t AddressReg = GetRnReg(Instr);
uint32_t ResultReg = GetRdReg(Instr); //TMP2
uint32_t ResultReg = GetRdReg(Instr); //TMP2
uint32_t DesiredReg = 0;
uint32_t ExpectedReg = 0;
for (size_t i = 1; i < 6; ++i) {
@@ -1675,7 +1755,7 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
}
//set up CASAL by doing mov(TMP2, Expected)
mcontext->regs[ResultReg] = mcontext->regs[ExpectedReg];
if(RunCASAL(_ucontext, _info, Size, DesiredReg, ResultReg, AddressReg)) {
return 7 * sizeof(uint32_t); //jump to mov to allocated register
} else {
@@ -2047,4 +2127,150 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
return NumInstructionsToSkip * 4;
}
bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
#ifdef _M_ARM_64
constexpr bool is_arm64 = true;
#else
constexpr bool is_arm64 = false;
#endif
if constexpr (is_arm64) {
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(ucontext);
uint32_t Instr = PC[0];
// 1 = 16bit
// 2 = 32bit
// 3 = 64bit
uint32_t Size = (Instr & 0xC000'0000) >> 30;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
if ((Instr & 0x3F'FF'FC'00) == 0x08'DF'FC'00 || // LDAR*
(Instr & 0x3F'FF'FC'00) == 0x38'BF'C0'00) { // LDAPR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t LDR = 0b0011'1000'0111'1111'0110'1000'0000'0000;
LDR |= Size << 30;
LDR |= AddrReg << 5;
LDR |= DataReg;
PC[-1] = DMB;
PC[0] = LDR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ( (Instr & 0x3F'FF'FC'00) == 0x08'9F'FC'00) { // STLR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t STR = 0b0011'1000'0011'1111'0110'1000'0000'0000;
STR |= Size << 30;
STR |= AddrReg << 5;
STR |= DataReg;
PC[-1] = DMB;
PC[0] = STR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXP_MASK) == FEXCore::ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
//Should be compare and swap pair only. LDAXP not used elsewhere
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleCASPAL_ARMv8(ucontext, info, Instr);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicVectorStore(ucontext, info, Instr)) {
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) { // STLXP
//Should not trigger - middle of an LDAXP/STAXP pair.
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASAL: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: {} Instruction: 0x{:08x}\n", Op, fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXR_MASK) == FEXCore::ArchHelpers::Arm64::LDAXR_INST) { // LDAXR*
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleAtomicLoadstoreExclusive(ucontext, info);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXR: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[-1], 16);
return true;
}
return false;
}
}
+10 -2
View File
@@ -34,10 +34,13 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t CCMP_MASK = 0x7F'E0'0C'10;
constexpr uint32_t CCMP_INST = 0x7A'40'00'00;
constexpr uint32_t CLREX_MASK = 0xFF'FF'F0'FF;
constexpr uint32_t CLREX_INST = 0xD5'03'30'5F;
enum ExclusiveAtomicPairType {
TYPE_SWAP,
TYPE_ADD,
@@ -65,6 +68,9 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t RN_OFFSET = 5;
constexpr uint32_t RM_OFFSET = 16;
constexpr uint32_t DMB = 0b1101'0101'0000'0011'0011'0000'1011'1111 |
0b1011'0000'0000; // Inner shareable all
inline uint32_t GetRdReg(uint32_t Instr) {
return (Instr >> RD_OFFSET) & REGISTER_MASK;
}
@@ -83,6 +89,8 @@ namespace FEXCore::ArchHelpers::Arm64 {
uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr);
[[nodiscard]] bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext);
}
+24 -12
View File
@@ -59,6 +59,20 @@ static inline mcontext_t* GetMContext(void* ucontext) {
#ifdef _M_ARM_64
constexpr uint32_t FPR_MAGIC = 0x46508001U;
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->sp;
}
@@ -91,19 +105,13 @@ static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
GetMContext(ucontext)->regs[id] = val;
}
constexpr uint32_t FPR_MAGIC = 0x46508001U;
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
auto MContext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&MContext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
struct HostCTXHeader {
uint32_t Magic;
uint32_t Size;
};
struct HostFPRState {
HostCTXHeader Head;
uint32_t FPSR;
uint32_t FPCR;
__uint128_t FPRs[32];
};
return HostState->FPRs[id];
}
using ContextBackup = ArmContextBackup;
template <typename T>
@@ -192,6 +200,10 @@ static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
ERROR_AND_DIE("Not impelented for x86 host");
}
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not implemented for x86 host");
}
using ContextBackup = X86ContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
+90 -14
View File
@@ -7,9 +7,12 @@ $end_info$
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include "Common/StringConv.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Utils/FileLoading.h"
#include "git_version.h"
#include <cstring>
@@ -45,6 +48,63 @@ static uint32_t GetCycleCounterFrequency() {
: [Res] "=r" (Result));
return Result;
}
static bool GetHostHybridFlag() {
int MaxCPUs = 64;
size_t AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
cpu_set_t *Set = CPU_ALLOC(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
int Result{};
for (;;) {
Result = sched_getaffinity(0, AllocSize, Set);
if (Result == 0 ||
(Result == -1 && errno != EINVAL)) {
break;
}
MaxCPUs <<= 1;
CPU_FREE(Set);
Set = CPU_ALLOC(MaxCPUs);
AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
}
if (Result != 0) {
return false;
}
int CPUs = CPU_COUNT_S(AllocSize, Set);
bool Hybrid = false;
uint64_t MIDR{};
for (int i = 0; i < CPUs; ++i) {
if (CPU_ISSET_S(i, AllocSize, Set)) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
break;
}
MIDR = NewMIDR;
}
}
}
}
}
CPU_FREE(Set);
return Hybrid;
}
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t eax, ebx, ecx, edx;
@@ -58,6 +118,19 @@ static uint32_t GetCycleCounterFrequency() {
}
return 0;
}
static bool GetHostHybridFlag() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x7) {
__cpuid(0x7, eax, ebx, ecx, edx);
// Bit 15 of edx claims hybrid CPU
return (edx & (1U << 15)) != 0;
}
return false;
}
#endif
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
@@ -308,7 +381,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(1 << 0) | // FS/GS support
(0 << 1) | // TSC adjust MSR
(0 << 2) | // SGX
(0 << 3) | // BMI1
(1 << 3) | // BMI1
(0 << 4) | // Intel Hardware Lock Elison
(0 << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
@@ -382,29 +455,29 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 6) | // Reserved
(0 << 7) | // Reserved
(0 << 8) | // AVX512_VP2INTERSECT
(0 << 9) | // Reserved
(0 << 9) | // SRBDS_CTRL (Special Register Buffer Data Sampling Mitigations)
(0 << 10) | // VERW clears CPU buffers
(0 << 11) | // Reserved
(0 << 12) | // Reserved
(0 << 13) | // Reserved
(0 << 13) | // TSX Force Abort (TSX will force abort if attempted)
(0 << 14) | // SERIALIZE instruction
(0 << 15) | // Reserved
(0 << 16) | // Reserved
((Hybrid ? 1U : 0U) << 15) | // Hybrid
(0 << 16) | // TSXLDTRK (TSX Suspend load address tracking) - Allows untracked memory loads inside TSX region
(0 << 17) | // Reserved
(0 << 18) | // Intel PCONFIG
(0 << 19) | // Intel Architectural LBR
(0 << 20) | // Intel CET
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(0 << 23) | // Reserved
(0 << 24) | // Reserved
(0 << 25) | // Reserved
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 22) | // AMX-BF16 - Tile computation on bfloat16
(0 << 23) | // AVX512_FP16 - FP16 AVX512 instructions
(0 << 24) | // AMX-tile - If AMX is implemented
(0 << 25) | // AMX-int8 - AMX on 8-bit integers
(0 << 26) | // IBRS_IBPB - Speculation control
(0 << 27) | // STIBP - Single Thread Indirect Branch Predictor, Part of IBC
(0 << 28) | // L1D Flush
(0 << 29) | // Arch capabilities
(0 << 30) | // Reserved
(0 << 31); // Reserved
(0 << 29) | // Arch capabilities - Speculative side channel mitigations
(0 << 30) | // Arch capabilities - MSR module specific
(0 << 31); // SSBD - Speculative Store Bypass Disable
}
return Res;
@@ -885,6 +958,9 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#endif
// 0x8000'001E: Extended APIC ID
// 0x8000'001F: AMD Secure Encryption
// Setup some state tracking
Hybrid = GetHostHybridFlag();
}
}
+1
View File
@@ -37,6 +37,7 @@ public:
}
private:
FEXCore::Context::Context *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
using FunctionHandler = std::function<FEXCore::CPUID::FunctionResults(uint32_t Leaf)>;
+64 -42
View File
@@ -249,6 +249,8 @@ namespace FEXCore::Context {
LocalLoader = Loader;
using namespace FEXCore::Core;
FEXCore::CPU::InitializeInterpreterOpHandlers();
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
@@ -1141,6 +1143,19 @@ namespace FEXCore::Context {
}
}
Context::AddrToFileMapType::iterator Context::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
@@ -1200,61 +1215,68 @@ namespace FEXCore::Context {
}
// The core managed to compile the code.
#if ENABLE_JITSYMBOLS
if (DebugData) {
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
Symbols.Register((void*)Subblock.HostCodeStart, GuestRIP, Subblock.HostCodeSize);
if (Config.BlockJITNaming()) {
if (DebugData) {
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
Symbols.Register((void*)Subblock.HostCodeStart, GuestRIP, Subblock.HostCodeSize);
}
} else {
Symbols.Register(CodePtr, GuestRIP, DebugData->HostCodeSize);
}
} else {
Symbols.Register(CodePtr, GuestRIP, DebugData->HostCodeSize);
}
}
#endif
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to AOT cache if aot generation is enabled
if ((Config.AOTIRCapture() || Config.AOTIRGenerate()) && RAData) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
std::shared_lock lk(AOTIRCacheLock);
auto file = FindAddrForFile(StartAddr, Length);
auto file = AddrToFile.lower_bound(StartAddr);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= StartAddr && (file->second.Start + file->second.Len) >= (StartAddr + Length)) {
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCache[fileid];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
uint64_t tag = 0xDEADBEEFC0D30004;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
});
}
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (DebugData && Config.LibraryJITNaming()) {
Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
}
if (Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && RAData &&
(Config.AOTIRCapture() || Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCache[fileid];
Thread->CPUBackend->ClearCache();
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
uint64_t tag = 0xDEADBEEFC0D30004;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
});
return (uintptr_t)CodePtr;
if (Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
Thread->CPUBackend->ClearCache();
return (uintptr_t)CodePtr;
}
}
}
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
}
if (DecrementRefCount)
@@ -1384,7 +1406,7 @@ namespace FEXCore::Context {
// TODO: Support overlapping maps and region splitting
auto base_filename = std::filesystem::path(filename).filename().string();
if (base_filename.size()) {
if (!base_filename.empty()) {
auto filename_hash = XXH3_64bits(filename.c_str(), filename.size());
auto fileid = base_filename + "-" + std::to_string(filename_hash) + "-";
@@ -25,9 +25,7 @@
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#ifdef ENABLE_JITSYMBOLS
#include <unistd.h>
#endif
namespace FEXCore::CPU {
@@ -342,24 +340,24 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
#if ENABLE_JITSYMBOLS
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
#endif
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
}
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext) {
for(int i = 0; i < SRA64.size(); i++) {
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
// TODO: Also recover FPRs, not sure where the neon context is
// This is usually not needed
/*
for(int i = 0; i < SRAFPR.size(); i++) {
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
State->State.State.xmm[i][0] = _mcontext.neon[SRAFPR[i].GetCode()];
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&ThreadState->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
}
*/
}
#ifdef _M_ARM_64
@@ -306,10 +306,13 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
Start = reinterpret_cast<uint64_t>(getCode());
End = Start + getSize();
#if ENABLE_JITSYMBOLS
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(Start), End-Start, Name);
#endif
}
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(Start), End-Start);
}
}
X86Dispatcher::~X86Dispatcher() {
+83 -11
View File
@@ -127,6 +127,54 @@ static uint32_t MapModRMToReg(uint8_t REX, uint8_t bits, bool HighBits, bool Has
return (*GPRs)[(REX << 3) | bits];
}
static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
using GPRArray = std::array<uint32_t, 16>;
static constexpr GPRArray GPRIndexes = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_RCX,
FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_RBX,
FEXCore::X86State::REG_RSP,
FEXCore::X86State::REG_RBP,
FEXCore::X86State::REG_RSI,
FEXCore::X86State::REG_RDI,
FEXCore::X86State::REG_R8,
FEXCore::X86State::REG_R9,
FEXCore::X86State::REG_R10,
FEXCore::X86State::REG_R11,
FEXCore::X86State::REG_R12,
FEXCore::X86State::REG_R13,
FEXCore::X86State::REG_R14,
FEXCore::X86State::REG_R15,
};
static constexpr GPRArray XMMIndexes = {
FEXCore::X86State::REG_XMM_0,
FEXCore::X86State::REG_XMM_1,
FEXCore::X86State::REG_XMM_2,
FEXCore::X86State::REG_XMM_3,
FEXCore::X86State::REG_XMM_4,
FEXCore::X86State::REG_XMM_5,
FEXCore::X86State::REG_XMM_6,
FEXCore::X86State::REG_XMM_7,
FEXCore::X86State::REG_XMM_8,
FEXCore::X86State::REG_XMM_9,
FEXCore::X86State::REG_XMM_10,
FEXCore::X86State::REG_XMM_11,
FEXCore::X86State::REG_XMM_12,
FEXCore::X86State::REG_XMM_13,
FEXCore::X86State::REG_XMM_14,
FEXCore::X86State::REG_XMM_15,
};
if (HasXMM) {
return XMMIndexes[vvvv];
} else {
return GPRIndexes[vvvv];
}
}
Decoder::Decoder(FEXCore::Context::Context *ctx)
: CTX {ctx}
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN } {
@@ -343,7 +391,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
}
bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op) {
bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op, DecodedHeader Options) {
DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
@@ -367,8 +415,9 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
"Group Ops should have been decoded before this!");
uint8_t DestSize{};
bool HasWideningDisplacement = FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST;
bool HasNarrowingDisplacement = FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST;
const bool HasWideningDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST) != 0 ||
Options.w;
const bool HasNarrowingDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
bool HasXMMSrc = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_GPR) &&
@@ -401,8 +450,8 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
// New instruction size decoding
{
// Decode destinations first
uint32_t DstSizeFlag = FEXCore::X86Tables::InstFlags::GetSizeDstFlags(Info->Flags);
uint32_t SrcSizeFlag = FEXCore::X86Tables::InstFlags::GetSizeSrcFlags(Info->Flags);
const auto DstSizeFlag = FEXCore::X86Tables::InstFlags::GetSizeDstFlags(Info->Flags);
const auto SrcSizeFlag = FEXCore::X86Tables::InstFlags::GetSizeSrcFlags(Info->Flags);
if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_8BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_8BIT);
@@ -546,6 +595,13 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
size_t CurrentSrc = 0;
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) != 0) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMSrc);
++CurrentSrc;
}
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM) {
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SF_MOD_DST) {
if (!ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest))
@@ -558,6 +614,13 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
++CurrentSrc;
}
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_2ND_SRC) != 0) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMSrc);
++CurrentSrc;
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_RAX)) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
@@ -571,6 +634,12 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
++CurrentSrc;
}
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_DST) != 0) {
CurrentDest->Type = DecodedOperand::OpType::GPR;
CurrentDest->Data.GPR.HighBits = false;
CurrentDest->Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMDst);
}
if (Bytes != 0) {
LOGMAN_THROW_A(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
@@ -703,16 +772,19 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
uint16_t map_select = 1;
uint16_t pp = 0;
uint8_t Byte1 = ReadByte();
const uint8_t Byte1 = ReadByte();
DecodedHeader options{};
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
}
else { // 0xC4 = Three byte VEX
uint8_t Byte2 = ReadByte();
const uint8_t Byte2 = ReadByte();
pp = Byte2 & 0b11;
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
options.w = (Byte2 & 0b10000000) != 0;
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::E("We don't understand a map_select of: %d", map_select);
return false;
@@ -740,10 +812,10 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
Op = OPD(LocalInfo->Type, pp, ModRM.reg);
#undef OPD
return NormalOp(&VEXTableGroupOps[Op], Op);
return NormalOp(&VEXTableGroupOps[Op], Op, options);
} else {
return NormalOp(LocalInfo, Op, options);
}
else
return NormalOp(LocalInfo, Op);
}
else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
FEXCORE_TELEMETRY_SET(EVEXOpTelem, 1);
+9 -1
View File
@@ -39,6 +39,13 @@ public:
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
private:
// To pass any information from instruction prefixes
// down into the actual instruction handling machinery.
struct DecodedHeader {
uint8_t vvvv; // Encoded operand in a VEX prefix.
bool w; // VEX.W bit.
};
FEXCore::Context::Context *CTX;
const FEXCore::HLE::SyscallOSABI OSABI{};
@@ -50,7 +57,8 @@ private:
uint8_t PeekByte(uint8_t Offset) const;
uint64_t ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) { InstructionSize += Size; }
bool NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op);
bool NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op);
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,796 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <FEXCore/Utils/BitUtils.h>
#include <cstdint>
namespace FEXCore::CPU {
#ifdef _M_X86_64
uint8_t AtomicFetchNeg(uint8_t *Addr) {
using Type = uint8_t;
std::atomic<Type> *MemData = reinterpret_cast<std::atomic<Type>*>(Addr);
Type Expected = MemData->load();
Type Desired = -Expected;
do {
Desired = -Expected;
} while (!MemData->compare_exchange_strong(Expected, Desired, std::memory_order_seq_cst));
return Expected;
}
uint16_t AtomicFetchNeg(uint16_t *Addr) {
using Type = uint16_t;
std::atomic<Type> *MemData = reinterpret_cast<std::atomic<Type>*>(Addr);
Type Expected = MemData->load();
Type Desired = -Expected;
do {
Desired = -Expected;
} while (!MemData->compare_exchange_strong(Expected, Desired, std::memory_order_seq_cst));
return Expected;
}
uint32_t AtomicFetchNeg(uint32_t *Addr) {
using Type = uint32_t;
std::atomic<Type> *MemData = reinterpret_cast<std::atomic<Type>*>(Addr);
Type Expected = MemData->load();
Type Desired = -Expected;
do {
Desired = -Expected;
} while (!MemData->compare_exchange_strong(Expected, Desired, std::memory_order_seq_cst));
return Expected;
}
uint64_t AtomicFetchNeg(uint64_t *Addr) {
using Type = uint64_t;
std::atomic<Type> *MemData = reinterpret_cast<std::atomic<Type>*>(Addr);
Type Expected = MemData->load();
Type Desired = -Expected;
do {
Desired = -Expected;
} while (!MemData->compare_exchange_strong(Expected, Desired, std::memory_order_seq_cst));
return Expected;
}
template<typename T>
T AtomicCompareAndSwap(T expected, T desired, T *addr)
{
std::atomic<T> *MemData = reinterpret_cast<std::atomic<T>*>(addr);
T Src1 = expected;
T Src2 = desired;
T Expected = Src1;
bool Result = MemData->compare_exchange_strong(Expected, Src2);
return Result ? Src1 : Expected;
}
template uint8_t AtomicCompareAndSwap<uint8_t>(uint8_t expected, uint8_t desired, uint8_t *addr);
template uint16_t AtomicCompareAndSwap<uint16_t>(uint16_t expected, uint16_t desired, uint16_t *addr);
template uint32_t AtomicCompareAndSwap<uint32_t>(uint32_t expected, uint32_t desired, uint32_t *addr);
template uint64_t AtomicCompareAndSwap<uint64_t>(uint64_t expected, uint64_t desired, uint64_t *addr);
#else
// Needs to match what the AArch64 JIT and unaligned signal handler expects
uint8_t AtomicFetchNeg(uint8_t *Addr) {
using Type = uint8_t;
Type Result{};
Type Tmp{};
Type TmpStatus{};
__asm__ volatile(
R"(
1:
ldaxrb %w[Result], [%[Memory]];
neg %w[Tmp], %w[Result];
stlxrb %w[TmpStatus], %w[Tmp], [%[Memory]];
cbnz %w[TmpStatus], 1b;
)"
: [Result] "=r" (Result)
, [Tmp] "=r" (Tmp)
, [TmpStatus] "=r" (TmpStatus)
, [Memory] "+r" (Addr)
:: "memory"
);
return Result;
}
uint16_t AtomicFetchNeg(uint16_t *Addr) {
using Type = uint16_t;
Type Result{};
Type Tmp{};
Type TmpStatus{};
__asm__ volatile(
R"(
1:
ldaxrh %w[Result], [%[Memory]];
neg %w[Tmp], %w[Result];
stlxrh %w[TmpStatus], %w[Tmp], [%[Memory]];
cbnz %w[TmpStatus], 1b;
)"
: [Result] "=r" (Result)
, [Tmp] "=r" (Tmp)
, [TmpStatus] "=r" (TmpStatus)
, [Memory] "+r" (Addr)
:: "memory"
);
return Result;
}
uint32_t AtomicFetchNeg(uint32_t *Addr) {
using Type = uint32_t;
Type Result{};
Type Tmp{};
Type TmpStatus{};
__asm__ volatile(
R"(
1:
ldaxr %w[Result], [%[Memory]];
neg %w[Tmp], %w[Result];
stlxr %w[TmpStatus], %w[Tmp], [%[Memory]];
cbnz %w[TmpStatus], 1b;
)"
: [Result] "=r" (Result)
, [Tmp] "=r" (Tmp)
, [TmpStatus] "=r" (TmpStatus)
, [Memory] "+r" (Addr)
:: "memory"
);
return Result;
}
uint64_t AtomicFetchNeg(uint64_t *Addr) {
using Type = uint64_t;
Type Result{};
Type Tmp{};
Type TmpStatus{};
__asm__ volatile(
R"(
1:
ldaxr %[Result], [%[Memory]];
neg %[Tmp], %[Result];
stlxr %w[TmpStatus], %[Tmp], [%[Memory]];
cbnz %w[TmpStatus], 1b;
)"
: [Result] "=r" (Result)
, [Tmp] "=r" (Tmp)
, [TmpStatus] "=r" (TmpStatus)
, [Memory] "+r" (Addr)
:: "memory"
);
return Result;
}
template<>
uint8_t AtomicCompareAndSwap(uint8_t expected, uint8_t desired, uint8_t *addr) {
using Type = uint8_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxrb %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected], uxtb;
b.ne 2f;
stlxrb %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint16_t AtomicCompareAndSwap(uint16_t expected, uint16_t desired, uint16_t *addr) {
using Type = uint16_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxrh %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected], uxth;
b.ne 2f;
stlxrh %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint32_t AtomicCompareAndSwap(uint32_t expected, uint32_t desired, uint32_t *addr) {
using Type = uint32_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxr %w[Tmp], [%[Memory]];
cmp %w[Tmp], %w[Expected];
b.ne 2f;
stlxr %w[Tmp2], %w[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %w[Result], %w[Expected];
b 3f;
2:
mov %w[Result], %w[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
template<>
uint64_t AtomicCompareAndSwap(uint64_t expected, uint64_t desired, uint64_t *addr) {
using Type = uint64_t;
//force Result to r9 (scratch register) or clang spills to stack
register Type Result asm("r9"){};
Type Tmp{};
Type Tmp2{};
__asm__ volatile(
R"(
1:
ldaxr %[Tmp], [%[Memory]];
cmp %[Tmp], %[Expected];
b.ne 2f;
stlxr %w[Tmp2], %[Desired], [%[Memory]];
cbnz %w[Tmp2], 1b;
mov %[Result], %[Expected];
b 3f;
2:
mov %[Result], %[Tmp];
clrex;
3:
)"
: [Tmp] "=r" (Tmp)
, [Tmp2] "=r" (Tmp2)
, [Desired] "+r" (desired)
, [Expected] "+r" (expected)
, [Result] "=r" (Result)
, [Memory] "+r" (addr)
:: "memory"
);
return Result;
}
#endif
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
// Size is the size of each pair element
switch (OpSize) {
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
);
break;
}
case 8: {
std::atomic<__uint128_t> *MemData = *GetSrc<std::atomic<__uint128_t> **>(Data->SSAData, Op->Header.Args[2]);
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
__uint128_t Expected = Src1;
bool Result = MemData->compare_exchange_strong(Expected, Src2);
memcpy(GDP, Result ? &Src1 : &Expected, 16);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", OpSize); break;
}
}
DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1: {
GD = AtomicCompareAndSwap(
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint8_t**>(Data->SSAData, Op->Header.Args[2])
);
break;
}
case 2: {
GD = AtomicCompareAndSwap(
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint16_t**>(Data->SSAData, Op->Header.Args[2])
);
break;
}
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint32_t**>(Data->SSAData, Op->Header.Args[2])
);
break;
}
case 8: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", OpSize); break;
}
}
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData += Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData += Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData += Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData += Src;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData -= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData -= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData -= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData -= Src;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData &= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData &= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData &= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData &= Src;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData |= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData |= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData |= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData |= Src;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData ^= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData ^= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData ^= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
*MemData ^= Src;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
switch (IROp->Size) {
case 1: {
using Type = uint8_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
break;
}
case 2: {
using Type = uint16_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
break;
}
case 4: {
using Type = uint32_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
break;
}
case 8: {
using Type = uint64_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
#undef DEF_OP
void InterpreterOps::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
REGISTER_OP(ATOMICSUB, AtomicSub);
REGISTER_OP(ATOMICAND, AtomicAnd);
REGISTER_OP(ATOMICOR, AtomicOr);
REGISTER_OP(ATOMICXOR, AtomicXor);
REGISTER_OP(ATOMICSWAP, AtomicSwap);
REGISTER_OP(ATOMICFETCHADD, AtomicFetchAdd);
REGISTER_OP(ATOMICFETCHSUB, AtomicFetchSub);
REGISTER_OP(ATOMICFETCHAND, AtomicFetchAnd);
REGISTER_OP(ATOMICFETCHOR, AtomicFetchOr);
REGISTER_OP(ATOMICFETCHXOR, AtomicFetchXor);
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
#undef REGISTER_OP
}
}
@@ -0,0 +1,157 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <cstdint>
namespace FEXCore::CPU {
[[noreturn]]
static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->SignalThread(Thread, FEXCore::Core::SignalEvent::Return);
LOGMAN_MSG_A_FMT("unreachable");
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
DEF_OP(CallbackReturn) {
Data->State->CTX->InterpreterCallbackReturn(Data->State, Data->StackEntry);
}
DEF_OP(ExitFunction) {
auto Op = IROp->C<IR::IROp_ExitFunction>();
uint8_t OpSize = IROp->Size;
uintptr_t* ContextPtr = reinterpret_cast<uintptr_t*>(Data->State->CurrentFrame);
void *ContextData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
memcpy(ContextData, Src, OpSize);
Data->BlockResults.Quit = true;
}
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
uintptr_t ListBegin = Data->CurrentIR->GetListData();
uintptr_t DataBegin = Data->CurrentIR->GetData();
Data->BlockIterator = IR::NodeIterator(ListBegin, DataBegin, Op->Header.Args[0]);
Data->BlockResults.Redo = true;
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
uintptr_t ListBegin = Data->CurrentIR->GetListData();
uintptr_t DataBegin = Data->CurrentIR->GetData();
bool CompResult;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp1);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp2);
if (Op->CompareSize == 4)
CompResult = IsConditionTrue<uint32_t, int32_t, float>(Op->Cond.Val, Src1, Src2);
else
CompResult = IsConditionTrue<uint64_t, int64_t, double>(Op->Cond.Val, Src1, Src2);
if (CompResult) {
Data->BlockIterator = IR::NodeIterator(ListBegin, DataBegin, Op->TrueBlock);
}
else {
Data->BlockIterator = IR::NodeIterator(ListBegin, DataBegin, Op->FalseBlock);
}
Data->BlockResults.Redo = true;
}
DEF_OP(Syscall) {
auto Op = IROp->C<IR::IROp_Syscall>();
FEXCore::HLE::SyscallArguments Args;
for (size_t j = 0; j < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++j) {
if (Op->Header.Args[j].IsInvalid()) break;
Args.Argument[j] = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[j]);
}
uint64_t Res = FEXCore::Context::HandleSyscall(Data->State->CTX->SyscallHandler, Data->State->CurrentFrame, &Args);
GD = Res;
}
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
auto thunkFn = Data->State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
thunkFn(*GetSrc<void**>(Data->SSAData, Op->Header.Args[0]));
}
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
auto CodePtr = Data->CurrentEntry + Op->Offset;
if (memcmp((void*)CodePtr, &Op->CodeOriginalLow, Op->CodeLength) != 0) {
GD = 1;
} else {
GD = 0;
}
}
DEF_OP(RemoveCodeEntry) {
Data->State->CTX->RemoveCodeEntry(Data->State, Data->CurrentEntry);
}
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
uint64_t *DstPtr = GetDest<uint64_t*>(Data->SSAData, Node);
uint64_t Arg = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Leaf = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
auto Results = Data->State->CTX->CPUID.RunFunction(Arg, Leaf);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
}
#undef DEF_OP
void InterpreterOps::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
}
@@ -0,0 +1,237 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
uint8_t OpSize = IROp->Size;
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Offset = Op->Index * Op->Header.ElementSize * 8;
__uint128_t Mask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
if (Op->Header.ElementSize == 8) {
Mask = ~0ULL;
}
Src2 = Src2 & Mask;
Mask <<= Offset;
Mask = ~Mask;
__uint128_t Dst = Src1 & Mask;
Dst |= Src2 << Offset;
memcpy(GDP, &Dst, OpSize);
}
DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[0]), Op->Header.ElementSize);
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
float Dst = (float)*GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0408: { // Float <- int64_t
float Dst = (float)*GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0804: { // Double <- int32_t
double Dst = (double)*GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0808: { // Double <- int64_t
double Dst = (double)*GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
}
}
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // Double <- Float
double Dst = (double)*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, 8);
break;
}
case 0x0408: { // Float <- Double
float Dst = (float)*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, &Dst, 4);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
}
}
DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
auto Func = [](auto a, auto min, auto max) { return a; };
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, float, int32_t, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, double, int64_t, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
auto Func = [](auto a, auto min, auto max) { return std::trunc(a); };
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, int32_t, float, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, int64_t, double, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
auto Func = [](auto a, auto min, auto max) { return std::nearbyint(a); };
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, int32_t, float, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, int64_t, double, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToF) {
auto Op = IROp->C<IR::IROp_Vector_FToF>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
auto Func = [](auto a, auto min, auto max) { return a; };
switch (Conv) {
case 0x0804: { // Double <- float
// Only the lower elements from the source
// This uses half the source elements
uint8_t Elements = OpSize / 8;
DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(double, float, Func, 0, 0)
break;
}
case 0x0408: { // Float <- Double
// Little bit tricky here
// Sometimes is used to convert from a 128bit vector register
// in to a 64bit vector register with different sized elements
// eg: %ssa5 i32v2 = Vector_FToF %ssa4 i128, #0x8
uint8_t Elements = (OpSize << 1) / Op->SrcElementSize;
DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(float, double, Func, 0, 0)
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Conversion Type : 0x{:04x}", Conv); break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
auto Func_Nearest = [](auto a) { return std::rint(a); };
auto Func_Neg = [](auto a) { return std::floor(a); };
auto Func_Pos = [](auto a) { return std::ceil(a); };
auto Func_Trunc = [](auto a) { return std::trunc(a); };
auto Func_Host = [](auto a) { return std::rint(a); };
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Nearest)
DO_VECTOR_1SRC_OP(8, double, Func_Nearest)
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Neg)
DO_VECTOR_1SRC_OP(8, double, Func_Neg)
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Pos)
DO_VECTOR_1SRC_OP(8, double, Func_Pos)
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Trunc)
DO_VECTOR_1SRC_OP(8, double, Func_Trunc)
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Host)
DO_VECTOR_1SRC_OP(8, double, Func_Host)
}
break;
}
memcpy(GDP, Tmp, OpSize);
}
#undef DEF_OP
void InterpreterOps::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
@@ -0,0 +1,443 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
namespace AES {
static __uint128_t InvShiftRows(uint8_t *State) {
uint8_t Shifted[16] = {
State[0], State[13], State[10], State[7],
State[4], State[1], State[14], State[11],
State[8], State[5], State[2], State[15],
State[12], State[9], State[6], State[3],
};
__uint128_t Res{};
memcpy(&Res, Shifted, 16);
return Res;
}
static __uint128_t InvSubBytes(uint8_t *State) {
// 16x16 matrix table
static const uint8_t InvSubstitutionTable[256] = {
0x52, 0x09, 0x6a, 0xd5, 0x30, 0x36, 0xa5, 0x38, 0xbf, 0x40, 0xa3, 0x9e, 0x81, 0xf3, 0xd7, 0xfb,
0x7c, 0xe3, 0x39, 0x82, 0x9b, 0x2f, 0xff, 0x87, 0x34, 0x8e, 0x43, 0x44, 0xc4, 0xde, 0xe9, 0xcb,
0x54, 0x7b, 0x94, 0x32, 0xa6, 0xc2, 0x23, 0x3d, 0xee, 0x4c, 0x95, 0x0b, 0x42, 0xfa, 0xc3, 0x4e,
0x08, 0x2e, 0xa1, 0x66, 0x28, 0xd9, 0x24, 0xb2, 0x76, 0x5b, 0xa2, 0x49, 0x6d, 0x8b, 0xd1, 0x25,
0x72, 0xf8, 0xf6, 0x64, 0x86, 0x68, 0x98, 0x16, 0xd4, 0xa4, 0x5c, 0xcc, 0x5d, 0x65, 0xb6, 0x92,
0x6c, 0x70, 0x48, 0x50, 0xfd, 0xed, 0xb9, 0xda, 0x5e, 0x15, 0x46, 0x57, 0xa7, 0x8d, 0x9d, 0x84,
0x90, 0xd8, 0xab, 0x00, 0x8c, 0xbc, 0xd3, 0x0a, 0xf7, 0xe4, 0x58, 0x05, 0xb8, 0xb3, 0x45, 0x06,
0xd0, 0x2c, 0x1e, 0x8f, 0xca, 0x3f, 0x0f, 0x02, 0xc1, 0xaf, 0xbd, 0x03, 0x01, 0x13, 0x8a, 0x6b,
0x3a, 0x91, 0x11, 0x41, 0x4f, 0x67, 0xdc, 0xea, 0x97, 0xf2, 0xcf, 0xce, 0xf0, 0xb4, 0xe6, 0x73,
0x96, 0xac, 0x74, 0x22, 0xe7, 0xad, 0x35, 0x85, 0xe2, 0xf9, 0x37, 0xe8, 0x1c, 0x75, 0xdf, 0x6e,
0x47, 0xf1, 0x1a, 0x71, 0x1d, 0x29, 0xc5, 0x89, 0x6f, 0xb7, 0x62, 0x0e, 0xaa, 0x18, 0xbe, 0x1b,
0xfc, 0x56, 0x3e, 0x4b, 0xc6, 0xd2, 0x79, 0x20, 0x9a, 0xdb, 0xc0, 0xfe, 0x78, 0xcd, 0x5a, 0xf4,
0x1f, 0xdd, 0xa8, 0x33, 0x88, 0x07, 0xc7, 0x31, 0xb1, 0x12, 0x10, 0x59, 0x27, 0x80, 0xec, 0x5f,
0x60, 0x51, 0x7f, 0xa9, 0x19, 0xb5, 0x4a, 0x0d, 0x2d, 0xe5, 0x7a, 0x9f, 0x93, 0xc9, 0x9c, 0xef,
0xa0, 0xe0, 0x3b, 0x4d, 0xae, 0x2a, 0xf5, 0xb0, 0xc8, 0xeb, 0xbb, 0x3c, 0x83, 0x53, 0x99, 0x61,
0x17, 0x2b, 0x04, 0x7e, 0xba, 0x77, 0xd6, 0x26, 0xe1, 0x69, 0x14, 0x63, 0x55, 0x21, 0x0c, 0x7d,
};
// Uses a byte substitution table with a constant set of values
// Needs to do a table look up
uint8_t Substituted[16];
for (size_t i = 0; i < 16; ++i) {
Substituted[i] = InvSubstitutionTable[State[i]];
}
__uint128_t Res{};
memcpy(&Res, Substituted, 16);
return Res;
}
static __uint128_t ShiftRows(uint8_t *State) {
uint8_t Shifted[16] = {
State[0], State[5], State[10], State[15],
State[4], State[9], State[14], State[3],
State[8], State[13], State[2], State[7],
State[12], State[1], State[6], State[11],
};
__uint128_t Res{};
memcpy(&Res, Shifted, 16);
return Res;
}
static __uint128_t SubBytes(uint8_t *State, size_t Bytes) {
// 16x16 matrix table
static const uint8_t SubstitutionTable[256] = {
0x63, 0x7c, 0x77, 0x7b, 0xf2, 0x6b, 0x6f, 0xc5, 0x30, 0x01, 0x67, 0x2b, 0xfe, 0xd7, 0xab, 0x76,
0xca, 0x82, 0xc9, 0x7d, 0xfa, 0x59, 0x47, 0xf0, 0xad, 0xd4, 0xa2, 0xaf, 0x9c, 0xa4, 0x72, 0xc0,
0xb7, 0xfd, 0x93, 0x26, 0x36, 0x3f, 0xf7, 0xcc, 0x34, 0xa5, 0xe5, 0xf1, 0x71, 0xd8, 0x31, 0x15,
0x04, 0xc7, 0x23, 0xc3, 0x18, 0x96, 0x05, 0x9a, 0x07, 0x12, 0x80, 0xe2, 0xeb, 0x27, 0xb2, 0x75,
0x09, 0x83, 0x2c, 0x1a, 0x1b, 0x6e, 0x5a, 0xa0, 0x52, 0x3b, 0xd6, 0xb3, 0x29, 0xe3, 0x2f, 0x84,
0x53, 0xd1, 0x00, 0xed, 0x20, 0xfc, 0xb1, 0x5b, 0x6a, 0xcb, 0xbe, 0x39, 0x4a, 0x4c, 0x58, 0xcf,
0xd0, 0xef, 0xaa, 0xfb, 0x43, 0x4d, 0x33, 0x85, 0x45, 0xf9, 0x02, 0x7f, 0x50, 0x3c, 0x9f, 0xa8,
0x51, 0xa3, 0x40, 0x8f, 0x92, 0x9d, 0x38, 0xf5, 0xbc, 0xb6, 0xda, 0x21, 0x10, 0xff, 0xf3, 0xd2,
0xcd, 0x0c, 0x13, 0xec, 0x5f, 0x97, 0x44, 0x17, 0xc4, 0xa7, 0x7e, 0x3d, 0x64, 0x5d, 0x19, 0x73,
0x60, 0x81, 0x4f, 0xdc, 0x22, 0x2a, 0x90, 0x88, 0x46, 0xee, 0xb8, 0x14, 0xde, 0x5e, 0x0b, 0xdb,
0xe0, 0x32, 0x3a, 0x0a, 0x49, 0x06, 0x24, 0x5c, 0xc2, 0xd3, 0xac, 0x62, 0x91, 0x95, 0xe4, 0x79,
0xe7, 0xc8, 0x37, 0x6d, 0x8d, 0xd5, 0x4e, 0xa9, 0x6c, 0x56, 0xf4, 0xea, 0x65, 0x7a, 0xae, 0x08,
0xba, 0x78, 0x25, 0x2e, 0x1c, 0xa6, 0xb4, 0xc6, 0xe8, 0xdd, 0x74, 0x1f, 0x4b, 0xbd, 0x8b, 0x8a,
0x70, 0x3e, 0xb5, 0x66, 0x48, 0x03, 0xf6, 0x0e, 0x61, 0x35, 0x57, 0xb9, 0x86, 0xc1, 0x1d, 0x9e,
0xe1, 0xf8, 0x98, 0x11, 0x69, 0xd9, 0x8e, 0x94, 0x9b, 0x1e, 0x87, 0xe9, 0xce, 0x55, 0x28, 0xdf,
0x8c, 0xa1, 0x89, 0x0d, 0xbf, 0xe6, 0x42, 0x68, 0x41, 0x99, 0x2d, 0x0f, 0xb0, 0x54, 0xbb, 0x16,
};
// Uses a byte substitution table with a constant set of values
// Needs to do a table look up
uint8_t Substituted[16];
Bytes = std::min(Bytes, (size_t)16);
for (size_t i = 0; i < Bytes; ++i) {
Substituted[i] = SubstitutionTable[State[i]];
}
__uint128_t Res{};
memcpy(&Res, Substituted, Bytes);
return Res;
}
static uint8_t FFMul02(uint8_t in) {
static const uint8_t FFMul02[256] = {
0x00, 0x02, 0x04, 0x06, 0x08, 0x0a, 0x0c, 0x0e, 0x10, 0x12, 0x14, 0x16, 0x18, 0x1a, 0x1c, 0x1e,
0x20, 0x22, 0x24, 0x26, 0x28, 0x2a, 0x2c, 0x2e, 0x30, 0x32, 0x34, 0x36, 0x38, 0x3a, 0x3c, 0x3e,
0x40, 0x42, 0x44, 0x46, 0x48, 0x4a, 0x4c, 0x4e, 0x50, 0x52, 0x54, 0x56, 0x58, 0x5a, 0x5c, 0x5e,
0x60, 0x62, 0x64, 0x66, 0x68, 0x6a, 0x6c, 0x6e, 0x70, 0x72, 0x74, 0x76, 0x78, 0x7a, 0x7c, 0x7e,
0x80, 0x82, 0x84, 0x86, 0x88, 0x8a, 0x8c, 0x8e, 0x90, 0x92, 0x94, 0x96, 0x98, 0x9a, 0x9c, 0x9e,
0xa0, 0xa2, 0xa4, 0xa6, 0xa8, 0xaa, 0xac, 0xae, 0xb0, 0xb2, 0xb4, 0xb6, 0xb8, 0xba, 0xbc, 0xbe,
0xc0, 0xc2, 0xc4, 0xc6, 0xc8, 0xca, 0xcc, 0xce, 0xd0, 0xd2, 0xd4, 0xd6, 0xd8, 0xda, 0xdc, 0xde,
0xe0, 0xe2, 0xe4, 0xe6, 0xe8, 0xea, 0xec, 0xee, 0xf0, 0xf2, 0xf4, 0xf6, 0xf8, 0xfa, 0xfc, 0xfe,
0x1b, 0x19, 0x1f, 0x1d, 0x13, 0x11, 0x17, 0x15, 0x0b, 0x09, 0x0f, 0x0d, 0x03, 0x01, 0x07, 0x05,
0x3b, 0x39, 0x3f, 0x3d, 0x33, 0x31, 0x37, 0x35, 0x2b, 0x29, 0x2f, 0x2d, 0x23, 0x21, 0x27, 0x25,
0x5b, 0x59, 0x5f, 0x5d, 0x53, 0x51, 0x57, 0x55, 0x4b, 0x49, 0x4f, 0x4d, 0x43, 0x41, 0x47, 0x45,
0x7b, 0x79, 0x7f, 0x7d, 0x73, 0x71, 0x77, 0x75, 0x6b, 0x69, 0x6f, 0x6d, 0x63, 0x61, 0x67, 0x65,
0x9b, 0x99, 0x9f, 0x9d, 0x93, 0x91, 0x97, 0x95, 0x8b, 0x89, 0x8f, 0x8d, 0x83, 0x81, 0x87, 0x85,
0xbb, 0xb9, 0xbf, 0xbd, 0xb3, 0xb1, 0xb7, 0xb5, 0xab, 0xa9, 0xaf, 0xad, 0xa3, 0xa1, 0xa7, 0xa5,
0xdb, 0xd9, 0xdf, 0xdd, 0xd3, 0xd1, 0xd7, 0xd5, 0xcb, 0xc9, 0xcf, 0xcd, 0xc3, 0xc1, 0xc7, 0xc5,
0xfb, 0xf9, 0xff, 0xfd, 0xf3, 0xf1, 0xf7, 0xf5, 0xeb, 0xe9, 0xef, 0xed, 0xe3, 0xe1, 0xe7, 0xe5,
};
return FFMul02[in];
}
static uint8_t FFMul03(uint8_t in) {
static const uint8_t FFMul03[256] = {
0x00, 0x03, 0x06, 0x05, 0x0c, 0x0f, 0x0a, 0x09, 0x18, 0x1b, 0x1e, 0x1d, 0x14, 0x17, 0x12, 0x11,
0x30, 0x33, 0x36, 0x35, 0x3c, 0x3f, 0x3a, 0x39, 0x28, 0x2b, 0x2e, 0x2d, 0x24, 0x27, 0x22, 0x21,
0x60, 0x63, 0x66, 0x65, 0x6c, 0x6f, 0x6a, 0x69, 0x78, 0x7b, 0x7e, 0x7d, 0x74, 0x77, 0x72, 0x71,
0x50, 0x53, 0x56, 0x55, 0x5c, 0x5f, 0x5a, 0x59, 0x48, 0x4b, 0x4e, 0x4d, 0x44, 0x47, 0x42, 0x41,
0xc0, 0xc3, 0xc6, 0xc5, 0xcc, 0xcf, 0xca, 0xc9, 0xd8, 0xdb, 0xde, 0xdd, 0xd4, 0xd7, 0xd2, 0xd1,
0xf0, 0xf3, 0xf6, 0xf5, 0xfc, 0xff, 0xfa, 0xf9, 0xe8, 0xeb, 0xee, 0xed, 0xe4, 0xe7, 0xe2, 0xe1,
0xa0, 0xa3, 0xa6, 0xa5, 0xac, 0xaf, 0xaa, 0xa9, 0xb8, 0xbb, 0xbe, 0xbd, 0xb4, 0xb7, 0xb2, 0xb1,
0x90, 0x93, 0x96, 0x95, 0x9c, 0x9f, 0x9a, 0x99, 0x88, 0x8b, 0x8e, 0x8d, 0x84, 0x87, 0x82, 0x81,
0x9b, 0x98, 0x9d, 0x9e, 0x97, 0x94, 0x91, 0x92, 0x83, 0x80, 0x85, 0x86, 0x8f, 0x8c, 0x89, 0x8a,
0xab, 0xa8, 0xad, 0xae, 0xa7, 0xa4, 0xa1, 0xa2, 0xb3, 0xb0, 0xb5, 0xb6, 0xbf, 0xbc, 0xb9, 0xba,
0xfb, 0xf8, 0xfd, 0xfe, 0xf7, 0xf4, 0xf1, 0xf2, 0xe3, 0xe0, 0xe5, 0xe6, 0xef, 0xec, 0xe9, 0xea,
0xcb, 0xc8, 0xcd, 0xce, 0xc7, 0xc4, 0xc1, 0xc2, 0xd3, 0xd0, 0xd5, 0xd6, 0xdf, 0xdc, 0xd9, 0xda,
0x5b, 0x58, 0x5d, 0x5e, 0x57, 0x54, 0x51, 0x52, 0x43, 0x40, 0x45, 0x46, 0x4f, 0x4c, 0x49, 0x4a,
0x6b, 0x68, 0x6d, 0x6e, 0x67, 0x64, 0x61, 0x62, 0x73, 0x70, 0x75, 0x76, 0x7f, 0x7c, 0x79, 0x7a,
0x3b, 0x38, 0x3d, 0x3e, 0x37, 0x34, 0x31, 0x32, 0x23, 0x20, 0x25, 0x26, 0x2f, 0x2c, 0x29, 0x2a,
0x0b, 0x08, 0x0d, 0x0e, 0x07, 0x04, 0x01, 0x02, 0x13, 0x10, 0x15, 0x16, 0x1f, 0x1c, 0x19, 0x1a,
};
return FFMul03[in];
}
static __uint128_t MixColumns(uint8_t *State) {
uint8_t In0[16] = {
State[0], State[4], State[8], State[12],
State[1], State[5], State[9], State[13],
State[2], State[6], State[10], State[14],
State[3], State[7], State[11], State[15],
};
uint8_t Out0[4]{};
uint8_t Out1[4]{};
uint8_t Out2[4]{};
uint8_t Out3[4]{};
for (size_t i = 0; i < 4; ++i) {
Out0[i] = FFMul02(In0[0 + i]) ^ FFMul03(In0[4 + i]) ^ In0[8 + i] ^ In0[12 + i];
Out1[i] = In0[0 + i] ^ FFMul02(In0[4 + i]) ^ FFMul03(In0[8 + i]) ^ In0[12 + i];
Out2[i] = In0[0 + i] ^ In0[4 + i] ^ FFMul02(In0[8 + i]) ^ FFMul03(In0[12 + i]);
Out3[i] = FFMul03(In0[0 + i]) ^ In0[4 + i] ^ In0[8 + i] ^ FFMul02(In0[12 + i]);
}
uint8_t OutArray[16] = {
Out0[0], Out1[0], Out2[0], Out3[0],
Out0[1], Out1[1], Out2[1], Out3[1],
Out0[2], Out1[2], Out2[2], Out3[2],
Out0[3], Out1[3], Out2[3], Out3[3],
};
__uint128_t Res{};
memcpy(&Res, OutArray, 16);
return Res;
}
static uint8_t FFMul09(uint8_t in) {
static const uint8_t FFMul09[256] = {
0x00, 0x09, 0x12, 0x1b, 0x24, 0x2d, 0x36, 0x3f, 0x48, 0x41, 0x5a, 0x53, 0x6c, 0x65, 0x7e, 0x77,
0x90, 0x99, 0x82, 0x8b, 0xb4, 0xbd, 0xa6, 0xaf, 0xd8, 0xd1, 0xca, 0xc3, 0xfc, 0xf5, 0xee, 0xe7,
0x3b, 0x32, 0x29, 0x20, 0x1f, 0x16, 0x0d, 0x04, 0x73, 0x7a, 0x61, 0x68, 0x57, 0x5e, 0x45, 0x4c,
0xab, 0xa2, 0xb9, 0xb0, 0x8f, 0x86, 0x9d, 0x94, 0xe3, 0xea, 0xf1, 0xf8, 0xc7, 0xce, 0xd5, 0xdc,
0x76, 0x7f, 0x64, 0x6d, 0x52, 0x5b, 0x40, 0x49, 0x3e, 0x37, 0x2c, 0x25, 0x1a, 0x13, 0x08, 0x01,
0xe6, 0xef, 0xf4, 0xfd, 0xc2, 0xcb, 0xd0, 0xd9, 0xae, 0xa7, 0xbc, 0xb5, 0x8a, 0x83, 0x98, 0x91,
0x4d, 0x44, 0x5f, 0x56, 0x69, 0x60, 0x7b, 0x72, 0x05, 0x0c, 0x17, 0x1e, 0x21, 0x28, 0x33, 0x3a,
0xdd, 0xd4, 0xcf, 0xc6, 0xf9, 0xf0, 0xeb, 0xe2, 0x95, 0x9c, 0x87, 0x8e, 0xb1, 0xb8, 0xa3, 0xaa,
0xec, 0xe5, 0xfe, 0xf7, 0xc8, 0xc1, 0xda, 0xd3, 0xa4, 0xad, 0xb6, 0xbf, 0x80, 0x89, 0x92, 0x9b,
0x7c, 0x75, 0x6e, 0x67, 0x58, 0x51, 0x4a, 0x43, 0x34, 0x3d, 0x26, 0x2f, 0x10, 0x19, 0x02, 0x0b,
0xd7, 0xde, 0xc5, 0xcc, 0xf3, 0xfa, 0xe1, 0xe8, 0x9f, 0x96, 0x8d, 0x84, 0xbb, 0xb2, 0xa9, 0xa0,
0x47, 0x4e, 0x55, 0x5c, 0x63, 0x6a, 0x71, 0x78, 0x0f, 0x06, 0x1d, 0x14, 0x2b, 0x22, 0x39, 0x30,
0x9a, 0x93, 0x88, 0x81, 0xbe, 0xb7, 0xac, 0xa5, 0xd2, 0xdb, 0xc0, 0xc9, 0xf6, 0xff, 0xe4, 0xed,
0x0a, 0x03, 0x18, 0x11, 0x2e, 0x27, 0x3c, 0x35, 0x42, 0x4b, 0x50, 0x59, 0x66, 0x6f, 0x74, 0x7d,
0xa1, 0xa8, 0xb3, 0xba, 0x85, 0x8c, 0x97, 0x9e, 0xe9, 0xe0, 0xfb, 0xf2, 0xcd, 0xc4, 0xdf, 0xd6,
0x31, 0x38, 0x23, 0x2a, 0x15, 0x1c, 0x07, 0x0e, 0x79, 0x70, 0x6b, 0x62, 0x5d, 0x54, 0x4f, 0x46,
};
return FFMul09[in];
}
static uint8_t FFMul0B(uint8_t in) {
static const uint8_t FFMul0B[256] = {
0x00, 0x0b, 0x16, 0x1d, 0x2c, 0x27, 0x3a, 0x31, 0x58, 0x53, 0x4e, 0x45, 0x74, 0x7f, 0x62, 0x69,
0xb0, 0xbb, 0xa6, 0xad, 0x9c, 0x97, 0x8a, 0x81, 0xe8, 0xe3, 0xfe, 0xf5, 0xc4, 0xcf, 0xd2, 0xd9,
0x7b, 0x70, 0x6d, 0x66, 0x57, 0x5c, 0x41, 0x4a, 0x23, 0x28, 0x35, 0x3e, 0x0f, 0x04, 0x19, 0x12,
0xcb, 0xc0, 0xdd, 0xd6, 0xe7, 0xec, 0xf1, 0xfa, 0x93, 0x98, 0x85, 0x8e, 0xbf, 0xb4, 0xa9, 0xa2,
0xf6, 0xfd, 0xe0, 0xeb, 0xda, 0xd1, 0xcc, 0xc7, 0xae, 0xa5, 0xb8, 0xb3, 0x82, 0x89, 0x94, 0x9f,
0x46, 0x4d, 0x50, 0x5b, 0x6a, 0x61, 0x7c, 0x77, 0x1e, 0x15, 0x08, 0x03, 0x32, 0x39, 0x24, 0x2f,
0x8d, 0x86, 0x9b, 0x90, 0xa1, 0xaa, 0xb7, 0xbc, 0xd5, 0xde, 0xc3, 0xc8, 0xf9, 0xf2, 0xef, 0xe4,
0x3d, 0x36, 0x2b, 0x20, 0x11, 0x1a, 0x07, 0x0c, 0x65, 0x6e, 0x73, 0x78, 0x49, 0x42, 0x5f, 0x54,
0xf7, 0xfc, 0xe1, 0xea, 0xdb, 0xd0, 0xcd, 0xc6, 0xaf, 0xa4, 0xb9, 0xb2, 0x83, 0x88, 0x95, 0x9e,
0x47, 0x4c, 0x51, 0x5a, 0x6b, 0x60, 0x7d, 0x76, 0x1f, 0x14, 0x09, 0x02, 0x33, 0x38, 0x25, 0x2e,
0x8c, 0x87, 0x9a, 0x91, 0xa0, 0xab, 0xb6, 0xbd, 0xd4, 0xdf, 0xc2, 0xc9, 0xf8, 0xf3, 0xee, 0xe5,
0x3c, 0x37, 0x2a, 0x21, 0x10, 0x1b, 0x06, 0x0d, 0x64, 0x6f, 0x72, 0x79, 0x48, 0x43, 0x5e, 0x55,
0x01, 0x0a, 0x17, 0x1c, 0x2d, 0x26, 0x3b, 0x30, 0x59, 0x52, 0x4f, 0x44, 0x75, 0x7e, 0x63, 0x68,
0xb1, 0xba, 0xa7, 0xac, 0x9d, 0x96, 0x8b, 0x80, 0xe9, 0xe2, 0xff, 0xf4, 0xc5, 0xce, 0xd3, 0xd8,
0x7a, 0x71, 0x6c, 0x67, 0x56, 0x5d, 0x40, 0x4b, 0x22, 0x29, 0x34, 0x3f, 0x0e, 0x05, 0x18, 0x13,
0xca, 0xc1, 0xdc, 0xd7, 0xe6, 0xed, 0xf0, 0xfb, 0x92, 0x99, 0x84, 0x8f, 0xbe, 0xb5, 0xa8, 0xa3,
};
return FFMul0B[in];
}
static uint8_t FFMul0D(uint8_t in) {
static const uint8_t FFMul0D[256] = {
0x00, 0x0d, 0x1a, 0x17, 0x34, 0x39, 0x2e, 0x23, 0x68, 0x65, 0x72, 0x7f, 0x5c, 0x51, 0x46, 0x4b,
0xd0, 0xdd, 0xca, 0xc7, 0xe4, 0xe9, 0xfe, 0xf3, 0xb8, 0xb5, 0xa2, 0xaf, 0x8c, 0x81, 0x96, 0x9b,
0xbb, 0xb6, 0xa1, 0xac, 0x8f, 0x82, 0x95, 0x98, 0xd3, 0xde, 0xc9, 0xc4, 0xe7, 0xea, 0xfd, 0xf0,
0x6b, 0x66, 0x71, 0x7c, 0x5f, 0x52, 0x45, 0x48, 0x03, 0x0e, 0x19, 0x14, 0x37, 0x3a, 0x2d, 0x20,
0x6d, 0x60, 0x77, 0x7a, 0x59, 0x54, 0x43, 0x4e, 0x05, 0x08, 0x1f, 0x12, 0x31, 0x3c, 0x2b, 0x26,
0xbd, 0xb0, 0xa7, 0xaa, 0x89, 0x84, 0x93, 0x9e, 0xd5, 0xd8, 0xcf, 0xc2, 0xe1, 0xec, 0xfb, 0xf6,
0xd6, 0xdb, 0xcc, 0xc1, 0xe2, 0xef, 0xf8, 0xf5, 0xbe, 0xb3, 0xa4, 0xa9, 0x8a, 0x87, 0x90, 0x9d,
0x06, 0x0b, 0x1c, 0x11, 0x32, 0x3f, 0x28, 0x25, 0x6e, 0x63, 0x74, 0x79, 0x5a, 0x57, 0x40, 0x4d,
0xda, 0xd7, 0xc0, 0xcd, 0xee, 0xe3, 0xf4, 0xf9, 0xb2, 0xbf, 0xa8, 0xa5, 0x86, 0x8b, 0x9c, 0x91,
0x0a, 0x07, 0x10, 0x1d, 0x3e, 0x33, 0x24, 0x29, 0x62, 0x6f, 0x78, 0x75, 0x56, 0x5b, 0x4c, 0x41,
0x61, 0x6c, 0x7b, 0x76, 0x55, 0x58, 0x4f, 0x42, 0x09, 0x04, 0x13, 0x1e, 0x3d, 0x30, 0x27, 0x2a,
0xb1, 0xbc, 0xab, 0xa6, 0x85, 0x88, 0x9f, 0x92, 0xd9, 0xd4, 0xc3, 0xce, 0xed, 0xe0, 0xf7, 0xfa,
0xb7, 0xba, 0xad, 0xa0, 0x83, 0x8e, 0x99, 0x94, 0xdf, 0xd2, 0xc5, 0xc8, 0xeb, 0xe6, 0xf1, 0xfc,
0x67, 0x6a, 0x7d, 0x70, 0x53, 0x5e, 0x49, 0x44, 0x0f, 0x02, 0x15, 0x18, 0x3b, 0x36, 0x21, 0x2c,
0x0c, 0x01, 0x16, 0x1b, 0x38, 0x35, 0x22, 0x2f, 0x64, 0x69, 0x7e, 0x73, 0x50, 0x5d, 0x4a, 0x47,
0xdc, 0xd1, 0xc6, 0xcb, 0xe8, 0xe5, 0xf2, 0xff, 0xb4, 0xb9, 0xae, 0xa3, 0x80, 0x8d, 0x9a, 0x97,
};
return FFMul0D[in];
}
static uint8_t FFMul0E(uint8_t in) {
static const uint8_t FFMul0E[256] = {
0x00, 0x0e, 0x1c, 0x12, 0x38, 0x36, 0x24, 0x2a, 0x70, 0x7e, 0x6c, 0x62, 0x48, 0x46, 0x54, 0x5a,
0xe0, 0xee, 0xfc, 0xf2, 0xd8, 0xd6, 0xc4, 0xca, 0x90, 0x9e, 0x8c, 0x82, 0xa8, 0xa6, 0xb4, 0xba,
0xdb, 0xd5, 0xc7, 0xc9, 0xe3, 0xed, 0xff, 0xf1, 0xab, 0xa5, 0xb7, 0xb9, 0x93, 0x9d, 0x8f, 0x81,
0x3b, 0x35, 0x27, 0x29, 0x03, 0x0d, 0x1f, 0x11, 0x4b, 0x45, 0x57, 0x59, 0x73, 0x7d, 0x6f, 0x61,
0xad, 0xa3, 0xb1, 0xbf, 0x95, 0x9b, 0x89, 0x87, 0xdd, 0xd3, 0xc1, 0xcf, 0xe5, 0xeb, 0xf9, 0xf7,
0x4d, 0x43, 0x51, 0x5f, 0x75, 0x7b, 0x69, 0x67, 0x3d, 0x33, 0x21, 0x2f, 0x05, 0x0b, 0x19, 0x17,
0x76, 0x78, 0x6a, 0x64, 0x4e, 0x40, 0x52, 0x5c, 0x06, 0x08, 0x1a, 0x14, 0x3e, 0x30, 0x22, 0x2c,
0x96, 0x98, 0x8a, 0x84, 0xae, 0xa0, 0xb2, 0xbc, 0xe6, 0xe8, 0xfa, 0xf4, 0xde, 0xd0, 0xc2, 0xcc,
0x41, 0x4f, 0x5d, 0x53, 0x79, 0x77, 0x65, 0x6b, 0x31, 0x3f, 0x2d, 0x23, 0x09, 0x07, 0x15, 0x1b,
0xa1, 0xaf, 0xbd, 0xb3, 0x99, 0x97, 0x85, 0x8b, 0xd1, 0xdf, 0xcd, 0xc3, 0xe9, 0xe7, 0xf5, 0xfb,
0x9a, 0x94, 0x86, 0x88, 0xa2, 0xac, 0xbe, 0xb0, 0xea, 0xe4, 0xf6, 0xf8, 0xd2, 0xdc, 0xce, 0xc0,
0x7a, 0x74, 0x66, 0x68, 0x42, 0x4c, 0x5e, 0x50, 0x0a, 0x04, 0x16, 0x18, 0x32, 0x3c, 0x2e, 0x20,
0xec, 0xe2, 0xf0, 0xfe, 0xd4, 0xda, 0xc8, 0xc6, 0x9c, 0x92, 0x80, 0x8e, 0xa4, 0xaa, 0xb8, 0xb6,
0x0c, 0x02, 0x10, 0x1e, 0x34, 0x3a, 0x28, 0x26, 0x7c, 0x72, 0x60, 0x6e, 0x44, 0x4a, 0x58, 0x56,
0x37, 0x39, 0x2b, 0x25, 0x0f, 0x01, 0x13, 0x1d, 0x47, 0x49, 0x5b, 0x55, 0x7f, 0x71, 0x63, 0x6d,
0xd7, 0xd9, 0xcb, 0xc5, 0xef, 0xe1, 0xf3, 0xfd, 0xa7, 0xa9, 0xbb, 0xb5, 0x9f, 0x91, 0x83, 0x8d,
};
return FFMul0E[in];
}
static __uint128_t InvMixColumns(uint8_t *State) {
uint8_t In0[16] = {
State[0], State[4], State[8], State[12],
State[1], State[5], State[9], State[13],
State[2], State[6], State[10], State[14],
State[3], State[7], State[11], State[15],
};
uint8_t Out0[4]{};
uint8_t Out1[4]{};
uint8_t Out2[4]{};
uint8_t Out3[4]{};
for (size_t i = 0; i < 4; ++i) {
Out0[i] = FFMul0E(In0[0 + i]) ^ FFMul0B(In0[4 + i]) ^ FFMul0D(In0[8 + i]) ^ FFMul09(In0[12 + i]);
Out1[i] = FFMul09(In0[0 + i]) ^ FFMul0E(In0[4 + i]) ^ FFMul0B(In0[8 + i]) ^ FFMul0D(In0[12 + i]);
Out2[i] = FFMul0D(In0[0 + i]) ^ FFMul09(In0[4 + i]) ^ FFMul0E(In0[8 + i]) ^ FFMul0B(In0[12 + i]);
Out3[i] = FFMul0B(In0[0 + i]) ^ FFMul0D(In0[4 + i]) ^ FFMul09(In0[8 + i]) ^ FFMul0E(In0[12 + i]);
}
uint8_t OutArray[16] = {
Out0[0], Out1[0], Out2[0], Out3[0],
Out0[1], Out1[1], Out2[1], Out3[1],
Out0[2], Out1[2], Out2[2], Out3[2],
Out0[3], Out1[3], Out2[3], Out3[3],
};
__uint128_t Res{};
memcpy(&Res, OutArray, 16);
return Res;
}
}
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
// Pseudo-code
// Dst = InvMixColumns(STATE)
__uint128_t Tmp{};
Tmp = AES::InvMixColumns(reinterpret_cast<uint8_t*>(&Src1));
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
// Pseudo-code
// STATE = Src1
// RoundKey = Src2
// STATE = ShiftRows(STATE)
// STATE = SubBytes(STATE)
// STATE = MixColumns(STATE)
// Dst = STATE XOR RoundKey
__uint128_t Tmp{};
Tmp = AES::ShiftRows(reinterpret_cast<uint8_t*>(&Src1));
Tmp = AES::SubBytes(reinterpret_cast<uint8_t*>(&Tmp), 16);
Tmp = AES::MixColumns(reinterpret_cast<uint8_t*>(&Tmp));
Tmp = Tmp ^ Src2;
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
// Pseudo-code
// STATE = Src1
// RoundKey = Src2
// STATE = ShiftRows(STATE)
// STATE = SubBytes(STATE)
// Dst = STATE XOR RoundKey
__uint128_t Tmp{};
Tmp = AES::ShiftRows(reinterpret_cast<uint8_t*>(&Src1));
Tmp = AES::SubBytes(reinterpret_cast<uint8_t*>(&Tmp), 16);
Tmp = Tmp ^ Src2;
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
// Pseudo-code
// STATE = Src1
// RoundKey = Src2
// STATE = InvShiftRows(STATE)
// STATE = InvSubBytes(STATE)
// STATE = InvMixColumns(STATE)
// Dst = STATE XOR RoundKey
__uint128_t Tmp{};
Tmp = AES::InvShiftRows(reinterpret_cast<uint8_t*>(&Src1));
Tmp = AES::InvSubBytes(reinterpret_cast<uint8_t*>(&Tmp));
Tmp = AES::InvMixColumns(reinterpret_cast<uint8_t*>(&Tmp));
Tmp = Tmp ^ Src2;
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
// Pseudo-code
// STATE = Src1
// RoundKey = Src2
// STATE = InvShiftRows(STATE)
// STATE = InvSubBytes(STATE)
// Dst = STATE XOR RoundKey
__uint128_t Tmp{};
Tmp = AES::InvShiftRows(reinterpret_cast<uint8_t*>(&Src1));
Tmp = AES::InvSubBytes(reinterpret_cast<uint8_t*>(&Tmp));
Tmp = Tmp ^ Src2;
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
// Pseudo-code
// X3 = Src1[127:96]
// X2 = Src1[95:64]
// X1 = Src1[63:32]
// X0 = Src1[31:30]
// RCON = (Zext)rcon
// Dest[31:0] = SubWord(X1)
// Dest[63:32] = RotWord(SubWord(X1)) XOR RCON
// Dest[95:64] = SubWord(X3)
// Dest[127:96] = RotWord(SubWord(X3)) XOR RCON
__uint128_t Tmp{};
uint32_t X1{};
uint32_t X3{};
memcpy(&X1, &Src1[4], 4);
memcpy(&X3, &Src1[12], 4);
uint32_t SubWord_X1 = AES::SubBytes(reinterpret_cast<uint8_t*>(&X1), 4);
uint32_t SubWord_X3 = AES::SubBytes(reinterpret_cast<uint8_t*>(&X3), 4);
auto Ror = [] (auto In, auto R) {
auto RotateMask = sizeof(In) * 8 - 1;
R &= RotateMask;
return (In >> R) | (In << (sizeof(In) * 8 - R));
};
uint32_t Rot_X1 = Ror(SubWord_X1, 8);
uint32_t Rot_X3 = Ror(SubWord_X3, 8);
Tmp = Rot_X3 ^ Op->RCON;
Tmp <<= 32;
Tmp |= SubWord_X3;
Tmp <<= 32;
Tmp |= Rot_X1 ^ Op->RCON;
Tmp <<= 32;
Tmp |= SubWord_X1;
memcpy(GDP, &Tmp, sizeof(Tmp));
}
#undef DEF_OP
void InterpreterOps::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
#undef REGISTER_OP
}
}
@@ -0,0 +1,389 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include "F80Ops.h"
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(F80LOADFCW) {
FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle(*GetSrc<uint16_t*>(Data->SSAData, IROp->Args[0]));
}
DEF_OP(F80ADD) {
auto Op = IROp->C<IR::IROp_F80Add>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FADD(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SUB) {
auto Op = IROp->C<IR::IROp_F80Sub>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSUB(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80MUL) {
auto Op = IROp->C<IR::IROp_F80Mul>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FMUL(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80DIV) {
auto Op = IROp->C<IR::IROp_F80Div>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FDIV(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FYL2X) {
auto Op = IROp->C<IR::IROp_F80FYL2X>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FYL2X(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80ATAN) {
auto Op = IROp->C<IR::IROp_F80ATAN>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FATAN(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FPREM1) {
auto Op = IROp->C<IR::IROp_F80FPREM1>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FREM1(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FPREM) {
auto Op = IROp->C<IR::IROp_F80FPREM>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FREM(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SCALE) {
auto Op = IROp->C<IR::IROp_F80SCALE>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSCALE(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80CVT) {
auto Op = IROp->C<IR::IROp_F80CVT>();
uint8_t OpSize = IROp->Size;
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
switch (OpSize) {
case 4: {
float Tmp = Src;
memcpy(GDP, &Tmp, OpSize);
break;
}
case 8: {
double Tmp = Src;
memcpy(GDP, &Tmp, OpSize);
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
}
DEF_OP(F80CVTINT) {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
uint8_t OpSize = IROp->Size;
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
switch (OpSize) {
case 2: {
int16_t Tmp = (Op->Truncate? FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t : FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2)(Src);
memcpy(GDP, &Tmp, sizeof(Tmp));
break;
}
case 4: {
int32_t Tmp = (Op->Truncate? FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t : FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4)(Src);
memcpy(GDP, &Tmp, sizeof(Tmp));
break;
}
case 8: {
int64_t Tmp = (Op->Truncate? FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t : FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8)(Src);
memcpy(GDP, &Tmp, sizeof(Tmp));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
}
DEF_OP(F80CVTTO) {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
case 4: {
float Src = *GetSrc<float *>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
case 8: {
double Src = *GetSrc<double *>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
}
}
DEF_OP(F80CVTTOINT) {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
case 2: {
int16_t Src = *GetSrc<int16_t*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
case 4: {
int32_t Src = *GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
}
}
DEF_OP(F80ROUND) {
auto Op = IROp->C<IR::IROp_F80Round>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FRNDINT(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80F2XM1) {
auto Op = IROp->C<IR::IROp_F80F2XM1>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::F2XM1(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80TAN) {
auto Op = IROp->C<IR::IROp_F80TAN>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FTAN(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SQRT) {
auto Op = IROp->C<IR::IROp_F80SQRT>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSQRT(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SIN) {
auto Op = IROp->C<IR::IROp_F80SIN>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSIN(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80COS) {
auto Op = IROp->C<IR::IROp_F80COS>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FCOS(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80XTRACT_EXP) {
auto Op = IROp->C<IR::IROp_F80XTRACT_EXP>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FXTRACT_EXP(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80XTRACT_SIG) {
auto Op = IROp->C<IR::IROp_F80XTRACT_SIG>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FXTRACT_SIG(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80CMP) {
auto Op = IROp->C<IR::IROp_F80Cmp>();
uint32_t ResultFlags{};
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
bool eq, lt, nan;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Op->Flags & (1 << IR::FCMP_FLAG_LT) &&
lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_UNORDERED) &&
nan) {
ResultFlags |= (1 << IR::FCMP_FLAG_UNORDERED);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ) &&
eq) {
ResultFlags |= (1 << IR::FCMP_FLAG_EQ);
}
GD = ResultFlags;
}
DEF_OP(F80BCDLOAD) {
auto Op = IROp->C<IR::IROp_F80BCDLoad>();
uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t BCD{};
// We walk through each uint8_t and pull out the BCD encoding
// Each 4bit split is a digit
// Only 0-9 is supported, A-F results in undefined data
// | 4 bit | 4 bit |
// | 10s place | 1s place |
// EG 0x48 = 48
// EG 0x4847 = 4847
// This gives us an 18digit value encoded in BCD
// The last byte lets us know if it negative or not
for (size_t i = 0; i < 9; ++i) {
uint8_t Digit = Src1[8 - i];
// First shift our last value over
BCD *= 100;
// Add the tens place digit
BCD += (Digit >> 4) * 10;
// Add the ones place digit
BCD += Digit & 0xF;
}
// Set negative flag once converted to x87
bool Negative = Src1[9] & 0x80;
X80SoftFloat Tmp;
Tmp = BCD;
Tmp.Sign = Negative;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80BCDSTORE) {
auto Op = IROp->C<IR::IROp_F80BCDStore>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
bool Negative = Src1.Sign;
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1;
uint8_t BCD[10]{};
for (size_t i = 0; i < 9; ++i) {
if (Tmp == 0) {
// Nothing left? Just leave
break;
}
// Extract the lower 100 values
uint8_t Digit = Tmp % 100;
// Now divide it for the next iteration
Tmp /= 100;
uint8_t UpperNibble = Digit / 10;
uint8_t LowerNibble = Digit % 10;
// Now store the BCD
BCD[i] = (UpperNibble << 4) | LowerNibble;
}
// Set negative flag once converted to x87
BCD[9] = Negative ? 0x80 : 0;
memcpy(GDP, BCD, 10);
}
#undef DEF_OP
void InterpreterOps::RegisterF80Handlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(F80LOADFCW, F80LOADFCW);
REGISTER_OP(F80ADD, F80ADD);
REGISTER_OP(F80SUB, F80SUB);
REGISTER_OP(F80MUL, F80MUL);
REGISTER_OP(F80DIV, F80DIV);
REGISTER_OP(F80FYL2X, F80FYL2X);
REGISTER_OP(F80ATAN, F80ATAN);
REGISTER_OP(F80FPREM1, F80FPREM1);
REGISTER_OP(F80FPREM, F80FPREM);
REGISTER_OP(F80SCALE, F80SCALE);
REGISTER_OP(F80CVT, F80CVT);
REGISTER_OP(F80CVTINT, F80CVTINT);
REGISTER_OP(F80CVTTO, F80CVTTO);
REGISTER_OP(F80CVTTOINT, F80CVTTOINT);
REGISTER_OP(F80ROUND, F80ROUND);
REGISTER_OP(F80F2XM1, F80F2XM1);
REGISTER_OP(F80TAN, F80TAN);
REGISTER_OP(F80SQRT, F80SQRT);
REGISTER_OP(F80SIN, F80SIN);
REGISTER_OP(F80COS, F80COS);
REGISTER_OP(F80XTRACT_EXP, F80XTRACT_EXP);
REGISTER_OP(F80XTRACT_SIG, F80XTRACT_SIG);
REGISTER_OP(F80CMP, F80CMP);
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
#undef REGISTER_OP
}
}
@@ -0,0 +1,330 @@
#pragma once
#include "Common/SoftFloat.h"
#include "Common/SoftFloat-3e/softfloat.h"
#include <FEXCore/IR/IR.h>
namespace FEXCore::CPU {
template<IR::IROps Op>
struct OpHandlers {
};
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
static X80SoftFloat handle4(float src) {
return src;
}
static X80SoftFloat handle8(double src) {
return src;
}
};
template<>
struct OpHandlers<IR::OP_F80CMP> {
template<uint32_t Flags>
static uint64_t handle(X80SoftFloat Src1, X80SoftFloat Src2) {
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Flags & (1 << IR::FCMP_FLAG_LT) &&
lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
if (Flags & (1 << IR::FCMP_FLAG_UNORDERED) &&
nan) {
ResultFlags |= (1 << IR::FCMP_FLAG_UNORDERED);
}
if (Flags & (1 << IR::FCMP_FLAG_EQ) &&
eq) {
ResultFlags |= (1 << IR::FCMP_FLAG_EQ);
}
return ResultFlags;
}
};
template<>
struct OpHandlers<IR::OP_F80CVT> {
static float handle4(X80SoftFloat src) {
return src;
}
static double handle8(X80SoftFloat src) {
return src;
}
};
template<>
struct OpHandlers<IR::OP_F80CVTINT> {
static int16_t handle2(X80SoftFloat src) {
return src;
}
static int32_t handle4(X80SoftFloat src) {
return src;
}
static int64_t handle8(X80SoftFloat src) {
return src;
}
static int16_t handle2t(X80SoftFloat src) {
auto rv = extF80_to_i32(src, softfloat_round_minMag, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
return INT16_MIN;
} else {
return rv;
}
}
static int32_t handle4t(X80SoftFloat src) {
return extF80_to_i32(src, softfloat_round_minMag, false);
}
static int64_t handle8t(X80SoftFloat src) {
return extF80_to_i64(src, softfloat_round_minMag, false);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTTOINT> {
static X80SoftFloat handle2(int16_t src) {
return src;
}
static X80SoftFloat handle4(int32_t src) {
return src;
}
};
template<>
struct OpHandlers<IR::OP_F80ROUND> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FRNDINT(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80F2XM1> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::F2XM1(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80TAN> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FTAN(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SQRT> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FSQRT(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SIN> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FSIN(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80COS> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FCOS(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_EXP> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_EXP(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_SIG> {
static X80SoftFloat handle(X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_SIG(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80ADD> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FADD(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SUB> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FSUB(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80MUL> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FMUL(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80DIV> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FDIV(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2X> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FYL2X(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FATAN(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM1> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FREM1(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FREM(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SCALE> {
static X80SoftFloat handle(X80SoftFloat Src1, X80SoftFloat Src2) {
return X80SoftFloat::FSCALE(Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
static X80SoftFloat handle(X80SoftFloat Src1) {
bool Negative = Src1.Sign;
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1;
X80SoftFloat Rv;
uint8_t *BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
for (size_t i = 0; i < 9; ++i) {
if (Tmp == 0) {
// Nothing left? Just leave
break;
}
// Extract the lower 100 values
uint8_t Digit = Tmp % 100;
// Now divide it for the next iteration
Tmp /= 100;
uint8_t UpperNibble = Digit / 10;
uint8_t LowerNibble = Digit % 10;
// Now store the BCD
BCD[i] = (UpperNibble << 4) | LowerNibble;
}
// Set negative flag once converted to x87
BCD[9] = Negative ? 0x80 : 0;
return Rv;
}
};
template<>
struct OpHandlers<IR::OP_F80BCDLOAD> {
static X80SoftFloat handle(X80SoftFloat Src) {
uint8_t *Src1 = reinterpret_cast<uint8_t *>(&Src);
uint64_t BCD{};
// We walk through each uint8_t and pull out the BCD encoding
// Each 4bit split is a digit
// Only 0-9 is supported, A-F results in undefined data
// | 4 bit | 4 bit |
// | 10s place | 1s place |
// EG 0x48 = 48
// EG 0x4847 = 4847
// This gives us an 18digit value encoded in BCD
// The last byte lets us know if it negative or not
for (size_t i = 0; i < 9; ++i) {
uint8_t Digit = Src1[8 - i];
// First shift our last value over
BCD *= 100;
// Add the tens place digit
BCD += (Digit >> 4) * 10;
// Add the ones place digit
BCD += Digit & 0xF;
}
// Set negative flag once converted to x87
bool Negative = Src1[9] & 0x80;
X80SoftFloat Tmp;
Tmp = BCD;
Tmp.Sign = Negative;
return Tmp;
}
};
template<>
struct OpHandlers<IR::OP_F80LOADFCW> {
static void handle(uint16_t NewFCW) {
auto PC = (NewFCW >> 8) & 3;
switch(PC) {
case 0: extF80_roundingPrecision = 32; break;
case 2: extF80_roundingPrecision = 64; break;
case 3: extF80_roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
}
auto RC = (NewFCW >> 10) & 3;
switch(RC) {
case 0:
softfloat_roundingMode = softfloat_round_near_even;
break;
case 1:
softfloat_roundingMode = softfloat_round_min;
break;
case 2:
softfloat_roundingMode = softfloat_round_max;
break;
case 3:
softfloat_roundingMode = softfloat_round_minMag;
break;
}
}
};
}
@@ -0,0 +1,27 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
GD = (*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]) >> Op->Flag) & 1;
}
#undef DEF_OP
void InterpreterOps::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
}
@@ -20,31 +20,38 @@ using DestMapType = std::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
std::string GetName() override { return "Interpreter"; }
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
explicit InterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
bool NeedsOpDispatch() override { return true; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
static void InitializeInterpreterOpHandlers();
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
uint32_t AllocateTmpSpace(size_t Size);
template<typename Res>
Res GetDest(void* SSAData, IR::OrderedNodeWrapper Op);
template<typename Res>
Res GetSrc(void* SSAData, IR::OrderedNodeWrapper Src);
std::unique_ptr<Dispatcher> Dispatcher{};
};
}
template<typename T>
T AtomicCompareAndSwap(T expected, T desired, T *addr);
uint8_t AtomicFetchNeg(uint8_t *Addr);
uint16_t AtomicFetchNeg(uint16_t *Addr);
uint32_t AtomicFetchNeg(uint32_t *Addr);
uint64_t AtomicFetchNeg(uint64_t *Addr);
} // namespace FEXCore::CPU
@@ -35,70 +35,27 @@ static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
bool InterpreterCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
#ifdef _M_ARM_64
constexpr bool is_arm64 = true;
#else
constexpr bool is_arm64 = false;
#endif
if constexpr (is_arm64) {
uint32_t *PC = reinterpret_cast<uint32_t*>(ArchHelpers::Context::GetPc(ucontext));
uint32_t Instr = PC[0];
if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x: PC: %p Instruction: 0x%08x\n", Op, PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXR_MASK) == FEXCore::ArchHelpers::Arm64::LDAXR_INST) { // LDAXR*
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleAtomicLoadstoreExclusive(ucontext, info);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS LDAXR: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
void InitializeInterpreterOpHandlers() {
for (uint32_t i = 0; i <= FEXCore::IR::IROps::OP_LAST; ++i) {
InterpreterOps::OpHandlers[i] = &InterpreterOps::Op_Unhandled;
}
return false;
InterpreterOps::RegisterALUHandlers();
InterpreterOps::RegisterAtomicHandlers();
InterpreterOps::RegisterBranchHandlers();
InterpreterOps::RegisterConversionHandlers();
InterpreterOps::RegisterFlagHandlers();
InterpreterOps::RegisterMemoryHandlers();
InterpreterOps::RegisterMiscHandlers();
InterpreterOps::RegisterMoveHandlers();
InterpreterOps::RegisterVectorHandlers();
InterpreterOps::RegisterEncryptionHandlers();
InterpreterOps::RegisterF80Handlers();
}
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: CTX {ctx}
, State {Thread} {
// Grab our space for temporary data
if (!CompileThread &&
CTX->Config.Core == FEXCore::Config::CONFIG_INTERPRETER) {
@@ -108,10 +65,12 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->HandleSIGBUS(Signal, info, ucontext);
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
@@ -13,6 +13,10 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
void InitializeInterpreterOpHandlers();
}
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
} // namespace FEXCore::CPU
@@ -0,0 +1,179 @@
#pragma once
#include <FEXCore/IR/IR.h>
#define GD *GetDest<uint64_t*>(Data->SSAData, Node)
#define GDP GetDest<void*>(Data->SSAData, Node)
#define DO_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(GDP); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
*Dst_d = func(*Src1_d, *Src2_d); \
break; \
}
#define DO_SCALAR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
Dst_d[0] = func(Src1_d[0], Src2_d[0]); \
break; \
}
#define DO_VECTOR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_PAIR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i*2], Src1_d[i*2 + 1]); \
Dst_d[i+Elements] = func(Src2_d[i*2], Src2_d[i*2 + 1]); \
} \
break; \
}
#define DO_VECTOR_SCALAR_OP(size, type, func)\
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], *Src2_d); \
} \
break; \
}
#define DO_VECTOR_0SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(); \
} \
break; \
}
#define DO_VECTOR_1SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src_d[i]); \
} \
break; \
}
#define DO_VECTOR_REDUCE_1SRC_OP(size, type, func, start_val) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
type begin = start_val; \
for (uint8_t i = 0; i < Elements; ++i) { \
begin = func(begin, Src_d[i]); \
} \
Dst_d[0] = begin; \
break; \
}
#define DO_VECTOR_SAT_OP(size, type, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(type, type2, func, min, max) \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src2); \
memcpy(Dst_d, Src1, Elements * sizeof(type2));\
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i+Elements] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP_SRC(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i+Elements], min, max); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i], (type)Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP_TOP_SRC(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i+Elements], (type)Src2_d[i+Elements]); \
} \
break; \
}
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID()];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, uint32_t Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID()];
return reinterpret_cast<Res>(DstPtr);
}
File diff suppressed because it is too large. Load diff
@@ -1,6 +1,9 @@
#pragma once
#include <stdint.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::Core {
struct InternalThreadState;
}
@@ -42,5 +45,366 @@ namespace FEXCore::CPU {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
static void RegisterALUHandlers();
static void RegisterAtomicHandlers();
static void RegisterBranchHandlers();
static void RegisterConversionHandlers();
static void RegisterFlagHandlers();
static void RegisterMemoryHandlers();
static void RegisterMiscHandlers();
static void RegisterMoveHandlers();
static void RegisterVectorHandlers();
static void RegisterEncryptionHandlers();
static void RegisterF80Handlers();
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
uint64_t CurrentEntry{};
FEXCore::IR::IRListView *CurrentIR{};
volatile void *StackEntry{};
void *SSAData{};
struct {
bool Quit;
bool Redo;
} BlockResults{};
IR::NodeIterator BlockIterator{0, 0};
};
using OpHandler = std::function<void(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)>;
static std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers;
#define DEF_OP(x) static void Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
///< Unhandled handler
DEF_OP(Unhandled);
///< No-op Handler
DEF_OP(NoOp);
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
DEF_OP(CycleCounter);
DEF_OP(Add);
DEF_OP(Sub);
DEF_OP(Neg);
DEF_OP(Mul);
DEF_OP(UMul);
DEF_OP(Div);
DEF_OP(UDiv);
DEF_OP(Rem);
DEF_OP(URem);
DEF_OP(MulH);
DEF_OP(UMulH);
DEF_OP(Or);
DEF_OP(And);
DEF_OP(Andn);
DEF_OP(Xor);
DEF_OP(Lshl);
DEF_OP(Lshr);
DEF_OP(Ashr);
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
DEF_OP(LURem);
DEF_OP(Zext);
DEF_OP(Not);
DEF_OP(Popcount);
DEF_OP(FindLSB);
DEF_OP(FindMSB);
DEF_OP(FindTrailingZeros);
DEF_OP(CountLeadingZeroes);
DEF_OP(Rev);
DEF_OP(Bfi);
DEF_OP(Bfe);
DEF_OP(Sbfe);
DEF_OP(Select);
DEF_OP(VExtractToGPR);
DEF_OP(Float_ToGPR_ZU);
DEF_OP(Float_ToGPR_ZS);
DEF_OP(Float_ToGPR_S);
DEF_OP(FCmp);
///< Atomic ops
DEF_OP(CASPair);
DEF_OP(CAS);
DEF_OP(AtomicAdd);
DEF_OP(AtomicSub);
DEF_OP(AtomicAnd);
DEF_OP(AtomicOr);
DEF_OP(AtomicXor);
DEF_OP(AtomicSwap);
DEF_OP(AtomicFetchAdd);
DEF_OP(AtomicFetchSub);
DEF_OP(AtomicFetchAnd);
DEF_OP(AtomicFetchOr);
DEF_OP(AtomicFetchXor);
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
DEF_OP(CondJump);
DEF_OP(Syscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_SToF);
DEF_OP(Vector_FToZS);
DEF_OP(Vector_FToS);
DEF_OP(Vector_FToF);
DEF_OP(Vector_FToI);
///< Flag ops
DEF_OP(GetHostFlag);
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
DEF_OP(LoadContextIndexed);
DEF_OP(StoreContextIndexed);
DEF_OP(SpillRegister);
DEF_OP(FillRegister);
DEF_OP(LoadFlag);
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
///< Misc ops
DEF_OP(EndBlock);
DEF_OP(Fence);
DEF_OP(Break);
DEF_OP(Phi);
DEF_OP(PhiValue);
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
///< Move ops
DEF_OP(ExtractElementPair);
DEF_OP(CreateElementPair);
DEF_OP(Mov);
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
DEF_OP(VOr);
DEF_OP(VXor);
DEF_OP(VAdd);
DEF_OP(VSub);
DEF_OP(VUQAdd);
DEF_OP(VUQSub);
DEF_OP(VSQAdd);
DEF_OP(VSQSub);
DEF_OP(VAddP);
DEF_OP(VAddV);
DEF_OP(VUMinV);
DEF_OP(VURAvg);
DEF_OP(VAbs);
DEF_OP(VPopcount);
DEF_OP(VFAdd);
DEF_OP(VFAddP);
DEF_OP(VFSub);
DEF_OP(VFMul);
DEF_OP(VFDiv);
DEF_OP(VFMin);
DEF_OP(VFMax);
DEF_OP(VFRecp);
DEF_OP(VFSqrt);
DEF_OP(VFRSqrt);
DEF_OP(VNeg);
DEF_OP(VFNeg);
DEF_OP(VNot);
DEF_OP(VUMin);
DEF_OP(VSMin);
DEF_OP(VUMax);
DEF_OP(VSMax);
DEF_OP(VZip);
DEF_OP(VUnZip);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
DEF_OP(VCMPGT);
DEF_OP(VCMPGTZ);
DEF_OP(VCMPLTZ);
DEF_OP(VFCMPEQ);
DEF_OP(VFCMPNEQ);
DEF_OP(VFCMPLT);
DEF_OP(VFCMPGT);
DEF_OP(VFCMPLE);
DEF_OP(VFCMPORD);
DEF_OP(VFCMPUNO);
DEF_OP(VUShl);
DEF_OP(VUShr);
DEF_OP(VSShr);
DEF_OP(VUShlS);
DEF_OP(VUShrS);
DEF_OP(VSShrS);
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
DEF_OP(VUShrNI);
DEF_OP(VUShrNI2);
DEF_OP(VBitcast);
DEF_OP(VSXTL);
DEF_OP(VSXTL2);
DEF_OP(VUXTL);
DEF_OP(VUXTL2);
DEF_OP(VSQXTN);
DEF_OP(VSQXTN2);
DEF_OP(VSQXTUN);
DEF_OP(VSQXTUN2);
DEF_OP(VUMul);
DEF_OP(VUMull);
DEF_OP(VSMul);
DEF_OP(VSMull);
DEF_OP(VUMull2);
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
///< Encryption ops
DEF_OP(AESImc);
DEF_OP(AESEnc);
DEF_OP(AESEncLast);
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
///< F80 ops
DEF_OP(F80LOADFCW);
DEF_OP(F80ADD);
DEF_OP(F80SUB);
DEF_OP(F80MUL);
DEF_OP(F80DIV);
DEF_OP(F80FYL2X);
DEF_OP(F80ATAN);
DEF_OP(F80FPREM1);
DEF_OP(F80FPREM);
DEF_OP(F80SCALE);
DEF_OP(F80CVT);
DEF_OP(F80CVTINT);
DEF_OP(F80CVTTO);
DEF_OP(F80CVTTOINT);
DEF_OP(F80ROUND);
DEF_OP(F80F2XM1);
DEF_OP(F80TAN);
DEF_OP(F80SQRT);
DEF_OP(F80SIN);
DEF_OP(F80COS);
DEF_OP(F80XTRACT_EXP);
DEF_OP(F80XTRACT_SIG);
DEF_OP(F80CMP);
DEF_OP(F80BCDLOAD);
DEF_OP(F80BCDSTORE);
#undef DEF_OP
template<typename unsigned_type, typename signed_type, typename float_type>
[[nodiscard]] static bool IsConditionTrue(uint8_t Cond, uint64_t Src1, uint64_t Src2) {
bool CompResult = false;
switch (Cond) {
case FEXCore::IR::COND_EQ:
CompResult = static_cast<unsigned_type>(Src1) == static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_NEQ:
CompResult = static_cast<unsigned_type>(Src1) != static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_SGE:
CompResult = static_cast<signed_type>(Src1) >= static_cast<signed_type>(Src2);
break;
case FEXCore::IR::COND_SLT:
CompResult = static_cast<signed_type>(Src1) < static_cast<signed_type>(Src2);
break;
case FEXCore::IR::COND_SGT:
CompResult = static_cast<signed_type>(Src1) > static_cast<signed_type>(Src2);
break;
case FEXCore::IR::COND_SLE:
CompResult = static_cast<signed_type>(Src1) <= static_cast<signed_type>(Src2);
break;
case FEXCore::IR::COND_UGE:
CompResult = static_cast<unsigned_type>(Src1) >= static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_ULT:
CompResult = static_cast<unsigned_type>(Src1) < static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_UGT:
CompResult = static_cast<unsigned_type>(Src1) > static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_ULE:
CompResult = static_cast<unsigned_type>(Src1) <= static_cast<unsigned_type>(Src2);
break;
case FEXCore::IR::COND_FLU:
CompResult = reinterpret_cast<float_type&>(Src1) < reinterpret_cast<float_type&>(Src2) || (std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_FGE:
CompResult = reinterpret_cast<float_type&>(Src1) >= reinterpret_cast<float_type&>(Src2) && !(std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_FLEU:
CompResult = reinterpret_cast<float_type&>(Src1) <= reinterpret_cast<float_type&>(Src2) || (std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_FGT:
CompResult = reinterpret_cast<float_type&>(Src1) > reinterpret_cast<float_type&>(Src2) && !(std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_FU:
CompResult = (std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_FNU:
CompResult = !(std::isnan(reinterpret_cast<float_type&>(Src1)) || std::isnan(reinterpret_cast<float_type&>(Src2)));
break;
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
default:
LOGMAN_MSG_A_FMT("Unsupported compare type");
break;
}
return CompResult;
}
static uint8_t GetOpSize(FEXCore::IR::IRListView *CurrentIR, IR::OrderedNodeWrapper Node) {
auto IROp = CurrentIR->GetOp<FEXCore::IR::IROp_Header>(Node);
return IROp->Size;
}
};
};
@@ -0,0 +1,289 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
namespace FEXCore::CPU {
static inline void CacheLineFlush(char *Addr) {
#ifdef _M_X86_64
__asm volatile (
"clflush (%[Addr]);"
:: [Addr] "r" (Addr)
: "memory");
#else
__builtin___clear_cache(Addr, Addr+64);
#endif
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->Offset;
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(ContextPtr); \
GD = *MemData; \
break; \
}
switch (OpSize) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16: {
void const *MemData = reinterpret_cast<void const*>(ContextPtr);
memcpy(GDP, MemData, OpSize);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
}
#undef LOAD_CTX
}
DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->Offset;
void *MemData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
memcpy(MemData, Src, OpSize);
}
DEF_OP(LoadRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(StoreRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(ContextPtr); \
GD = *MemData; \
break; \
}
switch (IROp->Size) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16: {
void const *MemData = reinterpret_cast<void const*>(ContextPtr);
memcpy(GDP, MemData, IROp->Size);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
}
#undef LOAD_CTX
}
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
void *MemData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
memcpy(MemData, Src, IROp->Size);
}
DEF_OP(SpillRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(FillRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(LoadFlag) {
auto Op = IROp->C<IR::IROp_LoadFlag>();
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += offsetof(FEXCore::Core::CPUState, flags[0]);
ContextPtr += Op->Flag;
uint8_t const *MemData = reinterpret_cast<uint8_t const*>(ContextPtr);
GD = *MemData;
}
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
uint8_t Arg = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += offsetof(FEXCore::Core::CPUState, flags[0]);
ContextPtr += Op->Flag;
uint8_t *MemData = reinterpret_cast<uint8_t*>(ContextPtr);
*MemData = Arg;
}
DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
uint8_t OpSize = IROp->Size;
uint8_t const *MemData = *GetSrc<uint8_t const**>(Data->SSAData, Op->Addr);
if (!Op->Offset.IsInvalid()) {
auto Offset = *GetSrc<uintptr_t const*>(Data->SSAData, Op->Offset) * Op->OffsetScale;
switch(Op->OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val: MemData += Offset; break;
case IR::MEM_OFFSET_UXTW.Val: MemData += (uint32_t)Offset; break;
case IR::MEM_OFFSET_SXTW.Val: MemData += (int32_t)Offset; break;
}
}
memset(GDP, 0, 16);
switch (OpSize) {
case 1: {
auto D = reinterpret_cast<const std::atomic<uint8_t>*>(MemData);
GD = D->load();
break;
}
case 2: {
auto D = reinterpret_cast<const std::atomic<uint16_t>*>(MemData);
GD = D->load();
break;
}
case 4: {
auto D = reinterpret_cast<const std::atomic<uint32_t>*>(MemData);
GD = D->load();
break;
}
case 8: {
auto D = reinterpret_cast<const std::atomic<uint64_t>*>(MemData);
GD = D->load();
break;
}
default:
memcpy(GDP, MemData, IROp->Size);
break;
}
}
DEF_OP(StoreMem) {
auto Op = IROp->C<IR::IROp_StoreMem>();
uint8_t OpSize = IROp->Size;
uint8_t *MemData = *GetSrc<uint8_t **>(Data->SSAData, Op->Addr);
if (!Op->Offset.IsInvalid()) {
auto Offset = *GetSrc<uintptr_t const*>(Data->SSAData, Op->Offset) * Op->OffsetScale;
switch(Op->OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val: MemData += Offset; break;
case IR::MEM_OFFSET_UXTW.Val: MemData += (uint32_t)Offset; break;
case IR::MEM_OFFSET_SXTW.Val: MemData += (int32_t)Offset; break;
}
}
switch (OpSize) {
case 1: {
reinterpret_cast<std::atomic<uint8_t>*>(MemData)->store(*GetSrc<uint8_t*>(Data->SSAData, Op->Value));
break;
}
case 2: {
reinterpret_cast<std::atomic<uint16_t>*>(MemData)->store(*GetSrc<uint16_t*>(Data->SSAData, Op->Value));
break;
}
case 4: {
reinterpret_cast<std::atomic<uint32_t>*>(MemData)->store(*GetSrc<uint32_t*>(Data->SSAData, Op->Value));
break;
}
case 8: {
reinterpret_cast<std::atomic<uint64_t>*>(MemData)->store(*GetSrc<uint64_t*>(Data->SSAData, Op->Value));
break;
}
default:
memcpy(MemData, GetSrc<void*>(Data->SSAData, Op->Value), IROp->Size);
break;
}
}
DEF_OP(VLoadMemElement) {
auto Op = IROp->C<IR::IROp_VLoadMemElement>();
void const *MemData = *GetSrc<void const**>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[1]), 16);
memcpy(reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(GDP) + (Op->Header.ElementSize * Op->Index)),
MemData, Op->Header.ElementSize);
}
DEF_OP(VStoreMemElement) {
#define STORE_DATA(x, y) \
case x: { \
y *MemData = *GetSrc<y**>(Data->SSAData, Op->Header.Args[0]); \
memcpy(MemData, &GetSrc<y*>(Data->SSAData, Op->Header.Args[1])[Op->Index], sizeof(y)); \
break; \
}
auto Op = IROp->C<IR::IROp_VStoreMemElement>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
STORE_DATA(1, uint8_t)
STORE_DATA(2, uint16_t)
STORE_DATA(4, uint32_t)
STORE_DATA(8, uint64_t)
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size"); break;
}
#undef STORE_DATA
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
char *MemData = *GetSrc<char **>(Data->SSAData, Op->Addr);
// 64-byte cache line clear
CacheLineFlush(MemData);
}
#undef DEF_OP
void InterpreterOps::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
REGISTER_OP(FILLREGISTER, FillRegister);
REGISTER_OP(LOADFLAG, LoadFlag);
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
#undef REGISTER_OP
}
}
@@ -0,0 +1,158 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
namespace FEXCore::CPU {
[[noreturn]]
static void StopThread(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->StopThread(Thread);
LOGMAN_MSG_A_FMT("unreachable");
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val:
std::atomic_thread_fence(std::memory_order_acquire);
break;
case IR::Fence_LoadStore.Val:
std::atomic_thread_fence(std::memory_order_seq_cst);
break;
case IR::Fence_Store.Val:
std::atomic_thread_fence(std::memory_order_release);
break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
}
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 4: // HLT
StopThread(Data->State);
break;
default: LOGMAN_MSG_A_FMT("Unknown Break Reason: {}", Op->Reason); break;
}
}
DEF_OP(GetRoundingMode) {
uint32_t GuestRounding{};
#ifdef _M_ARM_64
uint64_t Tmp{};
__asm(R"(
mrs %[Tmp], FPCR;
)"
: [Tmp] "=r" (Tmp));
// Extract the rounding
// On ARM the ordering is different than on x86
GuestRounding |= ((Tmp >> 24) & 1) ? IR::ROUND_MODE_FLUSH_TO_ZERO : 0;
uint8_t RoundingMode = (Tmp >> 22) & 0b11;
if (RoundingMode == 0)
GuestRounding |= IR::ROUND_MODE_NEAREST;
else if (RoundingMode == 1)
GuestRounding |= IR::ROUND_MODE_POSITIVE_INFINITY;
else if (RoundingMode == 2)
GuestRounding |= IR::ROUND_MODE_NEGATIVE_INFINITY;
else if (RoundingMode == 3)
GuestRounding |= IR::ROUND_MODE_TOWARDS_ZERO;
#else
GuestRounding = _mm_getcsr();
// Extract the rounding
GuestRounding = (GuestRounding >> 13) & 0b111;
#endif
memcpy(GDP, &GuestRounding, sizeof(GuestRounding));
}
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
uint8_t GuestRounding = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
#ifdef _M_ARM_64
uint64_t HostRounding{};
__asm volatile(R"(
mrs %[Tmp], FPCR;
)"
: [Tmp] "=r" (HostRounding));
// Mask out the rounding
HostRounding &= ~(0b111 << 22);
HostRounding |= (GuestRounding & IR::ROUND_MODE_FLUSH_TO_ZERO) ? (1U << 24) : 0;
uint8_t RoundingMode = GuestRounding & 0b11;
if (RoundingMode == IR::ROUND_MODE_NEAREST)
HostRounding |= (0b00U << 22);
else if (RoundingMode == IR::ROUND_MODE_POSITIVE_INFINITY)
HostRounding |= (0b01U << 22);
else if (RoundingMode == IR::ROUND_MODE_NEGATIVE_INFINITY)
HostRounding |= (0b10U << 22);
else if (RoundingMode == IR::ROUND_MODE_TOWARDS_ZERO)
HostRounding |= (0b11U << 22);
__asm volatile(R"(
msr FPCR, %[Tmp];
)"
:: [Tmp] "r" (HostRounding));
#else
uint32_t HostRounding = _mm_getcsr();
// Cut out the host rounding mode
HostRounding &= ~(0b111 << 13);
// Insert our new rounding mode
HostRounding |= GuestRounding << 13;
_mm_setcsr(HostRounding);
#endif
}
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
uint8_t OpSize = IROp->Size;
if (OpSize <= 8) {
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
LogMan::Msg::IFmt(">>>> Value in Arg: 0x{:x}, {}", Src, Src);
}
else if (OpSize == 16) {
__uint128_t Src = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src0 = Src;
uint64_t Src1 = Src >> 64;
LogMan::Msg::IFmt(">>>> Value[0] in Arg: 0x{:x}, {}", Src0, Src0);
LogMan::Msg::IFmt(" Value[1] in Arg: 0x{:x}, {}", Src1, Src1);
}
else
LOGMAN_MSG_A_FMT("Unknown value size: {}", OpSize);
}
#undef DEF_OP
void InterpreterOps::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
#undef REGISTER_OP
}
}
@@ -0,0 +1,50 @@
/*
$info$
tags: backend|interpreter
$end_info$
*/
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
uintptr_t Src = GetSrc<uintptr_t>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP,
reinterpret_cast<void*>(Src + Op->Header.Size * Op->Element), Op->Header.Size);
}
DEF_OP(CreateElementPair) {
auto Op = IROp->C<IR::IROp_CreateElementPair>();
void *Src_Lower = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src_Upper = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
uint8_t *Dst = GetDest<uint8_t*>(Data->SSAData, Node);
memcpy(Dst, Src_Lower, Op->Header.Size);
memcpy(Dst + Op->Header.Size, Src_Upper, Op->Header.Size);
}
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
uint8_t OpSize = IROp->Size;
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[0]), OpSize);
}
#undef DEF_OP
void InterpreterOps::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
#undef REGISTER_OP
}
}
File diff suppressed because it is too large. Load diff
+15 -1
View File
@@ -163,7 +163,7 @@ DEF_OP(Mul) {
case 8:
mul(Dst, GetReg<RA_64>(Op->Header.Args[0].ID()), GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown Mul size: %d", OpSize);
default: LOGMAN_MSG_A_FMT("Unknown Mul size: {}", OpSize);
}
}
@@ -390,6 +390,19 @@ DEF_OP(And) {
}
}
DEF_OP(Andn) {
auto Op = IROp->C<IR::IROp_Andn>();
const auto& Lhs = Op->Header.Args[0];
const auto& Rhs = Op->Header.Args[1];
uint64_t Const{};
if (IsInlineConstant(Rhs, &Const)) {
bic(GRS(Node), GRS(Lhs.ID()), Const);
} else {
bic(GRS(Node), GRS(Lhs.ID()), GRS(Rhs.ID()));
}
}
DEF_OP(Xor) {
auto Op = IROp->C<IR::IROp_Xor>();
uint64_t Const;
@@ -1079,6 +1092,7 @@ void Arm64JITCore::RegisterALUHandlers() {
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
@@ -219,17 +219,17 @@ DEF_OP(AtomicAdd) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: staddlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -266,7 +266,7 @@ DEF_OP(AtomicAdd) {
cbnz(TMP2, &LoopTop);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -278,17 +278,17 @@ DEF_OP(AtomicSub) {
if (SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1: staddlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: staddlh(TMP2.W(), MemOperand(MemSrc)); break;
case 4: staddl(TMP2.W(), MemOperand(MemSrc)); break;
case 8: staddl(TMP2.X(), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -325,7 +325,7 @@ DEF_OP(AtomicSub) {
cbnz(TMP2, &LoopTop);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -337,17 +337,17 @@ DEF_OP(AtomicAnd) {
if (SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1: stclrlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: stclrlh(TMP2.W(), MemOperand(MemSrc)); break;
case 4: stclrl(TMP2.W(), MemOperand(MemSrc)); break;
case 8: stclrl(TMP2.X(), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -384,7 +384,7 @@ DEF_OP(AtomicAnd) {
cbnz(TMP2, &LoopTop);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -395,17 +395,17 @@ DEF_OP(AtomicOr) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: stsetlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -442,7 +442,7 @@ DEF_OP(AtomicOr) {
cbnz(TMP2, &LoopTop);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -453,17 +453,17 @@ DEF_OP(AtomicXor) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: steorlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -500,7 +500,7 @@ DEF_OP(AtomicXor) {
cbnz(TMP2, &LoopTop);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -512,17 +512,17 @@ DEF_OP(AtomicSwap) {
if (SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1: swplb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swplh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpl(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpl(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -559,7 +559,7 @@ DEF_OP(AtomicSwap) {
mov(GetReg<RA_64>(Node), TMP2.X());
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -569,17 +569,17 @@ DEF_OP(AtomicFetchAdd) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: ldaddalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -620,7 +620,7 @@ DEF_OP(AtomicFetchAdd) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -631,17 +631,17 @@ DEF_OP(AtomicFetchSub) {
if (SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1: ldaddalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -682,7 +682,7 @@ DEF_OP(AtomicFetchSub) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -693,17 +693,17 @@ DEF_OP(AtomicFetchAnd) {
if (SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1: ldclralb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldclralh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldclral(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldclral(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -744,7 +744,7 @@ DEF_OP(AtomicFetchAnd) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -754,17 +754,17 @@ DEF_OP(AtomicFetchOr) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: ldsetalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -805,7 +805,7 @@ DEF_OP(AtomicFetchOr) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -815,17 +815,17 @@ DEF_OP(AtomicFetchXor) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
switch (Op->Size) {
switch (IROp->Size) {
case 1: ldeoralb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
else {
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -866,7 +866,7 @@ DEF_OP(AtomicFetchXor) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
}
@@ -876,7 +876,7 @@ DEF_OP(AtomicFetchNeg) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
// TMP2-TMP3
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
aarch64::Label LoopTop;
bind(&LoopTop);
@@ -917,7 +917,7 @@ DEF_OP(AtomicFetchNeg) {
mov(GetReg<RA_64>(Node), TMP2);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
+27 -153
View File
@@ -55,7 +55,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
LoadConstant(x1, (uintptr_t)Info.fn);
blr(x1);
@@ -112,7 +112,12 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
if (Info.ABI == FABI_F80_I16) {
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
else {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
LoadConstant(x1, (uintptr_t)Info.fn);
blr(x1);
@@ -133,7 +138,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -153,7 +158,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -173,7 +178,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -192,7 +197,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -211,7 +216,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -230,10 +235,10 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(x3, GetSrc(IROp->Args[1].ID()).V2D(), 1);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
@@ -252,7 +257,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
@@ -273,10 +278,10 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(x3, GetSrc(IROp->Args[1].ID()).V2D(), 1);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
@@ -316,6 +321,9 @@ Arm64JITCore::CodeBuffer Arm64JITCore::AllocateNewCodeBuffer(size_t Size) {
-1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
Dispatcher->RegisterCodeBuffer(Buffer.Ptr, Buffer.Size);
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
@@ -324,146 +332,6 @@ void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
Dispatcher->RemoveCodeBuffer(Buffer.Ptr);
}
bool Arm64JITCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
uint32_t *PC = (uint32_t*)ArchHelpers::Context::GetPc(ucontext);
uint32_t Instr = PC[0];
if (!Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
// 1 = 16bit
// 2 = 32bit
// 3 = 64bit
uint32_t Size = (Instr & 0xC000'0000) >> 30;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DMB = 0b1101'0101'0000'0011'0011'0000'1011'1111 |
0b1011'0000'0000; // Inner shareable all
if ((Instr & 0x3F'FF'FC'00) == 0x08'DF'FC'00 || // LDAR*
(Instr & 0x3F'FF'FC'00) == 0x38'BF'C0'00) { // LDAPR*
if (ParanoidTSO()) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t LDR = 0b0011'1000'0111'1111'0110'1000'0000'0000;
LDR |= Size << 30;
LDR |= AddrReg << 5;
LDR |= DataReg;
PC[-1] = DMB;
PC[0] = LDR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ( (Instr & 0x3F'FF'FC'00) == 0x08'9F'FC'00) { // STLR*
if (ParanoidTSO()) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t STR = 0b0011'1000'0011'1111'0110'1000'0000'0000;
STR |= Size << 30;
STR |= AddrReg << 5;
STR |= DataReg;
PC[-1] = DMB;
PC[0] = STR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXP_MASK) == FEXCore::ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
//Should be compare and swap pair only. LDAXP not used elsewhere
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleCASPAL_ARMv8(ucontext, info, Instr);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::STLXP_MASK) == FEXCore::ArchHelpers::Arm64::STLXP_INST) { // STLXP
//Should not trigger - middle of an LDAXP/STAXP pair.
LogMan::Msg::EFmt("Unhandled JIT SIGBUS STLXP: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASAL: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(ucontext, info, Instr)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: {} Instruction: 0x{:08x}\n", Op, fmt::ptr(PC), PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXR_MASK) == FEXCore::ArchHelpers::Arm64::LDAXR_INST) { // LDAXR*
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleAtomicLoadstoreExclusive(ucontext, info);
if (BytesToSkip) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + BytesToSkip);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAXR: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[-1], 16);
return true;
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: Arm64Emitter(0)
, CTX {ctx}
@@ -538,7 +406,13 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->HandleSIGBUS(Signal, info, ucontext);
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
+33 -24
View File
@@ -42,22 +42,26 @@ public:
size_t Size;
};
explicit Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
~Arm64JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] std::string GetName() override { return "JIT"; }
bool NeedsOpDispatch() override { return true; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
void ClearCache() override;
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
CodeBuffer AllocateNewCodeBuffer(size_t Size);
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
@@ -95,35 +99,39 @@ private:
constexpr static uint8_t RA_FPR = 2;
template<uint8_t RAType>
aarch64::Register GetReg(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg(uint32_t Node) const;
template<>
aarch64::Register GetReg<RA_32>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_32>(uint32_t Node) const;
template<>
aarch64::Register GetReg<RA_64>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_64>(uint32_t Node) const;
template<uint8_t RAType>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node) const;
template<>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node) const;
template<>
std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node) const;
aarch64::VRegister GetSrc(uint32_t Node) const;
aarch64::VRegister GetDst(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetSrc(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetDst(uint32_t Node) const;
FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node) const;
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node) const;
IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
bool IsFPR(uint32_t Node) const;
bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
MemOperand GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
[[nodiscard]] MemOperand GenerateMemOperand(uint8_t AccessSize,
aarch64::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
[[nodiscard]] bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
[[nodiscard]] bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
struct LiveRange {
uint32_t Begin;
@@ -213,6 +221,7 @@ private:
DEF_OP(UMulH);
DEF_OP(Or);
DEF_OP(And);
DEF_OP(Andn);
DEF_OP(Xor);
DEF_OP(Lshl);
DEF_OP(Lshr);
@@ -261,7 +261,7 @@ DEF_OP(StoreRegister) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = Op->Size;
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[0].ID());
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -288,7 +288,7 @@ DEF_OP(LoadContextIndexed) {
ldr(GetReg<RA_64>(Node), MemOperand(TMP1, Op->BaseOffset));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -335,7 +335,7 @@ DEF_OP(LoadContextIndexed) {
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -349,7 +349,7 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
size_t size = Op->Size;
size_t size = IROp->Size;
auto index = GetReg<RA_64>(Op->Header.Args[1].ID());
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -378,7 +378,7 @@ DEF_OP(StoreContextIndexed) {
str(value, MemOperand(TMP1, Op->BaseOffset));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -427,7 +427,7 @@ DEF_OP(StoreContextIndexed) {
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -571,11 +571,11 @@ DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GenerateMemOperand(Op->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
auto Dst = GetReg<RA_64>(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 1:
ldrb(Dst, MemSrc);
break;
@@ -588,12 +588,12 @@ DEF_OP(LoadMem) {
case 8:
ldr(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
}
}
else {
auto Dst = GetDst(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 1:
ldr(Dst.B(), MemSrc);
break;
@@ -609,7 +609,7 @@ DEF_OP(LoadMem) {
case 16:
ldr(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
}
}
}
@@ -624,7 +624,7 @@ DEF_OP(LoadMemTSO) {
}
if (SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldaprb(Dst, MemSrc);
@@ -633,7 +633,7 @@ DEF_OP(LoadMemTSO) {
// Aligned
auto Dst = GetReg<RA_64>(Node);
nop();
switch (Op->Size) {
switch (IROp->Size) {
case 2:
ldaprh(Dst, MemSrc);
break;
@@ -643,13 +643,13 @@ DEF_OP(LoadMemTSO) {
case 8:
ldapr(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", IROp->Size);
}
nop();
}
}
else if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldarb(Dst, MemSrc);
@@ -658,7 +658,7 @@ DEF_OP(LoadMemTSO) {
// Aligned
auto Dst = GetReg<RA_64>(Node);
nop();
switch (Op->Size) {
switch (IROp->Size) {
case 2:
ldarh(Dst, MemSrc);
break;
@@ -668,7 +668,7 @@ DEF_OP(LoadMemTSO) {
case 8:
ldar(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", IROp->Size);
}
nop();
}
@@ -676,7 +676,7 @@ DEF_OP(LoadMemTSO) {
else {
dmb(InnerShareable, BarrierAll);
auto Dst = GetDst(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 2:
ldr(Dst.H(), MemSrc);
break;
@@ -689,7 +689,7 @@ DEF_OP(LoadMemTSO) {
case 16:
ldr(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", IROp->Size);
}
dmb(InnerShareable, BarrierAll);
}
@@ -700,10 +700,10 @@ DEF_OP(StoreMem) {
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GenerateMemOperand(Op->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
switch (Op->Size) {
switch (IROp->Size) {
case 1:
strb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
@@ -716,12 +716,12 @@ DEF_OP(StoreMem) {
case 8:
str(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
break;
@@ -737,7 +737,7 @@ DEF_OP(StoreMem) {
case 16:
str(Src, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
}
@@ -751,13 +751,13 @@ DEF_OP(StoreMemTSO) {
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
}
else {
nop();
switch (Op->Size) {
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
@@ -767,7 +767,7 @@ DEF_OP(StoreMemTSO) {
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
nop();
}
@@ -775,7 +775,7 @@ DEF_OP(StoreMemTSO) {
else {
dmb(InnerShareable, BarrierAll);
auto Src = GetSrc(Op->Header.Args[1].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
break;
@@ -791,7 +791,7 @@ DEF_OP(StoreMemTSO) {
case 16:
str(Src, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
dmb(InnerShareable, BarrierAll);
}
@@ -807,14 +807,14 @@ DEF_OP(ParanoidLoadMemTSO) {
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldarb(Dst, MemSrc);
}
else {
auto Dst = GetReg<RA_64>(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 2:
ldarh(Dst, MemSrc);
break;
@@ -824,13 +824,13 @@ DEF_OP(ParanoidLoadMemTSO) {
case 8:
ldar(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", IROp->Size);
}
}
}
else {
auto Dst = GetDst(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 2:
ldarh(TMP1.W(), MemSrc);
fmov(Dst.H(), TMP1.W());
@@ -850,7 +850,7 @@ DEF_OP(ParanoidLoadMemTSO) {
mov(Dst.V2D(), 0, TMP1);
mov(Dst.V2D(), 1, TMP2);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", IROp->Size);
}
}
}
@@ -864,12 +864,12 @@ DEF_OP(ParanoidStoreMemTSO) {
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
}
else {
switch (Op->Size) {
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
@@ -879,19 +879,19 @@ DEF_OP(ParanoidStoreMemTSO) {
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", IROp->Size);
}
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
if (Op->Size == 1) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
mov(TMP1.W(), Src.V16B(), 0);
stlrb(TMP1, MemSrc);
}
else {
switch (Op->Size) {
switch (IROp->Size) {
case 2:
mov(TMP1.W(), Src.V8H(), 0);
stlrh(TMP1, MemSrc);
@@ -911,15 +911,13 @@ DEF_OP(ParanoidStoreMemTSO) {
Label B;
bind(&B);
nop(); // < Overwritten with DMB
// ldaxp must not have both the destination registers be the same
ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS
nop(); // < Overwritten with DMB
ldaxp(xzr, TMP3, MemSrc); // <- Can hit SIGBUS. Overwritten with DMB
stlxp(TMP3, TMP1, TMP2, MemSrc); // <- Can also hit SIGBUS
cbnz(TMP3, &B); // < Overwritten with DMB
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", IROp->Size);
}
}
}
+8 -3
View File
@@ -13,6 +13,11 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
}
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
} // namespace FEXCore::CPU
@@ -15,6 +15,11 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define GRS(Node) (IROp->Size <= 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -417,6 +422,25 @@ DEF_OP(And) {
mov(Dst, rax);
}
DEF_OP(Andn) {
auto Op = IROp->C<IR::IROp_Andn>();
const auto& Lhs = Op->Header.Args[0];
const auto& Rhs = Op->Header.Args[1];
auto Dst = GRD(Node);
uint64_t Const{};
if (IsInlineConstant(Rhs, &Const)) {
mov(Dst, GRS(Lhs.ID()));
and_(Dst, ~Const);
} else {
const auto Temp = IROp->Size <= 4 ? Xbyak::Reg{rax.cvt32()} : Xbyak::Reg{rax};
mov(Temp, GRS(Rhs.ID()));
not_(Temp);
and_(Temp, GRS(Lhs.ID()));
mov(Dst, Temp);
}
}
DEF_OP(Xor) {
auto Op = IROp->C<IR::IROp_Xor>();
auto Dst = GetDst<RA_64>(Node);
@@ -1048,10 +1072,6 @@ DEF_OP(Sbfe) {
}
}
#define GRS(Node) (IROp->Size <= 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
DEF_OP(Select) {
auto Op = IROp->C<IR::IROp_Select>();
auto Dst = GRD(Node);
@@ -1221,6 +1241,7 @@ void X86JITCore::RegisterALUHandlers() {
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
@@ -121,7 +121,7 @@ DEF_OP(AtomicAdd) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
lock();
switch (Op->Size) {
switch (IROp->Size) {
case 1:
add(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -134,7 +134,7 @@ DEF_OP(AtomicAdd) {
case 8:
add(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
}
@@ -143,7 +143,7 @@ DEF_OP(AtomicSub) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
lock();
switch (Op->Size) {
switch (IROp->Size) {
case 1:
sub(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -156,7 +156,7 @@ DEF_OP(AtomicSub) {
case 8:
sub(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
}
@@ -165,7 +165,7 @@ DEF_OP(AtomicAnd) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
lock();
switch (Op->Size) {
switch (IROp->Size) {
case 1:
and_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -178,7 +178,7 @@ DEF_OP(AtomicAnd) {
case 8:
and_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
}
@@ -187,7 +187,7 @@ DEF_OP(AtomicOr) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
lock();
switch (Op->Size) {
switch (IROp->Size) {
case 1:
or_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -200,7 +200,7 @@ DEF_OP(AtomicOr) {
case 8:
or_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
}
@@ -209,7 +209,7 @@ DEF_OP(AtomicXor) {
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
lock();
switch (Op->Size) {
switch (IROp->Size) {
case 1:
xor_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -222,7 +222,7 @@ DEF_OP(AtomicXor) {
case 8:
xor_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
}
@@ -232,7 +232,7 @@ DEF_OP(AtomicSwap) {
Xbyak::Reg MemReg = rax;
mov(MemReg, GetSrc<RA_64>(Op->Header.Args[0].ID()));
switch (Op->Size) {
switch (IROp->Size) {
case 1:
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Header.Args[1].ID()));
lock();
@@ -253,7 +253,7 @@ DEF_OP(AtomicSwap) {
lock();
xchg(qword [MemReg], GetDst<RA_64>(Node));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicSwap size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicSwap size: {}", IROp->Size);
}
}
@@ -261,7 +261,7 @@ DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1:
movzx(rcx, GetSrc<RA_8>(Op->Header.Args[1].ID()));
lock();
@@ -286,7 +286,7 @@ DEF_OP(AtomicFetchAdd) {
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchAdd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchAdd size: {}", IROp->Size);
}
}
@@ -294,7 +294,7 @@ DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1:
mov(cl, GetSrc<RA_8>(Op->Header.Args[1].ID()));
neg(cl);
@@ -323,7 +323,7 @@ DEF_OP(AtomicFetchSub) {
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchSub size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchSub size: {}", IROp->Size);
}
}
@@ -333,7 +333,7 @@ DEF_OP(AtomicFetchAnd) {
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -401,7 +401,7 @@ DEF_OP(AtomicFetchAnd) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchAnd size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchAnd size: {}", IROp->Size);
}
}
@@ -410,7 +410,7 @@ DEF_OP(AtomicFetchOr) {
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -478,7 +478,7 @@ DEF_OP(AtomicFetchOr) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchOr size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchOr size: {}", IROp->Size);
}
}
@@ -487,7 +487,7 @@ DEF_OP(AtomicFetchXor) {
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -555,7 +555,7 @@ DEF_OP(AtomicFetchXor) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchXor size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchXor size: {}", IROp->Size);
}
}
@@ -563,7 +563,7 @@ DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -631,7 +631,7 @@ DEF_OP(AtomicFetchNeg) {
mov(GetDst<RA_64>(Node), TMP3.cvt64());
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchNeg size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled AtomicFetchNeg size: {}", IROp->Size);
}
}
+7 -8
View File
@@ -44,7 +44,7 @@ $end_info$
namespace FEXCore::CPU {
CodeBuffer AllocateNewCodeBuffer(size_t Size) {
CodeBuffer AllocateNewCodeBuffer(FEXCore::Context::Context *CTX, size_t Size) {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(
@@ -54,6 +54,9 @@ CodeBuffer AllocateNewCodeBuffer(size_t Size) {
MAP_PRIVATE | MAP_ANONYMOUS,
-1, 0));
LOGMAN_THROW_A_FMT(Buffer.Ptr != reinterpret_cast<uint8_t*>(~0ULL), "Couldn't allocate code buffer");
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
@@ -61,10 +64,6 @@ void FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
}
namespace FEXCore::CPU {
void X86JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
@@ -418,7 +417,7 @@ void X86JITCore::ClearCache() {
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MAX_CODE_SIZE);
InitialCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
InitialCodeBuffer = AllocateNewCodeBuffer(CTX, CurrentCodeBuffer->Size);
setNewBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
}
}
@@ -426,7 +425,7 @@ void X86JITCore::ClearCache() {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(X86JITCore::INITIAL_CODE_SIZE);
auto NewCodeBuffer = AllocateNewCodeBuffer(CTX, X86JITCore::INITIAL_CODE_SIZE);
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
@@ -788,6 +787,6 @@ uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateF
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
}
}
+27 -22
View File
@@ -30,14 +30,9 @@ struct CodeBuffer {
size_t Size;
};
CodeBuffer AllocateNewCodeBuffer(size_t Size);
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void FreeCodeBuffer(CodeBuffer Buffer);
}
namespace FEXCore::CPU {
// Temp registers
// rax, rcx, rdx, rsi, r8, r9,
// r10, r11
@@ -62,14 +57,22 @@ const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
explicit X86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
CodeBuffer Buffer,
bool CompileThread);
~X86JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] std::string GetName() override { return "JIT"; }
bool NeedsOpDispatch() override { return true; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
void ClearCache() override;
@@ -111,26 +114,27 @@ private:
constexpr static uint8_t RA_64 = 3;
constexpr static uint8_t RA_XMM = 4;
IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
bool IsFPR(uint32_t Node) const;
bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
template<uint8_t RAType>
Xbyak::Reg GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetSrc(uint32_t Node) const;
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node) const;
template<uint8_t RAType>
Xbyak::Reg GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetDst(uint32_t Node) const;
Xbyak::Xmm GetSrc(uint32_t Node) const;
Xbyak::Xmm GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(uint32_t Node) const;
Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
[[nodiscard]] Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
[[nodiscard]] bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
[[nodiscard]] bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
@@ -216,6 +220,7 @@ private:
DEF_OP(UMulH);
DEF_OP(Or);
DEF_OP(And);
DEF_OP(Andn);
DEF_OP(Xor);
DEF_OP(Lshl);
DEF_OP(Lshr);
@@ -142,7 +142,7 @@ DEF_OP(StoreContext) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = Op->Size;
size_t size = IROp->Size;
Reg index = GetSrc<RA_64>(Op->Header.Args[0].ID());
if (Op->Class.Val == 0) {
@@ -166,7 +166,7 @@ DEF_OP(LoadContextIndexed) {
mov(GetDst<RA_64>(Node), qword [rax + index * Op->Stride]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -202,7 +202,7 @@ DEF_OP(LoadContextIndexed) {
vmovq(GetDst(Node), qword [rax + index * Op->Stride]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -231,7 +231,7 @@ DEF_OP(LoadContextIndexed) {
movups(GetDst(Node), xword [STATE + rax]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
break;
}
break;
@@ -246,7 +246,7 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
Reg index = GetSrc<RA_64>(Op->Header.Args[1].ID());
size_t size = Op->Size;
size_t size = IROp->Size;
if (Op->Class.Val == 0) {
auto value = GetSrc<RA_64>(Op->Header.Args[0].ID());
@@ -258,9 +258,9 @@ DEF_OP(StoreContextIndexed) {
case 4:
case 8: {
if (!(size == 1 || size == 2 || size == 4 || size == 8)) {
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", Op->Size);
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", IROp->Size);
}
mov(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value);
mov(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
break;
}
default:
@@ -278,16 +278,16 @@ DEF_OP(StoreContextIndexed) {
lea(rax, dword [STATE + Op->BaseOffset]);
switch (size) {
case 1:
pextrb(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value, 0);
pextrb(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value, 0);
break;
case 2:
pextrw(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value, 0);
pextrw(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value, 0);
break;
case 4:
vmovd(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value);
vmovd(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
break;
case 8:
vmovq(AddressFrame(Op->Size * 8) [rax + index * Op->Stride], value);
vmovq(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", size);
@@ -301,16 +301,16 @@ DEF_OP(StoreContextIndexed) {
lea(rax, dword [rax + Op->BaseOffset]);
switch (size) {
case 1:
pextrb(AddressFrame(Op->Size * 8) [STATE + rax], value, 0);
pextrb(AddressFrame(IROp->Size * 8) [STATE + rax], value, 0);
break;
case 2:
pextrw(AddressFrame(Op->Size * 8) [STATE + rax], value, 0);
pextrw(AddressFrame(IROp->Size * 8) [STATE + rax], value, 0);
break;
case 4:
vmovd(AddressFrame(Op->Size * 8) [STATE + rax], value);
vmovd(AddressFrame(IROp->Size * 8) [STATE + rax], value);
break;
case 8:
vmovq(AddressFrame(Op->Size * 8) [STATE + rax], value);
vmovq(AddressFrame(IROp->Size * 8) [STATE + rax], value);
break;
case 16:
if (Op->BaseOffset % 16 == 0)
@@ -472,7 +472,7 @@ DEF_OP(LoadMem) {
if (Op->Class.Val == 0) {
auto Dst = GetDst<RA_64>(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
movzx (Dst, byte [MemPtr]);
}
@@ -489,14 +489,14 @@ DEF_OP(LoadMem) {
mov(Dst, qword [MemPtr]);
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
}
}
else
{
auto Dst = GetDst(Node);
switch (Op->Size) {
switch (IROp->Size) {
case 1: {
movzx(eax, byte [MemPtr]);
vmovd(Dst, eax);
@@ -516,7 +516,7 @@ DEF_OP(LoadMem) {
}
break;
case 16: {
if (Op->Size == Op->Align)
if (IROp->Size == Op->Align)
movups(GetDst(Node), xword [MemPtr]);
else
movups(GetDst(Node), xword [MemPtr]);
@@ -525,7 +525,7 @@ DEF_OP(LoadMem) {
}
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
}
}
}
@@ -538,7 +538,7 @@ DEF_OP(StoreMem) {
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class.Val == 0) {
switch (Op->Size) {
switch (IROp->Size) {
case 1:
mov(byte [MemPtr], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
@@ -551,11 +551,11 @@ DEF_OP(StoreMem) {
case 8:
mov(qword [MemPtr], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
else {
switch (Op->Size) {
switch (IROp->Size) {
case 1:
pextrb(byte [MemPtr], GetSrc(Op->Header.Args[1].ID()), 0);
break;
@@ -569,12 +569,12 @@ DEF_OP(StoreMem) {
vmovq(qword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
break;
case 16:
if (Op->Size == Op->Align)
if (IROp->Size == Op->Align)
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
else
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", Op->Size);
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
}
+219 -31
View File
@@ -26,33 +26,7 @@ $end_info$
namespace FEXCore::IR {
auto OpToIndex = [](uint8_t Op) constexpr -> uint8_t {
switch (Op) {
// Group 1
case 0x80: return 0;
case 0x81: return 1;
case 0x82: return 2;
case 0x83: return 3;
// Group 2
case 0xC0: return 0;
case 0xC1: return 1;
case 0xD0: return 2;
case 0xD1: return 3;
case 0xD2: return 4;
case 0xD3: return 5;
// Group 3
case 0xF6: return 0;
case 0xF7: return 1;
// Group 4
case 0xFE: return 0;
// Group 5
case 0xFF: return 0;
// Group 11
case 0xC6: return 0;
case 0xC7: return 1;
}
return 0;
};
using X86Tables::OpToIndex;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
@@ -2156,6 +2130,209 @@ void OpDispatchBuilder::ROLImmediateOp(OpcodeArgs) {
GenerateFlags_RotateLeftImmediate(Op, ALUOp, Dest, Shift);
}
void OpDispatchBuilder::ANDNBMIOp(OpcodeArgs) {
auto* Src1 = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto* Src2 = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
auto Dest = _Andn(Src2, Src1);
StoreResult(GPRClass, Op, Dest, -1);
GenerateFlags_Logical(Op, Dest, Src1, Src2);
}
void OpDispatchBuilder::BEXTRBMIOp(OpcodeArgs) {
// Essentially (Src1 >> Start) & ((1 << Length) - 1)
// along with some edge-case handling and flag setting.
auto* Src1 = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto* Src2 = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
const auto SrcSize = GetSrcSize(Op) * 8;
const auto MaxSrcBit = SrcSize - 1;
auto MaxSrcBitOp = _Constant(SrcSize, MaxSrcBit);
// Shift the operand down to the starting bit
auto Start = _Bfe(8, 0, Src2);
auto Shifted = _Lshr(Src1, Start);
// Shifts larger than operand size need to be set to zero.
auto SanitizedShifted = _Select(IR::COND_ULE,
Start, MaxSrcBitOp,
Shifted, _Constant(SrcSize, 0));
// Now handle the length specifier.
auto Length = _Bfe(8, 8, Src2);
auto SanitizedLength = _Select(IR::COND_ULE,
Length, MaxSrcBitOp,
Length, MaxSrcBitOp);
// Now build up the mask
// (1 << SanitizedLength) - 1
auto One = _Constant(SrcSize, 1);
auto Mask = _Sub(_Lshl(One, SanitizedLength), One);
// Now put it all together and make the result.
auto Dest = _And(SanitizedShifted, Mask);
// Finally store the result.
StoreResult(GPRClass, Op, Dest, -1);
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
SetRFLAG<X86State::RFLAG_AF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_SF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_CF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_OF_LOC>(_Constant(0));
// PF
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(_Constant(0));
}
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Dest, _Constant(0),
_Constant(1), _Constant(0));
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroOp);
}
void OpDispatchBuilder::BLSIBMIOp(OpcodeArgs) {
// Equivalent to performing: SRC & -SRC
auto* Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto NegatedSrc = _Neg(Src);
auto Result = _And(Src, NegatedSrc);
// ...and we're done. Painless!
StoreResult(GPRClass, Op, Result, -1);
// Now for the flags:
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
//
// AF and PF are documented as being in an undefined state after
// a BLSI operation, however, we choose to reliably clear them.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant((GetSrcSize(Op) * 8) - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::BLSMSKBMIOp(OpcodeArgs) {
// Equivalent to: (Src - 1) ^ Src
auto Zero = _Constant(0);
auto One = _Constant(1);
auto* Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _Xor(_Sub(Src, One), Src);
StoreResult(GPRClass, Op, Result, -1);
// Now for the flags.
SetRFLAG<X86State::RFLAG_ZF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
void OpDispatchBuilder::BLSRBMIOp(OpcodeArgs) {
// Equivalent to: (Src - 1) & Src
auto Zero = _Constant(0);
auto One = _Constant(1);
auto* Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _And(_Sub(Src, One), Src);
StoreResult(GPRClass, Op, Result, -1);
// Now for flags.
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant((GetSrcSize(Op) * 8) - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::RCROp1Bit(OpcodeArgs) {
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto Size = GetSrcSize(Op) * 8;
@@ -2584,8 +2761,7 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
Result = _Lshr(Dest, BitSelect);
OrderedNode *BitMask = _Lshl(_Constant(1), BitSelect);
BitMask = _Not(BitMask);
Dest = _And(Dest, BitMask);
Dest = _Andn(Dest, BitMask);
StoreResult(GPRClass, Op, Dest, -1);
}
else {
@@ -2606,10 +2782,10 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
// Now add the addresses together and load the memory
OrderedNode *MemoryLocation = _Add(Dest, Src);
OrderedNode *BitMask = _Lshl(_Constant(1), BitSelect);
BitMask = _Not(BitMask);
if (DestIsLockedMem(Op)) {
HandledLock = true;
BitMask = _Not(BitMask);
// XXX: Technically this can optimize to an AArch64 ldclralb
// We don't current support this IR op though
Result = _AtomicFetchAnd(MemoryLocation, BitMask, 1);
@@ -2621,7 +2797,7 @@ void OpDispatchBuilder::BTROp(OpcodeArgs) {
// Now shift in to the correct bit location
Result = _Lshr(Value, BitSelect);
Value = _And(Value, BitMask);
Value = _Andn(Value, BitMask);
_StoreMemAutoTSO(GPRClass, 1, MemoryLocation, Value, 1);
}
}
@@ -5823,9 +5999,20 @@ constexpr uint16_t PF_F2 = 3;
{OPD(2, 0b01, 0x78), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(2, 0b01, 0x79), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(2, 0b00, 0xF2), 1, &OpDispatchBuilder::ANDNBMIOp},
{OPD(2, 0b00, 0xF7), 1, &OpDispatchBuilder::BEXTRBMIOp},
};
#undef OPD
#define OPD(group, pp, opcode) (((group - X86Tables::InstType::TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
const std::vector<std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr>> VEXGroupTable = {
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b001), 1, &OpDispatchBuilder::BLSRBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b010), 1, &OpDispatchBuilder::BLSMSKBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b011), 1, &OpDispatchBuilder::BLSIBMIOp},
};
#undef OPD
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> EVEXTable = {
{0x10, 2, &OpDispatchBuilder::UnimplementedOp},
{0x59, 1, &OpDispatchBuilder::UnimplementedOp},
@@ -5886,6 +6073,7 @@ constexpr uint16_t PF_F2 = 3;
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38Table);
InstallToTable(FEXCore::X86Tables::H0F3ATableOps, H0F3ATable);
InstallToTable(FEXCore::X86Tables::VEXTableOps, VEXTable);
InstallToTable(FEXCore::X86Tables::VEXTableGroupOps, VEXGroupTable);
InstallToTable(FEXCore::X86Tables::EVEXTableOps, EVEXTable);
}
+11 -4
View File
@@ -324,6 +324,13 @@ public:
template<size_t ElementSize>
void PSIGN(OpcodeArgs);
// BMI Ops
void ANDNBMIOp(OpcodeArgs);
void BEXTRBMIOp(OpcodeArgs);
void BLSIBMIOp(OpcodeArgs);
void BLSMSKBMIOp(OpcodeArgs);
void BLSRBMIOp(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -566,16 +573,16 @@ private:
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _StoreMemTSO(ssa0, ssa1, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _StoreMemTSO(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
else
return _StoreMem(ssa0, ssa1, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _StoreMem(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _LoadMemTSO(ssa0, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _LoadMemTSO(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
else
return _LoadMem(ssa0, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _LoadMem(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
@@ -33,9 +33,7 @@ void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, uint32_t Tag) {
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
Mask = _Lshl(Mask, TopOffset);
// XXX: This Neg can be removed if we support BIC
Mask = _Not(Mask);
OrderedNode *NewFTW = _And(FTW, Mask);
OrderedNode *NewFTW = _Andn(FTW, Mask);
if (Tag != 0) {
auto TagVal = _Lshl(_Constant(Tag), TopOffset);
NewFTW = _Or(NewFTW, TagVal);
@@ -15,7 +15,7 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeBaseTables(Context::OperatingMode Mode) {
const U8U8InfoStruct BaseOpTable[] = {
static constexpr U8U8InfoStruct BaseOpTable[] = {
// Prefixes
// Operand size overide
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
@@ -234,7 +234,7 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
};
const U8U8InfoStruct BaseOpTable_64[] = {
static constexpr U8U8InfoStruct BaseOpTable_64[] = {
{0x06, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x0E, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x16, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -258,7 +258,7 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xEA, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
const U8U8InfoStruct BaseOpTable_32[] = {
static constexpr U8U8InfoStruct BaseOpTable_32[] = {
{0x06, 1, X86InstInfo{"PUSH ES", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x07, 1, X86InstInfo{"POP ES", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x0E, 1, X86InstInfo{"PUSH CS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
@@ -14,7 +14,7 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeDDDTables() {
const U8U8InfoStruct DDDNowOpTable[] = {
static constexpr U8U8InfoStruct DDDNowOpTable[] = {
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
@@ -14,7 +14,7 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeEVEXTables() {
const U16U8InfoStruct EVEXTable[] = {
static constexpr U16U8InfoStruct EVEXTable[] = {
{0x10, 1, X86InstInfo{"VMOVUPS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x11, 1, X86InstInfo{"VMOVUPS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x18, 1, X86InstInfo{"VBROADCASTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -19,7 +19,7 @@ void InitializeH0F38Tables() {
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
const U16U8InfoStruct H0F38Table[] = {
static constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -20,7 +20,7 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = 1;
const U16U8InfoStruct H0F3ATable[] = {
static constexpr U16U8InfoStruct H0F3ATable[] = {
{OPD(0, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -52,7 +52,7 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
};
const U16U8InfoStruct H0F3ATable_64[] = {
static constexpr U16U8InfoStruct H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRQ", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
@@ -21,7 +21,7 @@ void InitializeSecondaryGroupTables() {
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
const U16U8InfoStruct SecondaryExtensionOpTable[] = {
static constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
// GROUP 2
// GROUP 3
@@ -14,7 +14,7 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeSecondaryModRMTables() {
const U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
static constexpr U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
// REG /1
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
@@ -15,7 +15,7 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeSecondaryTables(Context::OperatingMode Mode) {
const U8U8InfoStruct TwoByteOpTable[] = {
static constexpr U8U8InfoStruct TwoByteOpTable[] = {
// Instructions
{0x00, 1, X86InstInfo{"", TYPE_GROUP_6, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
{0x01, 1, X86InstInfo{"", TYPE_GROUP_7, FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -266,7 +266,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x3F, 1, X86InstInfo{"ALTINST", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0, nullptr}},
};
const U8U8InfoStruct TwoByteOpTable_32[] = {
static constexpr U8U8InfoStruct TwoByteOpTable_32[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -274,7 +274,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
const U8U8InfoStruct TwoByteOpTable_64[] = {
static constexpr U8U8InfoStruct TwoByteOpTable_64[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -282,7 +282,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
const U8U8InfoStruct RepModOpTable[] = {
static constexpr U8U8InfoStruct RepModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -362,7 +362,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
const U8U8InfoStruct RepNEModOpTable[] = {
static constexpr U8U8InfoStruct RepNEModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -435,7 +435,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xF8, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
const U8U8InfoStruct OpSizeModOpTable[] = {
static constexpr U8U8InfoStruct OpSizeModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -14,7 +14,7 @@ using namespace InstFlags;
void InitializeVEXTables() {
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
const U16U8InfoStruct VEXTable[] = {
static constexpr U16U8InfoStruct VEXTable[] = {
// Map 0 (Reserved)
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -386,7 +386,7 @@ void InitializeVEXTables() {
{OPD(2, 0b01, 0xDE), 1, X86InstInfo{"VAESDEC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDF), 1, X86InstInfo{"VAESDECLAST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b00, 0xF2), 1, X86InstInfo{"ANDN", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b00, 0xF2), 1, X86InstInfo{"ANDN", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b00, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b01, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
@@ -399,7 +399,7 @@ void InitializeVEXTables() {
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b00, 0xF7), 1, X86InstInfo{"BEXTR", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b00, 0xF7), 1, X86InstInfo{"BEXTR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF7), 1, X86InstInfo{"SHLX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b10, 0xF7), 1, X86InstInfo{"SARX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -486,7 +486,7 @@ void InitializeVEXTables() {
#undef OPD
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
const U8U8InfoStruct VEXGroupTable[] = {
static constexpr U8U8InfoStruct VEXGroupTable[] = {
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
@@ -503,9 +503,9 @@ void InitializeVEXTables() {
{OPD(TYPE_VEX_GROUP_15, 1, 0b010), 1, X86InstInfo{"VLDMXCSR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 1, 0b011), 1, X86InstInfo{"VSTMXCSR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b001), 1, X86InstInfo{"BLSR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b010), 1, X86InstInfo{"BLSMSK", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b011), 1, X86InstInfo{"BLSI", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b001), 1, X86InstInfo{"BLSR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b010), 1, X86InstInfo{"BLSMSK", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b011), 1, X86InstInfo{"BLSI", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
};
#undef OPD
@@ -15,7 +15,7 @@ using namespace InstFlags;
void InitializeX87Tables() {
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
#define OPDReg(op, reg) (((op - 0xD8) << 8) | (reg << 3))
const U16U8InfoStruct X87OpTable[] = {
static constexpr U16U8InfoStruct X87OpTable[] = {
// 0xD8
{OPDReg(0xD8, 0), 1, X86InstInfo{"FADD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 1), 1, X86InstInfo{"FMUL", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
@@ -20,7 +20,7 @@ void InitializeXOPTables() {
constexpr uint16_t XOP_GROUP_9 = 1;
constexpr uint16_t XOP_GROUP_A = 2;
const U16U8InfoStruct XOPTable[] = {
static constexpr U16U8InfoStruct XOPTable[] = {
// Group 8
{OPD(XOP_GROUP_8, 0, 0x85), 1, X86InstInfo{"VPMAXSSWW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x86), 1, X86InstInfo{"VPMACSSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -106,7 +106,7 @@ void InitializeXOPTables() {
#undef OPD
#define OPD(subgroup, opcode) (((subgroup - 1) << 3) | (opcode))
const U8U8InfoStruct XOPGroupTable[] = {
static constexpr U8U8InfoStruct XOPGroupTable[] = {
// Group 1
{OPD(1, 1), 1, X86InstInfo{"BLCFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 2), 1, X86InstInfo{"BLSFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
+48 -18
View File
@@ -558,8 +558,10 @@
"HasDest": true,
"DestClass": "Complex",
"DestSize": "Size",
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint32_t", "BaseOffset",
"uint32_t", "Stride",
"RegisterClassType", "Class"
@@ -573,12 +575,15 @@
],
"OpClass": "Memory",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Value",
"Index"
],
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint32_t", "BaseOffset",
"uint32_t", "Stride",
"RegisterClassType", "Class"
@@ -644,6 +649,7 @@
],
"OpClass": "Memory",
"SSAArgs": "1",
"DestSize": "1",
"SSANames": [
"Value"
],
@@ -692,8 +698,10 @@
"Addr",
"Offset"
],
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint8_t", "Align",
"RegisterClassType", "Class",
"MemOffsetType", "OffsetType",
@@ -709,13 +717,16 @@
"HasSideEffects": true,
"OpClass": "Memory",
"SSAArgs": "3",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value",
"Offset"
],
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint8_t", "Align",
"RegisterClassType", "Class",
"MemOffsetType", "OffsetType",
@@ -735,8 +746,10 @@
"Addr",
"Offset"
],
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint8_t", "Align",
"RegisterClassType", "Class",
"MemOffsetType", "OffsetType",
@@ -750,13 +763,16 @@
"HasSideEffects": true,
"OpClass": "Memory",
"SSAArgs": "3",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value",
"Offset"
],
"HelperArgs": [
"uint8_t", "Size"
],
"Args": [
"uint8_t", "Size",
"uint8_t", "Align",
"RegisterClassType", "Class",
"MemOffsetType", "OffsetType",
@@ -948,6 +964,15 @@
"SSAArgs": "2"
},
"Andn": {
"Desc": ["Integer binary AND NOT. Performs the equivalent of Src1 & ~Src2"],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "std::max<uint8_t>(4, GetOpSize(ssa0))",
"SSAArgs": "2"
},
"Xor": {
"Desc": ["Integer binary exclusive or"
],
@@ -1280,11 +1305,12 @@
],
"OpClass": "Atomic",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1295,11 +1321,12 @@
],
"OpClass": "Atomic",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1310,11 +1337,12 @@
],
"OpClass": "Atomic",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1325,11 +1353,12 @@
],
"OpClass": "Atomic",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1340,11 +1369,12 @@
],
"OpClass": "Atomic",
"SSAArgs": "2",
"DestSize": "Size",
"SSANames": [
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1363,7 +1393,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1383,7 +1413,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1404,7 +1434,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1424,7 +1454,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1444,7 +1474,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1464,7 +1494,7 @@
"Addr",
"Value"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
@@ -1482,7 +1512,7 @@
"SSANames": [
"Addr"
],
"Args": [
"HelperArgs": [
"uint8_t", "Size"
]
},
+4 -4
View File
@@ -531,7 +531,7 @@ bool ConstProp::ConstantPropagation(IREmitter *IREmit, const IRListView& Current
auto AddressHeader = IREmit->GetOpHeader(Op->Header.Args[0]);
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8) {
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, Op->Size, AddressHeader);
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, IROp->Size, AddressHeader);
Op->OffsetType = OffsetType;
Op->OffsetScale = OffsetScale;
@@ -548,7 +548,7 @@ bool ConstProp::ConstantPropagation(IREmitter *IREmit, const IRListView& Current
auto AddressHeader = IREmit->GetOpHeader(Op->Header.Args[0]);
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8) {
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, Op->Size, AddressHeader);
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, IROp->Size, AddressHeader);
Op->OffsetType = OffsetType;
Op->OffsetScale = OffsetScale;
@@ -941,7 +941,7 @@ bool ConstProp::ConstantInlining(IREmitter *IREmit, const IRListView& CurrentIR)
uint64_t Constant2{};
if (Op->OffsetType == MEM_OFFSET_SXTX && IREmit->IsValueConstant(Op->Header.Args[1], &Constant2)) {
if (IsImmMemory(Constant2, Op->Size)) {
if (IsImmMemory(Constant2, IROp->Size)) {
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->Header.Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, IREmit->_InlineConstant(Constant2));
@@ -958,7 +958,7 @@ bool ConstProp::ConstantInlining(IREmitter *IREmit, const IRListView& CurrentIR)
uint64_t Constant2{};
if (Op->OffsetType == MEM_OFFSET_SXTX && IREmit->IsValueConstant(Op->Header.Args[2], &Constant2)) {
if (IsImmMemory(Constant2, Op->Size)) {
if (IsImmMemory(Constant2, IROp->Size)) {
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->Header.Args[2]));
IREmit->ReplaceNodeArgument(CodeNode, 2, IREmit->_InlineConstant(Constant2));
@@ -257,19 +257,19 @@ namespace {
size_t ClassifiedStructSize{};
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
for (auto &it : *ContextClassification) {
LOGMAN_THROW_A(it.Class.Offset == ContextClassificationInfo->Lookup.size(), "Offset missmatch %d %d", it.Class.Offset == ContextClassificationInfo->Lookup.size());
LOGMAN_THROW_A_FMT(it.Class.Offset == ContextClassificationInfo->Lookup.size(), "Offset mismatch (offset={})", it.Class.Offset);
for (int i = 0; i < it.Class.Size; i++) {
ContextClassificationInfo->Lookup.push_back(&it);
}
ClassifiedStructSize += it.Class.Size;
}
LOGMAN_THROW_A(ClassifiedStructSize == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! %ld != %ld",
LOGMAN_THROW_A_FMT(ClassifiedStructSize == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! {} (classified) != {} (real)",
ClassifiedStructSize, sizeof(FEXCore::Core::CPUState));
LOGMAN_THROW_A(ContextClassificationInfo->Lookup.size() == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! %ld != %ld",
LOGMAN_THROW_A_FMT(ContextClassificationInfo->Lookup.size() == sizeof(FEXCore::Core::CPUState),
"Classified lookup size doesn't match real CPUState struct size! {} (classified) != {} (real)",
ContextClassificationInfo->Lookup.size(), sizeof(FEXCore::Core::CPUState));
}
@@ -1225,14 +1225,16 @@ namespace FEXCore::IR {
uint32_t ConstrainedRAPass::FindSpillSlot(uint32_t Node, FEXCore::IR::RegisterClassType RegisterClass) {
RegisterNode *CurrentNode = &Graph->Nodes[Node];
LiveRange *NodeLiveRange = &LiveRanges[Node];
for (uint32_t i = 0; i < Graph->SpillStack.size(); ++i) {
SpillStackUnit *SpillUnit = &Graph->SpillStack.at(i);
if (NodeLiveRange->Begin <= SpillUnit->SpillRange.End &&
SpillUnit->SpillRange.Begin <= NodeLiveRange->End) {
SpillUnit->SpillRange.Begin = std::min(SpillUnit->SpillRange.Begin, LiveRanges[Node].Begin);
SpillUnit->SpillRange.End = std::max(SpillUnit->SpillRange.End, LiveRanges[Node].End);
CurrentNode->Head.SpillSlot = i;
return i;
if (ReuseSpillSlots) {
for (uint32_t i = 0; i < Graph->SpillStack.size(); ++i) {
SpillStackUnit *SpillUnit = &Graph->SpillStack.at(i);
if (NodeLiveRange->Begin <= SpillUnit->SpillRange.End &&
SpillUnit->SpillRange.Begin <= NodeLiveRange->End) {
SpillUnit->SpillRange.Begin = std::min(SpillUnit->SpillRange.Begin, LiveRanges[Node].Begin);
SpillUnit->SpillRange.End = std::max(SpillUnit->SpillRange.End, LiveRanges[Node].End);
CurrentNode->Head.SpillSlot = i;
return i;
}
}
}
@@ -53,6 +53,9 @@ class RegisterAllocationPass : public FEXCore::IR::Pass {
protected:
bool HasSpills {};
// Debug option to disable split slot reuse
// Can be useful for testing if there is a bug with spill slots
constexpr static bool ReuseSpillSlots {true};
uint32_t SpillSlotCount {};
bool HadFullRA {};
};
+147
View File
@@ -1,5 +1,7 @@
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <sys/mman.h>
#ifdef ENABLE_JEMALLOC
#include <jemalloc/jemalloc.h>
@@ -85,4 +87,149 @@ namespace FEXCore::Allocator {
}
#pragma GCC diagnostic pop
FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
static constexpr std::array<uintptr_t, 7> TLBSizes = {
57,
52,
48,
47,
42,
39,
36,
};
for (auto Bits : TLBSizes) {
uintptr_t Size = 1ULL << Bits;
// Just try allocating
// We can't actually determine VA size on ARM safely
auto Find = [](uintptr_t Size) -> bool {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void *Ptr = ::mmap(reinterpret_cast<void*>(Size - PAGE_SIZE * i), PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, PAGE_SIZE);
if (Ptr == (void*)(Size - PAGE_SIZE * i)) {
return true;
}
}
}
return false;
};
if (Find(Size)) {
return Bits;
}
}
LOGMAN_MSG_A_FMT("Couldn't determine host VA size");
FEX_UNREACHABLE;
}
PtrCache* StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
PtrCache *Cache{};
uint64_t CacheSize{};
uint64_t CurrentCacheOffset = 0;
constexpr std::array<size_t, 10> ReservedVMARegionSizes = {{
// Anything larger than 64GB fails out
64ULL * 1024 * 1024 * 1024, // 64GB
32ULL * 1024 * 1024 * 1024, // 32GB
16ULL * 1024 * 1024 * 1024, // 16GB
4ULL * 1024 * 1024 * 1024, // 4GB
1ULL * 1024 * 1024 * 1024, // 1GB
512ULL * 1024 * 1024, // 512MB
128ULL * 1024 * 1024, // 128MB
32ULL * 1024 * 1024, // 32MB
1ULL * 1024 * 1024, // 1MB
4096ULL // One page
}};
constexpr size_t AllocationSizeMaxIndex = ReservedVMARegionSizes.size() - 1;
uint64_t CurrentSizeIndex = 0;
int PROT_FLAGS = PROT_READ | PROT_WRITE;
for (size_t MemoryOffset = Begin; MemoryOffset < End;) {
size_t AllocationSize = ReservedVMARegionSizes[CurrentSizeIndex];
size_t MemoryOffsetUpper = MemoryOffset + AllocationSize;
// If we would go above the upper bound on size then try the next size
if (MemoryOffsetUpper > End) {
++CurrentSizeIndex;
continue;
}
void *Ptr = ::mmap(reinterpret_cast<void*>(MemoryOffset), AllocationSize, PROT_FLAGS, MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE | MAP_FIXED_NOREPLACE, -1, 0);
// If we managed to allocate and not get the address we want then unmap it
// This happens with kernels older than 4.17
if (reinterpret_cast<uintptr_t>(Ptr) + AllocationSize > End) {
::munmap(Ptr, AllocationSize);
Ptr = reinterpret_cast<void*>(~0ULL);
}
// If we failed to allocate and we are on the smallest allocation size then just continue onward
// This page was unmappable
if (reinterpret_cast<uintptr_t>(Ptr) == ~0ULL && CurrentSizeIndex == AllocationSizeMaxIndex) {
CurrentSizeIndex = 0;
MemoryOffset += AllocationSize;
continue;
}
// Congratulations we were able to map this bit
// Reset and claim it was available
if (reinterpret_cast<uintptr_t>(Ptr) != ~0ULL) {
if (!Cache) {
Cache = reinterpret_cast<PtrCache *>(Ptr);
CacheSize = AllocationSize;
PROT_FLAGS = PROT_NONE;
}
else {
Cache[CurrentCacheOffset] = {
.Ptr = static_cast<uint64_t>(reinterpret_cast<uint64_t>(Ptr)),
.Size = static_cast<uint64_t>(AllocationSize)
};
++CurrentCacheOffset;
}
CurrentSizeIndex = 0;
MemoryOffset += AllocationSize;
continue;
}
// Couldn't allocate at this size
// Increase and continue
++CurrentSizeIndex;
}
Cache[CurrentCacheOffset] = {
.Ptr = static_cast<uint64_t>(reinterpret_cast<uint64_t>(Cache)),
.Size = CacheSize,
};
return Cache;
}
PtrCache* Steal48BitVA() {
size_t Bits = FEXCore::Allocator::DetermineVASize();
if (Bits < 48) {
return nullptr;
}
uintptr_t Begin48BitVA = 0x0'8000'0000'0000ULL;
uintptr_t End48BitVA = 0x1'0000'0000'0000ULL;
return StealMemoryRegion(Begin48BitVA, End48BitVA);
}
void ReclaimMemoryRegion(PtrCache* Regions) {
if (Regions == nullptr) {
return;
}
for (size_t i = 0;; ++i) {
void *Ptr = reinterpret_cast<void*>(Regions[i].Ptr);
size_t Size = Regions[i].Size;
::munmap(Ptr, Size);
if (Ptr == Regions) {
break;
}
}
}
}
+9 -128
View File
@@ -1,6 +1,7 @@
#include "Utils/Allocator/FlexBitSet.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/Allocator/IntrusiveArenaAllocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <algorithm>
@@ -139,49 +140,14 @@ namespace Alloc::OSAllocator {
}
// 32-bit old kernel workarounds
struct PtrCache {
uint32_t Ptr;
uint32_t Size;
};
PtrCache *Steal32BitIfOldKernel();
void Clear32BitOnOldKernel(PtrCache *Base);
FEXCore::Allocator::PtrCache *Steal32BitIfOldKernel();
};
void OSAllocator_64Bit::DetermineVASize() {
static constexpr std::array<uintptr_t, 7> TLBSizes = {
1ULL << 57,
1ULL << 52,
1ULL << 48,
1ULL << 47,
1ULL << 42,
1ULL << 39,
1ULL << 36,
};
for (auto Size : TLBSizes) {
// Just try allocating
// We can't actually determine VA size on ARM safely
auto Find = [](uintptr_t Size) -> bool {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void *Ptr = ::mmap(reinterpret_cast<void*>(Size - PAGE_SIZE * i), PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, PAGE_SIZE);
if (Ptr == (void*)(Size - PAGE_SIZE * i)) {
return true;
}
}
}
return false;
};
if (Find(Size)) {
UPPER_BOUND = Size;
UPPER_BOUND_PAGE = UPPER_BOUND / PAGE_SIZE;
break;
}
}
size_t Bits = FEXCore::Allocator::DetermineVASize();
uintptr_t Size = 1ULL << Bits;
UPPER_BOUND = Size;
UPPER_BOUND_PAGE = UPPER_BOUND / PAGE_SIZE;
}
void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
@@ -523,7 +489,7 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
return 0;
}
OSAllocator_64Bit::PtrCache *OSAllocator_64Bit::Steal32BitIfOldKernel() {
FEXCore::Allocator::PtrCache *OSAllocator_64Bit::Steal32BitIfOldKernel() {
// First calculate kernel version
struct utsname buf{};
if (uname(&buf) == -1) {
@@ -548,95 +514,10 @@ OSAllocator_64Bit::PtrCache *OSAllocator_64Bit::Steal32BitIfOldKernel() {
return nullptr;
}
OSAllocator_64Bit::PtrCache *Cache{};
uint32_t CacheSize{};
uint32_t CurrentCacheOffset = 0;
constexpr std::array<size_t, 6> ReservedVMARegionSizes = {{
1ULL * 1024 * 1024 * 1024, // 1GB
512ULL * 1024 * 1024, // 512MB
128ULL * 1024 * 1024, // 128MB
32ULL * 1024 * 1024, // 32MB
1ULL * 1024 * 1024, // 1MB
4096ULL // One page
}};
constexpr size_t AllocationSizeMaxIndex = ReservedVMARegionSizes.size() - 1;
uint64_t CurrentSizeIndex = 0;
constexpr size_t LOWER_BOUND_32 = 0x1'0000;
constexpr size_t UPPER_BOUND_32 = LOWER_BOUND;
for (size_t MemoryOffset = LOWER_BOUND_32; MemoryOffset < UPPER_BOUND_32;) {
size_t AllocationSize = ReservedVMARegionSizes[CurrentSizeIndex];
size_t MemoryOffsetUpper = MemoryOffset + AllocationSize;
// If we would go above the upper bound on size then try the next size
if (MemoryOffsetUpper > UPPER_BOUND_32) {
++CurrentSizeIndex;
continue;
}
void *Ptr = ::mmap(reinterpret_cast<void*>(MemoryOffset), AllocationSize, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0);
// If we managed to allocate and not get the address we want then unmap it
// This happens with kernels older than 4.17
if (reinterpret_cast<uintptr_t>(Ptr) + AllocationSize > UPPER_BOUND_32) {
::munmap(Ptr, AllocationSize);
Ptr = reinterpret_cast<void*>(~0ULL);
}
// If we failed to allocate and we are on the smallest allocation size then just continue onward
// This page was unmappable
if (reinterpret_cast<uintptr_t>(Ptr) == ~0ULL && CurrentSizeIndex == AllocationSizeMaxIndex) {
CurrentSizeIndex = 0;
MemoryOffset += AllocationSize;
continue;
}
// Congratulations we were able to map this bit
// Reset and claim it was available
if (reinterpret_cast<uintptr_t>(Ptr) != ~0ULL) {
if (!Cache) {
Cache = reinterpret_cast<OSAllocator_64Bit::PtrCache *>(Ptr);
CacheSize = AllocationSize;
}
else {
Cache[CurrentCacheOffset] = {
.Ptr = static_cast<uint32_t>(reinterpret_cast<uint64_t>(Ptr)),
.Size = static_cast<uint32_t>(AllocationSize)
};
++CurrentCacheOffset;
}
CurrentSizeIndex = 0;
MemoryOffset += AllocationSize;
continue;
}
// Couldn't allocate at this size
// Increase and continue
++CurrentSizeIndex;
}
Cache[CurrentCacheOffset] = {
.Ptr = static_cast<uint32_t>(reinterpret_cast<uint64_t>(Cache)),
.Size = CacheSize,
};
return Cache;
}
void OSAllocator_64Bit::Clear32BitOnOldKernel(OSAllocator_64Bit::PtrCache *Base) {
if (Base == nullptr) {
return;
}
for (size_t i = 0;; ++i) {
void *Ptr = reinterpret_cast<void*>(Base[i].Ptr);
size_t Size = Base[i].Size;
::munmap(Ptr, Size);
if (Ptr == Base) {
break;
}
}
return FEXCore::Allocator::StealMemoryRegion(LOWER_BOUND_32, UPPER_BOUND_32);
}
OSAllocator_64Bit::OSAllocator_64Bit() {
@@ -735,7 +616,7 @@ OSAllocator_64Bit::OSAllocator_64Bit() {
++CurrentSizeIndex;
}
Clear32BitOnOldKernel(ArrayPtr);
FEXCore::Allocator::ReclaimMemoryRegion(ArrayPtr);
}
OSAllocator_64Bit::~OSAllocator_64Bit() {
+51
View File
@@ -0,0 +1,51 @@
#include <string>
#include <vector>
#include <filesystem>
#include <fstream>
namespace FEXCore::FileLoading {
bool LoadFile(std::vector<char> &Data, const std::string &Filepath, size_t FixedSize) {
std::fstream ConfigFile;
ConfigFile.open(Filepath, std::ios::in);
if (!ConfigFile.is_open()) {
return false;
}
size_t FileSize{};
if (FixedSize == 0) {
if (!ConfigFile.seekg(0, std::fstream::end)) {
return false;
}
FileSize = ConfigFile.tellg();
if (ConfigFile.fail()) {
return false;
}
if (!ConfigFile.seekg(0, std::fstream::beg)) {
return false;
}
}
else {
FileSize = FixedSize;
}
if (FileSize > 0) {
Data.resize(FileSize);
if (!ConfigFile.read(&Data.at(0), FileSize)) {
// Probably means permissions aren't set. Just early exit
return false;
}
ConfigFile.close();
}
else {
return false;
}
return true;
}
}
+18
View File
@@ -0,0 +1,18 @@
#pragma once
#include <string>
#include <vector>
#include <filesystem>
#include <fstream>
namespace FEXCore::FileLoading {
/**
* @brief Loads a filepath in to a vector of data
*
* @param Data The vector to load the file data in to
* @param Filepath The filepath to load
*
* @return true on file loaded, false on failure
*/
bool LoadFile(std::vector<char> &Data, const std::string &Filepath, size_t FixedSize = 0);
}
+30 -6
View File
@@ -11,6 +11,30 @@
#include <unordered_map>
namespace FEXCore::Config {
namespace Handler {
static inline std::string_view CoreHandler(std::string_view Value) {
if (Value == "irint")
return "0";
else if (Value == "irjit")
return "1";
#ifdef _M_X86_64
else if (Value == "host")
return "2";
#endif
return "1";
}
static inline std::string_view SMCCheckHandler(std::string_view Value) {
if (Value == "none")
return "0";
else if (Value == "mman")
return "1";
else if (Value == "full")
return "2";
return "0";
}
}
enum ConfigOption {
#define OPT_BASE(type, group, enum, json, default) CONFIG_##enum,
#include <FEXCore/Config/ConfigValues.inl>
@@ -95,13 +119,13 @@ namespace Type {
return &it->second.front();
}
void Set(ConfigOption Option, std::string Data) {
OptionMap[Option].emplace_back(std::move(Data));
void Set(ConfigOption Option, std::string_view Data) {
OptionMap[Option].emplace_back(std::string(Data));
}
void EraseSet(ConfigOption Option, std::string Data) {
void EraseSet(ConfigOption Option, std::string_view Data) {
Erase(Option);
Set(Option, std::move(Data));
Set(Option, std::string(Data));
}
void Erase(ConfigOption Option) {
@@ -129,9 +153,9 @@ namespace Type {
FEX_DEFAULT_VISIBILITY std::optional<LayerValue*> All(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<std::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string Data);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void EraseSet(ConfigOption Option, std::string Data);
FEX_DEFAULT_VISIBILITY void EraseSet(ConfigOption Option, std::string_view Data);
template<typename T>
class FEX_DEFAULT_VISIBILITY Value {
+7 -4
View File
@@ -36,7 +36,7 @@ class LLVMCore;
/**
* @return The name of this backend
*/
virtual std::string GetName() = 0;
[[nodiscard]] virtual std::string GetName() = 0;
/**
* @brief Tells this CPUBackend to compile code for the provided IR and DebugData
*
@@ -54,14 +54,17 @@ class LLVMCore;
* @return An executable function pointer that is theoretically compiled from this point.
* Is actually a function pointer of type `void (FEXCore::Core::ThreadState *Thread)
*/
virtual void *CompileCode(uint64_t Entry, FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) = 0;
[[nodiscard]] virtual void *CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) = 0;
/**
* @brief Function for mapping memory in to the CPUBackend's visible space. Allows setting up virtual mappings if required
*
* @return Currently unused
*/
virtual void *MapRegion(void *HostPtr, uint64_t GuestPtr, uint64_t Size) = 0;
[[nodiscard]] virtual void *MapRegion(void *HostPtr, uint64_t GuestPtr, uint64_t Size) = 0;
/**
* @brief This is post-setup initialization that is called just before code executino
@@ -79,7 +82,7 @@ class LLVMCore;
*
* @return true if it needs the IR
*/
virtual bool NeedsOpDispatch() = 0;
[[nodiscard]] virtual bool NeedsOpDispatch() = 0;
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
+67 -56
View File
@@ -59,13 +59,15 @@ constexpr uint32_t FLAG_OPADDR_MASK = (((1 << FLAG_OPADDR_STACKSIZE) - 1) << FLA
constexpr uint32_t FLAG_OPERAND_SIZE_LAST = 0b01;
constexpr uint32_t FLAG_WIDENING_SIZE_LAST = 0b10;
inline uint32_t GetSizeDstFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_DST_OFF) & SIZE_MASK; }
inline uint32_t GetSizeSrcFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_SRC_OFF) & SIZE_MASK; }
constexpr uint32_t GetSizeDstFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_DST_OFF) & SIZE_MASK; }
constexpr uint32_t GetSizeSrcFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_SRC_OFF) & SIZE_MASK; }
inline uint32_t GenSizeDstSize(uint32_t Size) { return Size << FLAG_SIZE_DST_OFF; }
inline uint32_t GenSizeSrcSize(uint32_t Size) { return Size << FLAG_SIZE_SRC_OFF; }
constexpr uint32_t GenSizeDstSize(uint32_t Size) { return Size << FLAG_SIZE_DST_OFF; }
constexpr uint32_t GenSizeSrcSize(uint32_t Size) { return Size << FLAG_SIZE_SRC_OFF; }
inline uint32_t GetOpAddr(uint32_t Flags, int Index) { return (((Flags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) >> (Index * 2)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1); }
constexpr uint32_t GetOpAddr(uint32_t Flags, uint32_t Index) {
return (((Flags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) >> (Index * 2)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
}
inline void PushOpAddr(uint32_t *Flags, uint32_t Flag) {
uint32_t TmpFlags = *Flags;
@@ -270,93 +272,102 @@ enum InstType {
};
namespace InstFlags {
constexpr uint32_t FLAGS_NONE = 0;
constexpr uint32_t FLAGS_DEBUG = (1 << 1);
constexpr uint32_t FLAGS_DEBUG_MEM_ACCESS = (1 << 2);
constexpr uint32_t FLAGS_SUPPORTS_REP = (1 << 3);
constexpr uint32_t FLAGS_BLOCK_END = (1 << 4);
constexpr uint32_t FLAGS_SETS_RIP = (1 << 5);
constexpr uint32_t FLAGS_DISPLACE_SIZE_MUL_2 = (1 << 6);
constexpr uint32_t FLAGS_DISPLACE_SIZE_DIV_2 = (1 << 7);
constexpr uint32_t FLAGS_SRC_SEXT = (1 << 8);
constexpr uint32_t FLAGS_MEM_OFFSET = (1 << 9);
using InstFlagType = uint64_t;
constexpr InstFlagType FLAGS_NONE = 0;
constexpr InstFlagType FLAGS_DEBUG = (1ULL << 1);
constexpr InstFlagType FLAGS_DEBUG_MEM_ACCESS = (1ULL << 2);
constexpr InstFlagType FLAGS_SUPPORTS_REP = (1ULL << 3);
constexpr InstFlagType FLAGS_BLOCK_END = (1ULL << 4);
constexpr InstFlagType FLAGS_SETS_RIP = (1ULL << 5);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_MUL_2 = (1ULL << 6);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_DIV_2 = (1ULL << 7);
constexpr InstFlagType FLAGS_SRC_SEXT = (1ULL << 8);
constexpr InstFlagType FLAGS_MEM_OFFSET = (1ULL << 9);
// Enables XMM based subflags
// Current reserved range for this SF is [10, 15]
constexpr uint32_t FLAGS_XMM_FLAGS = (1 << 10);
constexpr InstFlagType FLAGS_XMM_FLAGS = (1ULL << 10);
// X87 flags aliased to XMM flags selection
// Allows X87 instruction table that is abusing the flag for 64BIT selection to work
constexpr uint32_t FLAGS_X87_FLAGS = (1 << 10);
constexpr InstFlagType FLAGS_X87_FLAGS = (1ULL << 10);
// Non-XMM subflags
constexpr uint32_t FLAGS_SF_DST_RAX = (1 << 11);
constexpr uint32_t FLAGS_SF_DST_RDX = (1 << 12);
constexpr uint32_t FLAGS_SF_SRC_RAX = (1 << 13);
constexpr uint32_t FLAGS_SF_SRC_RCX = (1 << 14);
constexpr uint32_t FLAGS_SF_REX_IN_BYTE = (1 << 15);
constexpr InstFlagType FLAGS_SF_DST_RAX = (1ULL << 11);
constexpr InstFlagType FLAGS_SF_DST_RDX = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_RAX = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_SRC_RCX = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_REX_IN_BYTE = (1ULL << 15);
// XMM subflags
constexpr uint32_t FLAGS_SF_HIGH_XMM_REG = (1 << 11);
constexpr uint32_t FLAGS_SF_DST_GPR = (1 << 12);
constexpr uint32_t FLAGS_SF_SRC_GPR = (1 << 13);
constexpr uint32_t FLAGS_SF_MMX = (3 << 14); // MMX_DST | MMX_SRC
constexpr uint32_t FLAGS_SF_MMX_DST = (1 << 14);
constexpr uint32_t FLAGS_SF_MMX_SRC = (1 << 15);
constexpr InstFlagType FLAGS_SF_HIGH_XMM_REG = (1ULL << 11);
constexpr InstFlagType FLAGS_SF_DST_GPR = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_GPR = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_MMX_DST = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_MMX_SRC = (1ULL << 15);
constexpr InstFlagType FLAGS_SF_MMX = FLAGS_SF_MMX_DST | FLAGS_SF_MMX_SRC;
// Enables MODRM specific subflags
// Current reserved range for this SF is [14, 17]
constexpr uint32_t FLAGS_MODRM = (1 << 16);
constexpr InstFlagType FLAGS_MODRM = (1ULL << 16);
// With ModRM SF flag enabled
// Direction of ModRM. Dst ^ Src
// Set means destination is rm bits
// Unset means src is rm bits
constexpr uint32_t FLAGS_SF_MOD_DST = (1 << 17);
constexpr InstFlagType FLAGS_SF_MOD_DST = (1ULL << 17);
// If the instruction is restricted to mem or reg only
// 0b00 = Regular ModRM support
// 0b01 = Memory accesses only
// 0b10 = Register accesses only
// 0b11 = <Reserved>
constexpr uint32_t FLAGS_SF_MOD_MEM_ONLY = (1 << 18);
constexpr uint32_t FLAGS_SF_MOD_REG_ONLY = (1 << 19);
constexpr InstFlagType FLAGS_SF_MOD_MEM_ONLY = (1ULL << 18);
constexpr InstFlagType FLAGS_SF_MOD_REG_ONLY = (1ULL << 19);
// The secondary Opcode Map uses prefix bytes to overlay more instruction
// But some instructions need to ignore this overlay and consume these prefixes.
constexpr uint32_t FLAGS_NO_OVERLAY = (1 << 20);
constexpr InstFlagType FLAGS_NO_OVERLAY = (1ULL << 20);
// Some instructions partially ignore overlay
// Ignore OpSize (0x66) in this case
constexpr uint32_t FLAGS_NO_OVERLAY66 = (1 << 21);
constexpr InstFlagType FLAGS_NO_OVERLAY66 = (1ULL << 21);
// x87
constexpr uint32_t FLAGS_POP = (1 << 22);
constexpr InstFlagType FLAGS_POP = (1ULL << 22);
// Only SEXT if the instruction is operating in 64bit operand size
constexpr uint32_t FLAGS_SRC_SEXT64BIT = (1 << 23);
constexpr InstFlagType FLAGS_SRC_SEXT64BIT = (1ULL << 23);
constexpr uint32_t FLAGS_SIZE_DST_OFF = 26;
constexpr uint32_t FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
// Whether or not the instruction has a VEX prefix for the first source operand
constexpr InstFlagType FLAGS_VEX_1ST_SRC = (1ULL << 24);
// Whether or not the instruction has a VEX prefix for the second source operand
constexpr InstFlagType FLAGS_VEX_2ND_SRC = (1ULL << 25);
// Whether or not the instruction has a VEX prefix for the destination
constexpr InstFlagType FLAGS_VEX_DST = (1ULL << 26);
constexpr uint32_t SIZE_MASK = 0b111;
constexpr uint32_t SIZE_DEF = 0b000;
constexpr uint32_t SIZE_8BIT = 0b001;
constexpr uint32_t SIZE_16BIT = 0b010;
constexpr uint32_t SIZE_32BIT = 0b011;
constexpr uint32_t SIZE_64BIT = 0b100;
constexpr uint32_t SIZE_128BIT = 0b101;
constexpr uint32_t SIZE_256BIT = 0b110;
constexpr uint32_t SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
inline uint32_t GetSizeDstFlags(uint32_t Flags) { return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK; }
inline uint32_t GetSizeSrcFlags(uint32_t Flags) { return (Flags >> FLAGS_SIZE_SRC_OFF) & SIZE_MASK; }
constexpr InstFlagType SIZE_MASK = 0b111;
constexpr InstFlagType SIZE_DEF = 0b000;
constexpr InstFlagType SIZE_8BIT = 0b001;
constexpr InstFlagType SIZE_16BIT = 0b010;
constexpr InstFlagType SIZE_32BIT = 0b011;
constexpr InstFlagType SIZE_64BIT = 0b100;
constexpr InstFlagType SIZE_128BIT = 0b101;
constexpr InstFlagType SIZE_256BIT = 0b110;
constexpr InstFlagType SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
inline uint32_t GenFlagsDstSize(uint32_t Size) { return Size << FLAGS_SIZE_DST_OFF; }
inline uint32_t GenFlagsSrcSize(uint32_t Size) { return Size << FLAGS_SIZE_SRC_OFF; }
inline uint32_t GenFlagsSameSize(uint32_t Size) {return (Size << FLAGS_SIZE_DST_OFF) | (Size << FLAGS_SIZE_SRC_OFF); }
inline uint32_t GenFlagsSizes(uint32_t Dest, uint32_t Src) {return (Dest << FLAGS_SIZE_DST_OFF) | (Src << FLAGS_SIZE_SRC_OFF); }
constexpr InstFlagType GetSizeDstFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK; }
constexpr InstFlagType GetSizeSrcFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_SRC_OFF) & SIZE_MASK; }
constexpr InstFlagType GenFlagsDstSize(InstFlagType Size) { return Size << FLAGS_SIZE_DST_OFF; }
constexpr InstFlagType GenFlagsSrcSize(InstFlagType Size) { return Size << FLAGS_SIZE_SRC_OFF; }
constexpr InstFlagType GenFlagsSameSize(InstFlagType Size) { return (Size << FLAGS_SIZE_DST_OFF) | (Size << FLAGS_SIZE_SRC_OFF); }
constexpr InstFlagType GenFlagsSizes(InstFlagType Dest, InstFlagType Src) { return (Dest << FLAGS_SIZE_DST_OFF) | (Src << FLAGS_SIZE_SRC_OFF); }
// If it has an xmm subflag
#define HAS_XMM_SUBFLAG(x, flag) (((x) & (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag))) == (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag)))
@@ -365,7 +376,7 @@ inline uint32_t GenFlagsSizes(uint32_t Dest, uint32_t Src) {return (Dest << FLAG
#define HAS_NON_XMM_SUBFLAG(x, flag) (((x) & (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag))) == (flag))
}
auto OpToIndex = [](uint8_t Op) constexpr -> uint8_t {
constexpr uint8_t OpToIndex(uint8_t Op) {
switch (Op) {
// Group 1
case 0x80: return 0;
@@ -391,7 +402,7 @@ auto OpToIndex = [](uint8_t Op) constexpr -> uint8_t {
case 0xC7: return 1;
}
return 0;
};
}
using DecodedOp = DecodedInst const*;
using OpDispatchPtr = void (IR::OpDispatchBuilder::*)(DecodedOp);
@@ -418,7 +429,7 @@ void InstallDebugInfo();
struct X86InstInfo {
char const *Name;
InstType Type;
uint32_t Flags; ///< Must be larger than InstFlags enum
InstFlags::InstFlagType Flags; ///< Must be larger than InstFlags enum
uint8_t MoreBytes;
OpDispatchPtr OpcodeDispatcher;
#ifndef NDEBUG
+10 -4
View File
@@ -80,23 +80,29 @@ friend class FEXCore::IR::PassManager;
return _Bfi(ssa0, ssa1, Width, lsb, DestSize);
}
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Align = 1) {
return _StoreMem(ssa0, ssa1, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _StoreMem(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
IRPair<IROp_StoreMemTSO> _StoreMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Align = 1) {
return _StoreMemTSO(ssa0, ssa1, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _StoreMemTSO(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
IRPair<IROp_VStoreMemElement> _VStoreMemElement(uint8_t RegisterSize, uint8_t ElementSize, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Index, uint8_t Align = 1) {
return _VStoreMemElement(ssa0, ssa1, Index, Align, RegisterSize, ElementSize);
}
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
return _LoadMem(ssa0, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _LoadMem(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
IRPair<IROp_LoadMemTSO> _LoadMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
return _LoadMemTSO(ssa0, Invalid(), Size, Align, Class, MEM_OFFSET_SXTX, 1);
return _LoadMemTSO(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
}
IRPair<IROp_VLoadMemElement> _VLoadMemElement(uint8_t RegisterSize, uint8_t ElementSize, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Index, uint8_t Align = 1) {
return _VLoadMemElement(ssa0, ssa1, Index, Align, RegisterSize, ElementSize);
}
IRPair<IROp_LoadContextIndexed> _LoadContextIndexed(OrderedNode *ssa0, uint8_t Size, uint32_t BaseOffset, uint32_t Stride, RegisterClassType Class) {
return _LoadContextIndexed(ssa0, BaseOffset, Stride, Class, Size);
}
IRPair<IROp_StoreContextIndexed> _StoreContextIndexed(OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Size, uint32_t BaseOffset, uint32_t Stride, RegisterClassType Class) {
return _StoreContextIndexed(ssa0, ssa1, BaseOffset, Stride, Class, Size);
}
IRPair<IROp_Select> _Select(uint8_t Cond, OrderedNode *ssa0, OrderedNode *ssa1, OrderedNode *ssa2, OrderedNode *ssa3, uint8_t CompareSize = 0) {
if (CompareSize == 0)
CompareSize = std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(ssa0), GetOpSize(ssa1)));
+17
View File
@@ -21,4 +21,21 @@ namespace FEXCore::Allocator {
FEX_DEFAULT_VISIBILITY void SetupHooks();
FEX_DEFAULT_VISIBILITY void ClearHooks();
FEX_DEFAULT_VISIBILITY size_t DetermineVASize();
// 48-bit VA handling
struct PtrCache {
uint64_t Ptr;
uint64_t Size;
};
FEX_DEFAULT_VISIBILITY PtrCache* StealMemoryRegion(uintptr_t Begin, uintptr_t End);
FEX_DEFAULT_VISIBILITY void ReclaimMemoryRegion(PtrCache* Regions);
// When running a 64-bit executable on ARM then userspace guest only gets 47 bits of VA
// This is a feature of x86-64 where the kernel gets a full 128TB of VA space
// x86-64 canonical addresses with bit 48 set will sign extend the address (Ignoring LA57)
// AArch64 canonical addresses are only up to bits 48/52 with the remainder being other things
// Use this to reserve the top 128TB of VA so the guest never see it
// Returns nullptr on host VA < 48bits
FEX_DEFAULT_VISIBILITY PtrCache* Steal48BitVA();
}
+4
View File
@@ -32,6 +32,10 @@ namespace FEXCore::Telemetry {
TYPE_16BYTE_SPLIT,
TYPE_USES_VEX_OPS,
TYPE_USES_EVEX_OPS,
TYPE_CAS_16BIT_TEAR,
TYPE_CAS_32BIT_TEAR,
TYPE_CAS_64BIT_TEAR,
TYPE_CAS_128BIT_TEAR,
TYPE_LAST,
};
+2 -2
View File
@@ -60,8 +60,8 @@ with open(sys.argv[1]) as cpuinfo_file:
current_part = int(re.findall(r'0x[0-9A-F]+', line, re.I)[0], 16)
cpuinfo += {tuple([current_implementer, current_part])}
largest_big = "native"
largest_little = "native"
largest_big = "cortex-a57"
largest_little = "cortex-a53"
for core in cpuinfo:
if BigCoreIDs.get(core):
+1 -24
View File
@@ -1,34 +1,11 @@
#include "Common/ArgumentLoader.h"
#include <FEXCore/Config/Config.h>
#include "OptionParser.h"
#include "git_version.h"
#include <stdint.h>
namespace FEX::Handler {
std::string CoreHandler(std::string &Value) {
if (Value == "irint")
return "0";
else if (Value == "irjit")
return "1";
#ifdef _M_X86_64
else if (Value == "host")
return "2";
#endif
return "1";
}
std::string SMCCheckHandler(std::string &Value) {
if (Value == "none")
return "0";
else if (Value == "mman")
return "1";
else if (Value == "full")
return "2";
return "0";
}
}
namespace FEX::ArgLoader {
std::vector<std::string> RemainingArgs;
std::vector<std::string> ProgramArguments;
+5 -1
View File
@@ -474,12 +474,15 @@ int main(int argc, char **argv, char **const envp) {
return -ENOEXEC;
}
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_FILENAME, std::filesystem::canonical(Program));
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_FILENAME, std::filesystem::canonical(Program).string());
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Loader.Is64BitMode() ? "1" : "0");
std::unique_ptr<FEX::HLE::x32::MemAllocator> Allocator;
FEXCore::Allocator::PtrCache *Base48Bit{};
if (Loader.Is64BitMode()) {
// Destroy the 48th bit if it exists
Base48Bit = FEXCore::Allocator::Steal48BitVA();
if (!Loader.MapMemory([](void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
return FEXCore::Allocator::mmap(addr, length, prot, flags, fd, offset);
}, [](void *addr, size_t length) {
@@ -619,6 +622,7 @@ int main(int argc, char **argv, char **const envp) {
LogMan::Msg::UnInstallHandlers();
FEXCore::Allocator::ClearHooks();
FEXCore::Allocator::ReclaimMemoryRegion(Base48Bit);
// Allocator is now original system allocator
FEXCore::Telemetry::Shutdown(ProgramName);
@@ -365,7 +365,8 @@ uint64_t FileManager::Stat(const char *pathname, void *buf) {
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
// Stat follows symlinks
auto Path = GetEmulatedPath(SelfPath, true);
if (!Path.empty()) {
uint64_t Result = ::stat(Path.c_str(), reinterpret_cast<struct stat*>(buf));
if (Result != -1)
@@ -378,7 +379,8 @@ uint64_t FileManager::Lstat(const char *pathname, void *buf) {
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
// lstat does not follow symlinks
auto Path = GetEmulatedPath(SelfPath, false);
if (!Path.empty()) {
uint64_t Result = ::lstat(Path.c_str(), reinterpret_cast<struct stat*>(buf));
if (Result != -1)
@@ -392,7 +394,8 @@ uint64_t FileManager::Access(const char *pathname, [[maybe_unused]] int mode) {
auto NewPath = GetSelf(pathname);
const char *SelfPath = NewPath ? NewPath->c_str() : nullptr;
auto Path = GetEmulatedPath(SelfPath);
// Access follows symlinks
auto Path = GetEmulatedPath(SelfPath, true);
if (!Path.empty()) {
uint64_t Result = ::access(Path.c_str(), mode);
if (Result != -1)
+2 -1
View File
@@ -63,12 +63,13 @@ public:
void UpdatePID(uint32_t PID) { CurrentPID = PID; }
std::string GetEmulatedPath(const char *pathname, bool FollowSymlink = false);
private:
FEX::EmulatedFile::EmulatedFDManager EmuFD;
std::mutex FDLock;
std::unordered_map<int32_t, std::string> FDToNameMap;
std::string GetEmulatedPath(const char *pathname, bool FollowSymlink = false);
std::map<std::string, std::string, std::less<>> ThunkOverlays;
FEX_CONFIG_OPT(Filename, APP_FILENAME);
@@ -82,30 +82,8 @@ namespace FEX::HLE {
void SignalDelegator::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *Info, void *UContext) {
// Let the host take first stab at handling the signal
siginfo_t *SigInfo = static_cast<siginfo_t*>(Info);
SignalHandler &Handler = HostHandlers[Signal];
if (Signal == SIGCHLD) {
bool StopOrResume = SigInfo->si_code == CLD_STOPPED || SigInfo->si_code == CLD_CONTINUED || SigInfo->si_code == CLD_TRAPPED;
// Do some special handling around this signal
// If the guest has a signal handler installed with SA_NOCLDSTOP or SA_NOCHLDWAIT then
// handle carefully
if (Handler.GuestAction.sa_flags & SA_NOCLDSTOP &&
StopOrResume) {
// SA_NOCLDSTOP blocks SIGCHLD when si_code is CLD_STOPPED/CLD_CONTINUED/CLD_TRAPPED
// in that case, drop the signal
return;
}
if (Handler.GuestAction.sa_flags & SA_NOCLDWAIT) {
// Linux will still generate a signal for this
// POSIX leaves it unspecific
// "do not transform children in to zombies when they terminate"
// XXX: Handle this
}
}
ucontext_t* _context = (ucontext_t*)UContext;
// Remove the pending signal
+9 -4
View File
@@ -116,12 +116,14 @@ uint64_t ExecveHandler(const char *pathname, char* const* argv, char* const* env
std::error_code ec;
std::string RootFS = FEX::HLE::_SyscallHandler->RootFSPath();
// Check the rootfs if it is available first
if (pathname[0] == '/') {
Filename = RootFS + pathname;
bool exists = std::filesystem::exists(Filename, ec);
if (ec || !exists) {
auto Path = FEX::HLE::_SyscallHandler->FM.GetEmulatedPath(pathname, true);
if (!Path.empty() && std::filesystem::exists(Path, ec)) {
Filename = Path;
}
else {
Filename = pathname;
}
}
@@ -150,9 +152,12 @@ uint64_t ExecveHandler(const char *pathname, char* const* argv, char* const* env
// If we don't have the interpreter installed we need to be extra careful for ENOEXEC
// Reasoning is that if we try executing a file from FEXLoader then this process loses the ENOEXEC flag
// Kernel does its own checks for file format support for this
// We can only call execve directly if we both have an interpreter installed AND were ran with the interpreter
// If the user ran FEX through FEXLoader then we must go down the emulated path
ELFLoader::ELFContainer::ELFType Type = ELFLoader::ELFContainer::GetELFType(Filename);
uint64_t Result{};
if (FEX::HLE::_SyscallHandler->IsInterpreterInstalled() &&
FEX::HLE::_SyscallHandler->IsInterpreter() &&
(Type == ELFLoader::ELFContainer::ELFType::TYPE_X86_32 ||
Type == ELFLoader::ELFContainer::ELFType::TYPE_X86_64)) {
// If the FEX interpreter is installed then just execve the ELF file
@@ -16,6 +16,7 @@
#include <drm/nouveau_drm.h>
#include <drm/vc4_drm.h>
#include <drm/v3d_drm.h>
#include <drm/virtgpu_drm.h>
#include <sys/ioctl.h>
#define CPYT(x) val.x = x
@@ -0,0 +1,9 @@
_BASIC_META(DRM_IOCTL_VIRTGPU_MAP)
_BASIC_META(DRM_IOCTL_VIRTGPU_EXECBUFFER)
_BASIC_META(DRM_IOCTL_VIRTGPU_GETPARAM)
_BASIC_META(DRM_IOCTL_VIRTGPU_RESOURCE_CREATE)
_BASIC_META(DRM_IOCTL_VIRTGPU_RESOURCE_INFO)
_BASIC_META(DRM_IOCTL_VIRTGPU_TRANSFER_FROM_HOST)
_BASIC_META(DRM_IOCTL_VIRTGPU_TRANSFER_TO_HOST)
_BASIC_META(DRM_IOCTL_VIRTGPU_WAIT)
_BASIC_META(DRM_IOCTL_VIRTGPU_GET_CAPS)
@@ -409,6 +409,31 @@ namespace FEX::HLE::x32 {
return -EPERM;
}
uint32_t Virtio_Handler(int fd, uint32_t cmd, uint32_t args) {
switch (_IOC_NR(cmd)) {
#define _BASIC_META(x) case _IOC_NR(x):
#define _BASIC_META_VAR(x, args...) case _IOC_NR(x):
#define _CUSTOM_META(name, ioctl_num)
#define _CUSTOM_META_OFFSET(name, ioctl_num, offset)
// DRM
#include "Tests/LinuxSyscalls/x32/Ioctl/virtio_drm.inl"
{
uint64_t Result = ::ioctl(fd, cmd, args);
SYSCALL_ERRNO();
break;
}
default:
UnhandledIoctl("Virtio", fd, cmd, args);
return -EPERM;
break;
}
#undef _BASIC_META
#undef _BASIC_META_VAR
#undef _CUSTOM_META
#undef _CUSTOM_META_OFFSET
return -EPERM;
}
void AssignDeviceTypeToFD(int fd, drm_version const &Version) {
if (Version.name) {
if (strcmp(Version.name, "amdgpu") == 0) {
@@ -435,6 +460,9 @@ namespace FEX::HLE::x32 {
else if (strcmp(Version.name, "v3d") == 0) {
FDToHandler.SetFDHandler(fd, V3D_Handler);
}
else if (strcmp(Version.name, "virtio_gpu") == 0) {
FDToHandler.SetFDHandler(fd, Virtio_Handler);
}
else {
LogMan::Msg::E("Unknown DRM device: '%s'", Version.name);
}
@@ -589,6 +617,7 @@ namespace FEX::HLE::x32 {
#include "Tests/LinuxSyscalls/x32/Ioctl/nouveau_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/vc4_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/v3d_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/virtio_drm.inl"
#undef _BASIC_META
#undef _BASIC_META_VAR
+2 -2
View File
@@ -100,9 +100,9 @@ namespace FEX::HLE::x32 {
uint64_t Result = 0;
if (req) {
const struct timespec req64 = *req;
Result = ::nanosleep(&req64, &rem64);
Result = ::nanosleep(&req64, rem64_ptr);
} else {
Result = ::nanosleep(nullptr, &rem64);
Result = ::nanosleep(nullptr, rem64_ptr);
}
if (rem) {
@@ -5,6 +5,7 @@ $end_info$
*/
#include <FEXCore/Utils/LogManager.h>
#include "Tests/LinuxSyscalls/Syscalls.h"
#include "Tests/LinuxSyscalls/x64/Syscalls.h"
#include <errno.h>
+9 -6
View File
@@ -56,11 +56,11 @@ void MsgHandler(LogMan::DebugLevels Level, char const *Message) {
CharLevel = "???";
break;
}
printf("[%s] %s\n", CharLevel, Message);
fmt::print("[{}] {}\n", CharLevel, Message);
}
void AssertHandler(char const *Message) {
printf("[ASSERT] %s\n", Message);
fmt::print("[ASSERT] {}\n", Message);
}
int main(int argc, char **argv, char **const envp) {
@@ -73,7 +73,10 @@ int main(int argc, char **argv, char **const envp) {
auto Args = FEX::ArgLoader::Get();
LOGMAN_THROW_A(Args.size() > 1, "Not enough arguments");
if (Args.size() < 2) {
LogMan::Msg::EFmt("Not enough arguments");
return -1;
}
FEX::HarnessHelper::HarnessCodeLoader Loader{Args[0], Args[1].c_str()};
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Loader.Is64BitMode() ? "1" : "0");
@@ -111,7 +114,7 @@ int main(int argc, char **argv, char **const envp) {
return Allocator->munmap(addr, length);
})) {
// failed to map
LogMan::Msg::E("Failed to map 32-bit elf file.");
LogMan::Msg::EFmt("Failed to map 32-bit elf file.");
return -ENOEXEC;
}
}
@@ -140,8 +143,8 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Context::GetCPUState(CTX, &State);
bool Passed = !DidFault && Loader.CompareStates(&State, nullptr);
LogMan::Msg::I("Faulted? %s", DidFault ? "Yes" : "No");
LogMan::Msg::I("Passed? %s", Passed ? "Yes" : "No");
LogMan::Msg::IFmt("Faulted? {}", DidFault ? "Yes" : "No");
LogMan::Msg::IFmt("Passed? {}", Passed ? "Yes" : "No");
SyscallHandler.reset();
SignalDelegation.reset();
+2 -28
View File
@@ -23,6 +23,8 @@ $end_info$
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/LogManager.h>
using FEXCore::X86Tables::OpToIndex;
constexpr std::array<std::pair<int16_t, int16_t>, 3> Disp8Ranges = {{
{static_cast<int16_t>(-16), 16},
{static_cast<int16_t>(-128), static_cast<int16_t>(-112)},
@@ -79,34 +81,6 @@ uint32_t GetModRMMapping(uint32_t Register) {
return Register;
};
auto OpToIndex = [](uint8_t Op) constexpr -> uint8_t {
switch (Op) {
// Group 1
case 0x80: return 0;
case 0x81: return 1;
case 0x82: return 2;
case 0x83: return 3;
// Group 2
case 0xC0: return 0;
case 0xC1: return 1;
case 0xD0: return 2;
case 0xD1: return 3;
case 0xD2: return 4;
case 0xD3: return 5;
// Group 3
case 0xF6: return 0;
case 0xF7: return 1;
// Group 4
case 0xFE: return 0;
// Group 5
case 0xFF: return 0;
// Group 11
case 0xC6: return 0;
case 0xC7: return 1;
}
return 0;
};
auto PrimaryIndexToOp = [](uint16_t Op) constexpr -> uint32_t {
#define OPD(group, prefix, Reg) ((((group) - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
switch (Op & ~0b111) {
+3 -4
View File
@@ -129,15 +129,14 @@ namespace {
while (!INotifyShutdown) {
constexpr size_t DATA_SIZE = (16 * (sizeof(struct inotify_event) + NAME_MAX + 1));
char buf[DATA_SIZE];
struct timeval tv{};
// 50 ms
tv.tv_usec = 50000;
int Ret{};
do {
fd_set Set{};
FD_ZERO(&Set);
FD_SET(INotifyFD, &Set);
struct timeval tv{};
// 50 ms
tv.tv_usec = 50000;
Ret = select(INotifyFD + 1, &Set, nullptr, nullptr, &tv);
} while (Ret == 0 && INotifyFD != -1);
+3
View File
@@ -333,6 +333,8 @@ def print_types():
print_remaining_base_types()
# Structs third
print("#pragma GCC diagnostic push")
print("#pragma GCC diagnostic ignored \"-Wattributes\"") # Suppress error spam about GCC not recognizing fex-match annotations
for i in range(0, 2):
for StructName, Struct in StructDefs.items():
# First walk the struct members and ensure any dependency is already emitted
@@ -342,6 +344,7 @@ def print_types():
# Now print this struct
print_struct(Struct.Name)
print("#pragma GCC diagnostic pop")
# Walks the commands element in the XML and pulls out all functions
# This will be used to generate the thunks that we need to hit
+3 -2
View File
@@ -71,8 +71,8 @@ function(add_guest_lib_with_name NAME LIBNAME)
target_compile_options(${LIBNAME}-guest PRIVATE -DLIB_NAME=${LIBNAME} -DLIBLIB_NAME=lib${LIBNAME})
add_custom_target(${LIBNAME}-guest-install
COMMAND ${CMAKE_COMMAND} -E make_directory ${DATA_DIRECTORY}/GuestThunks/
COMMAND ${CMAKE_COMMAND} -E copy_if_different ${CMAKE_BINARY_DIR}/lib${LIBNAME}-guest.so ${DATA_DIRECTORY}/GuestThunks/)
COMMAND ${CMAKE_COMMAND} -E make_directory $ENV{DESTDIR}/${DATA_DIRECTORY}/GuestThunks/
COMMAND ${CMAKE_COMMAND} -E copy_if_different ${CMAKE_BINARY_DIR}/lib${LIBNAME}-guest.so $ENV{DESTDIR}/${DATA_DIRECTORY}/GuestThunks/)
add_dependencies(ThunkGuestsInstall ${LIBNAME}-guest-install)
endfunction()
@@ -240,3 +240,4 @@ add_guest_lib(xshmfence)
generate(libdrm thunks function_packs function_packs_public)
add_guest_lib(drm)
target_include_directories(drm-guest PRIVATE /usr/include/drm/)
target_include_directories(drm-guest PRIVATE /usr/include/libdrm/)
+3 -2
View File
@@ -60,8 +60,8 @@ function(add_host_lib_with_name NAME LIBNAME)
target_compile_options(${LIBNAME}-host PRIVATE -DLIB_NAME=${LIBNAME} -DLIBLIB_NAME=lib${LIBNAME})
add_custom_target(${LIBNAME}-host-install
COMMAND ${CMAKE_COMMAND} -E make_directory ${DATA_DIRECTORY}/HostThunks/
COMMAND ${CMAKE_COMMAND} -E copy_if_different ${CMAKE_BINARY_DIR}/lib${LIBNAME}-host.so ${DATA_DIRECTORY}/HostThunks/)
COMMAND ${CMAKE_COMMAND} -E make_directory $ENV{DESTDIR}/${DATA_DIRECTORY}/HostThunks/
COMMAND ${CMAKE_COMMAND} -E copy_if_different ${CMAKE_BINARY_DIR}/lib${LIBNAME}-host.so $ENV{DESTDIR}/${DATA_DIRECTORY}/HostThunks/)
add_dependencies(ThunkHostsInstall ${LIBNAME}-host-install)
endfunction()
@@ -159,3 +159,4 @@ add_host_lib(xshmfence)
generate(libdrm function_unpacks tab_function_unpacks ldr ldr_ptrs)
add_host_lib(drm)
target_include_directories(drm-host PRIVATE /usr/include/drm/)
target_include_directories(drm-host PRIVATE /usr/include/libdrm/)
+8
View File
@@ -0,0 +1,8 @@
#pragma once
#include <cstdint>
struct CBWork {
uintptr_t cb;
void *argsv;
};
+14 -1
View File
@@ -1,4 +1,4 @@
# FEX-2110
# FEX-2111
## External/FEXCore
See [FEXCore/Readme.md](../External/FEXCore/Readme.md) for more details
@@ -29,6 +29,19 @@ IR to host code generation
- [MoveOps.cpp](../External/FEXCore/Source/Interface/Core/JIT/Arm64/MoveOps.cpp)
- [VectorOps.cpp](../External/FEXCore/Source/Interface/Core/JIT/Arm64/VectorOps.cpp)
#### interpreter
- [ALUOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/ALUOps.cpp)
- [AtomicOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/AtomicOps.cpp)
- [BranchOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/BranchOps.cpp)
- [ConversionOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/ConversionOps.cpp)
- [EncryptionOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/EncryptionOps.cpp)
- [F80Ops.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/F80Ops.cpp)
- [FlagOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/FlagOps.cpp)
- [MemoryOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/MemoryOps.cpp)
- [MiscOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/MiscOps.cpp)
- [MoveOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/MoveOps.cpp)
- [VectorOps.cpp](../External/FEXCore/Source/Interface/Core/Interpreter/VectorOps.cpp)
#### shared
- [CPUBackend.h](../External/FEXCore/include/FEXCore/Core/CPUBackend.h)
+22
View File
@@ -0,0 +1,22 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0xFFFFFFFFFFFFFFFF",
"RBX": "0",
"RCX": "0xFFFFFFFF",
"RDX": "0"
}
}
%endif
mov rax, 0
mov rbx, -1
andn rax, rax, rbx
andn rbx, rbx, rax
mov rcx, 0
mov rdx, -1
andn ecx, ecx, edx
andn edx, edx, ecx
hlt
+34
View File
@@ -0,0 +1,34 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x7F",
"RBX": "0",
"RDX": "0xFF",
"RSI": "0"
}
}
%endif
; General extraction
mov rax, 0x7FFFFFFFFFFFFFFF
mov rbx, 0x838 ; Start at bit 56 and extract 8 bits
bextr rax, rax, rbx ; This results in 0x7F being placed into RAX
; Extraction with 0 bits should clear the destination
mov rbx, -1
mov rcx, 0
bextr rbx, rbx, rcx
; Same tests as above but with 32-bit registers
; General extraction
mov rdx, 0x7FFFFFFFFFFFFFFF
mov rsi, 0x818 ; Start at bit 24 and extract 8 bits
bextr edx, edx, esi ; This results in 0xFF being placed into EDX
; Extraction with 0 bits should clear RSI to 0
mov rsi, -1
mov rdi, 0
bextr esi, esi, edi
hlt
+34
View File
@@ -0,0 +1,34 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "1",
"RBX": "0xFF00000000000000",
"RCX": "0x0100000000000000",
"RDX": "1",
"RSI": "0xFF000000",
"RDI": "0x01000000"
}
}
%endif
; Trivial test, this should result in 1.
mov rax, 11
blsi rax, rax
; Results in the lowest set bit (bit 56) being extracted
mov rbx, 0xFF00000000000000
mov rcx, 0
blsi rcx, rbx
; Same tests but with 32-bit registers
; Trivial test, this should result in 1.
mov edx, 11
blsi edx, edx
; Results in the lowest set bit (bit 24) being extracted
mov rsi, 0xFF000000
mov rdi, 0
blsi edi, esi
hlt
Loaded 100 of 108 files, more files were not shown because too many files have changed in this diff. Show more