Compare commits

..
213 Commits
Author SHA1 Message Date
Ryan Houdek d0943322b8 Docs: Update for release FEX-2112 2021-12-05 17:44:04 -08:00
Ryan Houdek db0740becc Merge pull request #1421 from lioncash/symbols
JitSymbols: Make use of fmt
2021-12-03 00:41:05 -08:00
lioncash ec8f077e60 JitSymbols: Take std::string_view instead of std::string
Now that we're using fmt, we can make the API itself non-allocating and
allow passing any kind of character buffer to it.
2021-12-03 02:49:30 -05:00
lioncash 2b190f1713 JitSymbols: Make HostAddr const
We can seamlessly format const pointers, so we can allow this in the
interface.
2021-12-03 02:49:30 -05:00
lioncash 69413d51bf JitSymbols: Make use of fmt
Makes the constructed strings a little quicker to read.
2021-12-03 02:49:26 -05:00
Ryan Houdek 8c38f936da Merge pull request #1420 from lioncash/const
IR: Make some interface functions accept pointers to const
2021-12-02 21:14:05 -08:00
lioncash e3fbe48f7c JitSymbols: Make use of unique_ptr
Ensures that the FILE pointer will always be handled.
2021-12-02 23:58:28 -05:00
lioncash c1acd9ad59 IR: Mark relevant API functions as [[nodiscard]] 2021-12-02 23:48:53 -05:00
lioncash 381d07707c IntrusiveIRList: Mark APIs [[nodiscard]] where applicable
Allows us to be more aggressively warned by the compiler in cases it's
definitely a bug to ignore the returned value.
2021-12-02 23:38:34 -05:00
lioncash a07e77b028 IntrusiveIRList: Remove some redundant reinterpret_casts
GetListData() and GetData() already return uintptr_t values, so we don't
need to cast these anymore.
2021-12-02 23:32:40 -05:00
lioncash 6b5ccb8244 IR: Make some interface functions accept pointers to const
Some interface functions are just querying state, so we can allow them
being used with const qualified data.
2021-12-02 23:32:36 -05:00
Ryan Houdek 597e4e1d57 Merge pull request #1419 from lioncash/ra-cast
RegisterAllocationPass: Resolve sign comparison mismatch in CalculateNodeInterference()
2021-12-02 20:22:08 -08:00
lioncash cdab71e612 RegisterAllocationPass: Resolve sign comparison mismatch in CalculateNodeInterference()
GetSSACount() returns a size_t, rather than a signed value.
2021-12-02 18:13:01 -05:00
Ryan Houdek 8ef0278338 Merge pull request #1418 from lioncash/branch
JIT: Eliminate redundant jump target map lookups
2021-12-02 14:45:44 -08:00
lioncash 98f54eab9c Arm64/BranchOps: Mark MapBranchCC as static
This isn't used outside of this translation unit, so we can make that
explicit.
2021-12-02 17:18:44 -05:00
Ryan Houdek 919e2bc3bd Merge pull request #1408 from Sonicadvance1/add_timetrace
CMake: Adds option to enable time-trace compile option
2021-12-02 14:14:38 -08:00
Ryan Houdek ffee22d3fd Merge pull request #1417 from lioncash/cpuid
CPUID: Call handler functions directly
2021-12-02 14:14:20 -08:00
lioncash e648e90c2e JIT: Eliminate redundant jump target map lookups
try_emplace() will return the existing entry in the map if it
exists already, so we don't need to perform a find() and then
try_emplace(). We can just use try_emplace() by itself.
2021-12-02 17:03:04 -05:00
Ryan Houdek b93871ff55 Merge pull request #1416 from lioncash/strong
IR: Convert NodeID into a strong type
2021-12-02 13:39:24 -08:00
lioncash 352f0a1133 CPUID: Call handler functions directly
Given the calls are all internal, we can store the function pointers
directly and call them with the this pointer instead of indirecting
through std::bind and std::function.
2021-12-02 16:34:43 -05:00
Ryan Houdek 6c93b6f775 Merge pull request #1412 from Sonicadvance1/socket_logger
Adds Socket logger and tool
2021-12-02 13:32:04 -08:00
lioncash f806577c46 RegisterAllocationPass: Use std::optional instead of magic value
Instead of using UINT32_MAX to communicate not found, we can return an
empty optional instance.
2021-12-02 15:17:19 -05:00
lioncash fac4022302 IR: Convert NodeID into a strong type
Converts the NodeID alias into a strongly-typed structure, which will
allow catching attempts to pass NodeIDs into incorrect APIs at
compile-time.
2021-12-02 15:04:57 -05:00
lioncash b5b3cb252a BucketList: Pass T to Next explicitly
Without this, the internal BucketList type will always be allocated with
T as a uint32_t, due to the default type for T, even if T is specified
differently in other code.
2021-12-02 14:55:44 -05:00
Ryan Houdek d5bf6414f3 Merge pull request #1413 from Sonicadvance1/fix_sched_getaffinity
Linux: Fixes sched_getaffinity
2021-11-30 18:06:04 -08:00
Ryan Houdek fc858ca0a5 Adds FEXLogServer
This tool works in both gui-less and gui modes.
FEXLogServer executed alone will open a socket on `localhost:8087` and
pass log output to stderr.

If passed the `-g` argument then it will open the IMGui UI which has
some more options.
It's fairly basic right now but eventually should allow filtering by
PIDs, TIDs, and log level type
2021-11-30 14:58:06 -08:00
Ryan Houdek 8ddf1ccbfe Linux: Fixes sched_getaffinity
Noticed an issue where some applications were thinking that applications
were running on a system with 512 CPU cores.
Specifically freedreno/turnip was trying to create 512 worker threads
which would cause my ARM devices to run out of memory.

Looks like sysconf behaviour has changed around this where it is
actually expecting the uint64_t alignment that the Linux kernel does.
So follow that behaviour as well and resolve the issue.
2021-11-30 14:53:10 -08:00
Ryan Houdek f2292c08e9 FEXLoader: Adds Socket logging path
Adds a new OutputSocket config option that when set will force all logs
to go through a socket.

This allows us to avoid polluting the guest application's output through
a socket instead of a file. Alleviating the issue of a file output
overwriting when multiple applications are ran.
2021-11-30 12:05:45 -08:00
Ryan Houdek b1b616087b Merge pull request #1380 from neobrain/fix_thunk_fails
Thunks: Fail loudly if thunking is enabled for a library that isn't installed
2021-11-30 11:45:11 -08:00
Tony Wasserka 5542360948 Thunks: Fail if thunking is enabled for a library that isn't installed
Previously, the host thunk would seemingly succeed to load even when
the corresponding host library could not be loaded, leading to obscure
crashes later throughout execution.
2021-11-30 17:39:01 +01:00
Tony Wasserka 7490d15369 Thunks: Clean up helper macro formatting 2021-11-30 17:39:01 +01:00
Tony Wasserka 68587cd02c unittests/ASM: Drop OP_THUNK test
This test doesn't increase coverage significantly, since OP_THUNK is called
with an invalid library name. An ideal test should verify that thunk symbols
are loaded properly, whereas currently it's only ensured the opcode 0xF 0x3F
is recognized by FEX at all. That's better than nothing, but a regression
here would likely show up in other tests anyway.
2021-11-30 17:39:01 +01:00
Ryan Houdek a9e1d5f2a2 FEX/Common: Adds a SocketLogging namespace
This will be used for sending messages over a socket
2021-11-29 13:14:33 -08:00
Ryan Houdek a8b9b351cb FEXConfig: Splits GUI setup in to a header
This will be used in another tool soon
2021-11-29 13:14:33 -08:00
Ryan Houdek 559f7128c3 Netstream: Use MSG_NOSIGNAL to avoid signals on socket
If the socket was closed then Linux would raise a EPIPE signal.
We don't want this, so avoid it with this flag
2021-11-29 13:14:33 -08:00
Ryan Houdek 5c1f14a155 FEXCore: Moves NetStream to Utils
The Frontend will want to use this
2021-11-29 13:14:32 -08:00
Ryan Houdek 82ebdf52be Merge pull request #1411 from Sonicadvance1/non_fata_32bit_mapping
FEXLoader: Change 32-bit memory check error to warning
2021-11-27 00:02:12 -08:00
Ryan Houdek 86684f9033 FEXLoader: Change 32-bit memory check error to warning
Instead of early exiting, allow the application to continue running but
throw error messages anyway. Should allow some users to still run FEX
even if an application steals a page in the lower 32-bits
2021-11-26 23:52:38 -08:00
Ryan Houdek e6d285a716 Merge pull request #1410 from Sonicadvance1/print_memory_map
FEXLoader: Print memory map when 32-bit intersect happens
2021-11-26 22:56:05 -08:00
Ryan Houdek 5c98c782f6 FEXLoader: Print memory map when 32-bit intersect happens
To get more information as to what intersected
2021-11-26 22:46:46 -08:00
Ryan Houdek 9b43f94f51 Merge pull request #1407 from lioncash/irgen
IR.json/json_ir_generator: Minor touchups to generated IR utilities
2021-11-26 15:01:00 -08:00
Ryan Houdek 19a0651944 Merge pull request #1406 from lioncash/regnode
RegisterAllocationPass: Reduce usages of NodeIDs where applicable
2021-11-26 14:57:08 -08:00
Ryan Houdek 436d20d660 CMake: Adds option to enable time-trace compile option
Useful for running something like aras-p/ClangBuildAnalyzer on the
source and finding compile bottlenecks
2021-11-26 14:43:39 -08:00
Ryan Houdek e027b83a06 Merge pull request #1399 from Sonicadvance1/gdbstub_improvements
GDBStub: Fixes a few hangs and crashes
2021-11-26 13:52:47 -08:00
lioncash d0aa785e6f json_ir_generator: Mark generated IR functions [[nodiscard]] where applicable
Marks generated functions as [[nodiscard]] where it would otherwise be a
logic bug to ignore their return values.
2021-11-26 14:09:17 -05:00
lioncash b2f44d607f json_ir_generator: Add helper functions for accessing header SSA args
In several parts of the IR we have cases where code accesses header SSA
arguments via:

Op->Header.Args[index]

which can be quite noisy when repeated quite a bit.

This adds member functions to the IR ops that have SSA arguments, that
allow it to simply be:

Op->Args(index)
2021-11-26 13:59:04 -05:00
lioncash 0aa53f726a json_ir_generator: Remove extra tab for enum entries
Makes the generated indentation consistent with the rest of the file.
2021-11-26 13:28:25 -05:00
lioncash b7e2dcc6c1 json_ir_generator: Make some utilities take a pointer to const
A few of these functions are just querying state or returning it, so we
can allow pointers to const to make the functions a little more
flexible.
2021-11-26 13:24:14 -05:00
lioncash 90e9010b17 json_ir_generator: Use _v equivalent of type traits
Same behavior, less writing.
2021-11-26 13:09:34 -05:00
lioncash c62945e615 IR.json: Remove semicolons from MEM_OFFSET constants
Also avoids some -Wextra-semi warnings, given the generation script will
add semicolons for these lines.
2021-11-26 12:57:17 -05:00
lioncash c78976adf1 json_ir_generator: Don't print out semicolons for empty entries
These can cause -Wextra-semi warnings on higher warning levels, so we
can just print out a newline and move on.
2021-11-26 12:55:34 -05:00
Ryan Houdek 13ad5b66bb Merge pull request #1405 from lioncash/interparr
InterpreterOps: Mark F80CMP array as static constexpr
2021-11-26 09:38:59 -08:00
lioncash 25b9bcc652 IR.json: Mark TypeDefinition instances as constexpr
Create() is constexpr, so we can also mark the constants as such to
avoid potentially having static constructors.
2021-11-26 12:23:38 -05:00
lioncash e2e2687a85 IR.json: Remove static from constants
Namespace-scope variables declared as const or constexpr have internal
linkage by default.

Also makes the output a little more consistent in terms of const/static
ordering.
2021-11-26 12:20:53 -05:00
lioncash 2d3aca2398 RegisterAllocationPass: Eliminate trivial casts
We can just specify LiveRange** as the return type of the lambdas to
allow the compiler to assume all returns are of this type.
2021-11-26 12:06:54 -05:00
lioncash a1d1674033 RegisterAllocatorPass: Convert macro functions into concrete ones
Gets rid of some preprocessor leakage across the file and makes the
functions more statically typed.
2021-11-26 12:06:54 -05:00
lioncash a281b6b38b RegisterAllocationPass: Reduce usages of NodeIDs where applicable
Deduplicates some code that makes use of NodeIDs to make migration to a
strongly-typed NodeID a little more straightforward.
2021-11-26 12:06:50 -05:00
lioncash 9e0c652a9a RegisterAllocationPass: Extract remat cost calc to its own function
Keeps it nicely separated from the live-range calculation code.
2021-11-26 10:52:55 -05:00
lioncash fdc0ce7101 InterpreterOps: Mark GetFallbackInfo as static
This is only used within this translation unit.
2021-11-26 10:31:41 -05:00
lioncash 600b1ad88a InterpreterOps: Mark F80CMP array as static constexpr
We don't need to construct this every time this code is executed.
2021-11-26 10:26:37 -05:00
Ryan Houdek b45603d174 Merge pull request #1404 from FEX-Emu/skmp/fix-steamwebhelper
jit/arm64: COND_FGT should map to gt, not hi, as it is not FGTU
2021-11-25 12:41:35 -08:00
Stefanos Kornilios Mitsis Poiitidis 28cb1240b4 jit/arm64: COND_FGT should map to gt, not hi, as it is not FGTU 2021-11-25 22:32:02 +02:00
Ryan Houdek 63f9e0b410 Merge pull request #1403 from lioncash/interp
Interpreter: Build op handler table at compile-time
2021-11-24 22:33:31 -08:00
lioncash e4c1a285ce Interpreter: Build op handler table at compile-time
We can create the table at compile-time, since we already have the
function labels available ahead of time.

This also centralizes the op table in one read-only location
2021-11-25 00:18:04 -05:00
Ryan Houdek 1f306d666c Merge pull request #1401 from lioncash/nodeid
IR: Add type alias for Node IDs
2021-11-24 17:12:58 -08:00
Ryan Houdek baadf0b98e Merge pull request #1402 from lioncash/bucket
BucketList: Minor API touchups
2021-11-24 17:07:44 -08:00
lioncash 0c697af5fb Passes: Make use of NodeID alias where applicable
Unifies the passes so that they're using the NodeID alias as well.
2021-11-24 17:15:09 -05:00
lioncash e6ad608226 IR: Add alias for Node IDs
In quite a few places we have a raw primitive to represent an IR node's
ID. This can make reading some bits of the API (or the passes) a little
confusing to take in, since there's no meaningful type name for some
data structure members.

We can provide an alias that communicates this directly to the reader.

This also has the nice benefit of providing a single point of definition
for node IDs which can allow for easier changes in the future (e.g.
making Node IDs strongly-typed etc).
2021-11-24 17:15:06 -05:00
lioncash 1cf0c335a6 BucketList: Compare against T{} instead of zero directly
Allows BucketList to work with types that aren't a direct numeric
primitive, so long as the object has equality operators defined and are
default constructible
2021-11-24 17:09:20 -05:00
lioncash 52ed4a97cf BucketList: Generify Append() and Erase()
BucketList allows choosing an arbitrary type, but the interface was
assuming uint32_t was only desirable
2021-11-24 17:01:52 -05:00
lioncash db6cbbe7d4 BucketList: Remove dereference to static member
We can reference this directly, since it'll be the same for the lifetime
of the class. Other member functions already do this as well.
2021-11-24 16:57:00 -05:00
lioncash 70a66cdff2 BucketList: Resolve signed/unsigned mismatches in API
The size of the bucket list is expressed as an unsigned value, but all
indexing was taking place with a signed value.
2021-11-24 16:51:12 -05:00
Ryan Houdek ef772c3ef4 Merge pull request #1400 from lioncash/service
CompileService: Store WorkItem instances as unique_ptr
2021-11-24 10:41:05 -08:00
Ryan Houdek 8ab6830fef Merge pull request #1388 from Sonicadvance1/reduce_signal_stalls
Arm64: Reduce the chance of hanging on reentrant allocations
2021-11-24 10:36:07 -08:00
lioncash 5d62a8cedc CompileService: Store WorkItem instances as unique_ptr
Lets us simplify our GC management a little bit and also makes it so we
won't potentially leak memory if an exception occurs anywhere.
2021-11-24 11:48:46 -05:00
Ryan Houdek 262c1c55d2 GDBStub: Fixes a few hangs and crashes
Enough to get a backtrace sometimes but otherwise still pretty finicky.
2021-11-23 17:45:23 -08:00
Ryan Houdek 6913a8b5ec Dispatcher: Initialize Context member states in StoreThreadState
In the case of storing a thread state without a guest signal then this
would be filled with garbage data.
This would then cause a crash on thread restart.

Only happens using gdbstub
2021-11-23 17:42:32 -08:00
Ryan Houdek 5345f7fa37 Merge pull request #1392 from Sonicadvance1/race_condition_compileservice
FEXCore: Fixes CompileService race condition on thread creation
2021-11-23 13:05:34 -08:00
Ryan Houdek 966272ceaf Merge pull request #1398 from lioncash/math
FEXCore: Centralize alignment utility functions in one header
2021-11-23 11:33:48 -08:00
lioncash bb881c5c9b FEXCore: Centralize alignment utility functions
Previously, these alignment functions were in four separate places. We
can centralize these in one predictable spot to remove a little
duplication.
2021-11-23 14:21:48 -05:00
Ryan Houdek 92ae6ae2e4 Merge pull request #1393 from Sonicadvance1/fix_map32bit_error
Linux: Fixes MAP_32BIT error case
2021-11-23 10:42:16 -08:00
Ryan Houdek 7bb789121b Merge pull request #1387 from Sonicadvance1/syscall_fix_sigprocask
Linux: Fixes sigprocmask with new and old set being the same address
2021-11-23 10:39:40 -08:00
Ryan Houdek 97e3a36427 Merge pull request #1386 from Sonicadvance1/ensure_compileservice_signals
JIT: Ensures signals in compileservice JIT space is handled
2021-11-23 10:39:20 -08:00
Ryan Houdek 0560e9be4e Merge pull request #1397 from lioncash/alloc
Syscalls: Move construction of syscall name map into GetSyscallName()
2021-11-23 10:38:46 -08:00
Ryan Houdek 513b02fda3 Merge pull request #1396 from lioncash/fmtstr
General: Migrate logging over to fmt where possible
2021-11-23 10:20:44 -08:00
lioncash 0511764df3 Syscalls: Move construction of syscall name map into GetSyscallName()
These maps are primarily used for debugging, so we don't need to
construct and allocate the data until GetSyscallName() is called.

Removes a bit of memory usage and a static constructor for release
builds.
2021-11-23 13:17:24 -05:00
Lioncash 38b9e85f4f LogManager: Remove now unused portions of the printf-style logger
Now with many of the facilities from the printf logger removed, we can
narrow the exposed functions and in other cases, remove them completely.
2021-11-23 12:51:59 -05:00
Lioncash 75b2f226f6 General: Migrate over to fmt where possible
Migrates lingering instances of the old logger over to fmt where
applicable. This allows removing some of the old defines and functions.

The only remaining usages of the printf-based variant of the logger is
in Tests/LinuxSyscalls/Syscalls.cpp for the strace handling.
2021-11-23 12:51:57 -05:00
Ryan Houdek c4d05fb8e3 Merge pull request #1395 from Sonicadvance1/update_linux_5.16
Updates syscalls to 5.16 syscalls
2021-11-23 09:16:08 -08:00
Ryan Houdek e0d40dd403 Merge pull request #1394 from Sonicadvance1/update_thunksdb
ThunksDB: Updates file to include local libs
2021-11-23 09:16:00 -08:00
Ryan Houdek f97fdd4593 Linux: Fixes MAP_32BIT error case
If MAP_32BIT was failing then we weren't checking error result
correctly.
This was causing us to crash.

Now add a helper function for checking syscall results since it is easy
to mess up.

Replaces the few cases where we were manually checking with the helper.

Fixes a Dota: Underlords crash
2021-11-22 14:15:51 -08:00
Ryan Houdek 6b6b5e880c Linux: Bumps emulated Linux kernel version to 5.16
Only if available as usual.

Fixes #1159
2021-11-22 14:03:28 -08:00
Ryan Houdek 83dc458019 Linux: Adds syscalls for 5.16
Only futex_waitv
2021-11-22 14:03:28 -08:00
Ryan Houdek 4623e4ca21 Linux: Adds syscalls for 5.15
Only process_mrelease
2021-11-22 14:03:28 -08:00
Ryan Houdek bc442871b4 Linux: Adds syscalls for 5.14
memfd_secret and quotactl_fd
2021-11-22 14:03:28 -08:00
Ryan Houdek abaddcccf4 Linux: Adds syscalls for 5.13
Only adds the landlock syscalls
2021-11-22 14:03:28 -08:00
Ryan Houdek ee912c1bb8 Linux: Switches syscall definition usage to FEX's definitions
This works around the problem where some syscall defines may not exist
depending on your compilation environment.
No change in behaviour.
2021-11-22 14:03:28 -08:00
Ryan Houdek f69013fd44 Linux: Updates Arm64 syscall definitions
Fixes the missing definitions from the previous script change
2021-11-22 14:03:28 -08:00
Ryan Houdek 33be4fa98d Scripts: Updates syscall definition extractor
We were missing one of the definition types which means we missed
definitions
2021-11-22 14:03:28 -08:00
Ryan Houdek 99f760dcab ThunksDB: Updates file to include local libs
We were missing installed things in /usr/local, Update the DB to include
those.

Fixes the Ender Lilies vulkan thunk hang that I was encountering
2021-11-22 09:13:15 -08:00
Ryan Houdek 16f1ad4692 Merge pull request #1385 from Sonicadvance1/pressure_vessel_workaround
RootFS: Keep rootfs around while in container
2021-11-22 07:28:07 -08:00
Ryan Houdek e90602f79c Linux: Fixes sigprocmask with new and old set being the same address
If set and oldset point to the same location. We need to be careful to
store the old mask before returning it. Otherwise we will will store the
current mask over the new mask and not change anything.
2021-11-22 06:36:32 -08:00
Ryan Houdek f2d06ecf67 RootFS: Keep rootfs around while in container
For some reason the full container path isn't actually setup right now.
Until this problem is resolved, keep the FEX rootfs configuration
around.

This lets applications still execute albeit still wrapped with the FEX
rootfs rather than the container's
2021-11-22 06:32:56 -08:00
Ryan Houdek 23e583e131 Merge pull request #1390 from Sonicadvance1/packaging_improvements
Various packaging improvements
2021-11-21 18:02:51 -08:00
Ryan Houdek d883b4bf1c Merge pull request #1391 from Sonicadvance1/oops_message
OpcodeDispatcher: Remove debug message
2021-11-21 18:02:37 -08:00
Ryan Houdek f584f16ca4 FEXCore: Fixes CompileService race condition on thread creation
The CompileService was spinning up with the incoming thread mask and
then setting the mask once running.

Instead set the mask, which the thread inherits, then set it back once
it is created.
2021-11-21 11:08:02 -08:00
Ryan Houdek fe3164924c OpcodeDispatcher: Remove debug message 2021-11-21 10:51:41 -08:00
Ryan Houdek 826818192d CPack: Adds more library dependencies
FEXConfig mostly needs these

squashfuse is for FEXMountDaemon
2021-11-21 10:41:52 -08:00
Ryan Houdek 46e30495dc CPack: Add an ldconfig trigger
lintian was complaining about this
2021-11-21 10:38:56 -08:00
Ryan Houdek 702c9e97e4 CPack: Update conflicts and change package name when built statically
Will allow users to choose between a static package or a non-static
package

These packages will conflict with each other, so you can only choose one
or the other.

fex-emu-static: Necessary for Chroots, Can't use thunking.
fex-emu: Necessary for thunking, Can't as easily be used for chroots
2021-11-21 10:36:02 -08:00
Ryan Houdek 2bdd100d2c CPack: Update package contact
lintian was complaining about short description
2021-11-21 10:35:18 -08:00
Ryan Houdek 97ba05ca5b CPack: Update package name
Didn't expose a real version before
2021-11-21 10:34:44 -08:00
Ryan Houdek 508d72d6cc CPack: Add description file
lintian complains about this
2021-11-21 10:32:36 -08:00
Ryan Houdek 3183cf79e5 FEXCore: Adds library soversion
lintian complains about this
2021-11-21 10:31:34 -08:00
Ryan Houdek 5d77a64c8f FEXCore: Compress man page
lintiant complains about this
2021-11-21 10:31:05 -08:00
Ryan Houdek 005c3ce4b3 Tools: Strip binaries on release
lintian complains about this
2021-11-21 10:30:26 -08:00
Ryan Houdek 41adfebf9d JIT: Ensures signals in compileservice JIT space is handled
If a SIGBUS is received in compile service code then we weren't handling
it correctly. Instead we would fail the JIT space check and hand it off
to the guest.
2021-11-20 13:49:54 -08:00
Ryan Houdek e1604fb32f Arm64: Reduce the chance of hanging on reentrant allocations
When compiling code and linking blocks, disable all signals then
reenable once back in the dispatcher.

This adds 15-40 microseconds on to each of these steps and some
additional branching overhead.

This is a problem where our allocator is not reentrant safe and when
receiving a signal we need to jump to a new code region and compile.
Depending on the mutexes that are currently live, we will just stall out
forever.

This is /very/ commonly picked up when attempting to run pressure-vessel
2021-11-20 13:47:30 -08:00
Ryan Houdek 3730a4284c Merge pull request #1382 from neobrain/refactor_thunk_cleanup
Various thunk library cleanups
2021-11-20 09:49:33 -08:00
Ryan Houdek fbd14b65f7 Merge pull request #1383 from Sonicadvance1/fix_error_and_die
FEXCore: Fixes ERROR_AND_DIE
2021-11-19 20:03:00 -08:00
Ryan Houdek 6e1ea92c09 Merge pull request #1384 from Sonicadvance1/support_guest_sigill
FEXCore: Supports guest SIGILL
2021-11-19 20:02:53 -08:00
Ryan Houdek e68d4c52f7 unittests: Updates IR tests for new Break argument
Textual rather than numerical
2021-11-19 12:49:36 -08:00
Ryan Houdek 5759b0d503 FEXCore: Supports guest SIGILL
Fixes #1217

Instead of throwing an error and closing down FEX. Instead pass the
SIGILL to the guest application.

On unhandled instruction implementation the instruction, we instead emit
a _Break IR op at that location.

A _Break IR op will ensure the context state is synchronized at the
point of of the fault and has fairly low overhead. We branch to the
dispatcher which does the SRA spilling.

Tested this with an application that attempts an AVX512 instruction,
catches the fault, and continues onward.
With #1383 in place, we also won't pass spurious ERROR_AND_DIE to the guest anymore.
2021-11-19 12:49:23 -08:00
Ryan Houdek b10ee67525 RCLSE: Don't optimize through a BREAK irop
The end of this optimization pass is getting to a bit long.
It might be worth adding a new flag to IR ops soon for ops we can't
optimize context loadstores through.
2021-11-19 12:19:37 -08:00
Ryan Houdek 47d04ae807 FEXCore: Fixes ERROR_AND_DIE
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.

If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.

Now we can remove this usage of __builtin_trap and switch over to our
own handler.

Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
2021-11-19 12:05:30 -08:00
Ryan Houdek 3dc938b8c5 Merge pull request #1377 from Sonicadvance1/inline_32bit_syscalls
Linux: Passthrough 32-bit syscalls that can be
2021-11-18 20:23:59 -08:00
Ryan Houdek d81d5e21ba Merge pull request #1378 from Sonicadvance1/remove_jit_signalframe
Dispatcher: Removes usage of SignalFrame stack on JITs
2021-11-18 20:23:34 -08:00
Ryan Houdek a92233c7f0 Merge pull request #1379 from Sonicadvance1/update_drm
Updates 32-bit DRM emulation
2021-11-18 19:42:00 -08:00
Ryan Houdek 61ccec3726 Merge pull request #1381 from neobrain/fix_slow_stack
Dispatcher: Use std::vector as the underlying container for std::stack
2021-11-18 13:21:36 -08:00
Tony Wasserka 08bfc5e0c0 Thunks: Catch unresolved function references at build time
This could also prove useful for guest libraries, however the implicit
dependency of libEGL on libGL (via glXGetProcAddress) cannot be resolved
easily with the current build setup.
2021-11-18 19:27:09 +01:00
Tony Wasserka f62ec61e4f Thunks: Add missing includes 2021-11-18 19:27:09 +01:00
Tony Wasserka 06881af363 Thunks/xcb: Remove unused function
The callback mechanism is not used for this API since the function pointer
argument is called on the guest itself (if through some indirection).
2021-11-18 19:27:09 +01:00
Tony Wasserka 598533555d Thunks/vulkan: Small cleanup for pointer lookup logic 2021-11-18 19:27:09 +01:00
Tony Wasserka 85651ad090 Thunks/vulkan: Simplify API entrypoint lookup
This avoids the need for the heap-allocated string->symbol map that had
to be kept around previously.
2021-11-18 19:27:09 +01:00
Tony Wasserka 25545dcd66 Thunks/vulkan: Use more efficient unordered_map initialization
Dynamically adding each entry involves unneeded map rehashing that can be
avoided by constructing the container using an initializer_list.
2021-11-18 19:27:09 +01:00
Tony Wasserka a5ac66c1ce Thunks/vulkan: Use std::string_view to look up API entrypoints
This eliminates unneeded memory allocations on library initialization
without change of behavior.
2021-11-18 19:27:09 +01:00
Tony Wasserka b6effa7a9b Thunks/vulkan: Ignore null-instances when looking up symbols and remove obsolete code 2021-11-18 19:27:08 +01:00
Tony Wasserka 730ba42cf3 Dispatcher: Use std::vector as the underlying container for std::stack
By default std::stack uses the heavyweight std::deque container for storage,
but std::vector is perfectly suitable for our purposes and it has better
performance characteristics.
2021-11-18 19:15:11 +01:00
Ryan Houdek 0165329285 Ioctl: Update i915 drm 2021-11-18 03:27:21 -08:00
Ryan Houdek f12b4aa3d1 Ioctl: Update v3d drm 2021-11-18 03:22:48 -08:00
Ryan Houdek a9852d31e7 Ioctl: Update virtio drm 2021-11-18 03:13:59 -08:00
Ryan Houdek 265e8b4d39 Updates External drm-headers repo 2021-11-18 03:13:42 -08:00
Ryan Houdek ef3338ec0d Dispatcher: Removes usage of SignalFrame stack on JITs
Only the interpreter requires this currently since the stack register
doesn't match.
The JITs correctly set up their stack register and matches what the API
wants.

This is guaranteed to work. signal + rt_sigreturn is a handshake with
the kernel to return the data that the kernel wants.
Only time this doesn't align is when the guest signal handler longjumps
out of the signal handler. Which leaves a dangling entry in the
SignalFrame stack on the interpreter.
2021-11-18 03:10:14 -08:00
Ryan Houdek d47182b631 Merge pull request #1376 from lioncash/vex
Frontend: Handle VEX RXB bits
2021-11-18 01:48:24 -08:00
Ryan Houdek 9ac9b89ab3 Linux: Passthrough 32-bit syscalls that can be
Most of these were already marked for passthrough, just needed to be
enabled.

Some syscalls can be passed through but need to be renamed for 32-bit.
This adds another define which does that for us.
2021-11-18 00:37:28 -08:00
Ryan Houdek 50c8d9edab Merge pull request #1375 from lioncash/bzhi
OpcodeDispatcher: Implement BZHI
2021-11-18 00:31:15 -08:00
lioncash 77c969b424 Frontend: Handle VEX RXB bits
These are equivalent to the REX prefix's RXB bits, except that they're
in 1's complement form.

These are trivial to handle and fix usages of the upper range of
registers for ModRM encoded fields.

To ensure that we have coverage for this, I've altered the BEXTR test to
make use of R14 and R15.
2021-11-17 14:50:21 -05:00
lioncash 95aabc0947 OpcodeDispatcher: Implement BZHI 2021-11-17 13:44:57 -05:00
Ryan Houdek a04cc4dc96 Merge pull request #1373 from lioncash/unused
InterpreterCore/Dispatcher: Resolve unused variable warnings
2021-11-17 07:49:11 -08:00
Ryan Houdek 2a894bc111 Merge pull request #1372 from lioncash/copy
Tests: Minor cleanup
2021-11-17 07:49:02 -08:00
Ryan Houdek f746870356 Merge pull request #1371 from Sonicadvance1/inline_syscall
Arm64: Adds an inline syscall optimization
2021-11-17 07:48:51 -08:00
lioncash 70f8779793 Dispatcher: Mark guest_siginfo as [[maybe_unused]] 2021-11-17 08:43:44 -05:00
lioncash 1615ed9ed7 InterpreterCore: Remove unused variable
This became unused once SIGBUS handling was centralized in one location.
2021-11-17 08:38:48 -05:00
lioncash e3a7c14740 HarnessHelpers: Construct fstream directly in ReadFile()
We can also make use of data() here to avoid undefined behavior if any
inputs are 0 for whatever reason (unlikely, but still nice to have the
reassurance).
2021-11-17 07:42:20 -05:00
lioncash 66a5057914 HarnessHelpers: Mark lookup tables as static
We don't need to keep pushing these onto the stack over and over
(especially given how many tests we have).
2021-11-17 07:42:20 -05:00
lioncash 62dfccc989 HarnessHelpers: Migrate to fmt
May as well, given we're in the area
2021-11-17 07:42:17 -05:00
lioncash 51d8bb9020 ELFCodeLoader2: Migrate to fmt
May as well, given we're in the area.
2021-11-17 07:06:27 -05:00
lioncash 65e28ba6cf IRLoader: Migrate logging to fmt
Given we're in proximity of FEXLoader, we may as well.
2021-11-17 06:58:57 -05:00
lioncash c181033263 FEXLoader: Take string as const in RanAsInterpreter()
This doesn't modify the input string.
2021-11-17 06:55:43 -05:00
lioncash f371b04d7c FEXLoader: Construct fstreams directly
We can use the constructor to directly open the fstream
2021-11-17 06:55:43 -05:00
lioncash 98f42aafa9 FEXLoader: Migrate logging to fmt
While we're in the area, lets move logging over to fmt to get it out of
the way.

This also gives us an opportunity to merge some CMakeLists things
together to organize it.
2021-11-17 06:55:40 -05:00
Lioncash 69437eaf11 FEXLoader: Take sections by const reference in loop
Given LoadedSection instances are 72 bytes, this avoids a minor bit of
copy churn
2021-11-17 06:01:35 -05:00
Ryan Houdek c63c5fb664 Merge pull request #1369 from lioncash/cast
OpcodeDispatcher: Amend return types of BitSize helpers
2021-11-17 02:00:50 -08:00
Ryan Houdek ec8e24e84a IR: Adds documentation description for syscall and inlinesyscall 2021-11-16 22:53:44 -08:00
Ryan Houdek c6b26739db Interpreter: Implements the InlineSyscall for the interpreter
This isn't optimal, just here to work
2021-11-16 22:53:44 -08:00
Ryan Houdek f9e22432d2 Arm64: Adds an inline syscall optimization
In a syscall microbench this improves performance by ~19% on my
Snapdragon 888.
Going from ~9.6 million syscalls per second to ~11.5 million.

Macbook Pro is less effective here due to high syscall overhead due to
VM. Going form  7.2M/s to 7.6M/s, ~6% improvement

We can also inline some 32-bit syscalls but that will need some more
work which isn't done yet. Even though the op in the JIT supports it.
2021-11-16 22:53:44 -08:00
Ryan Houdek d54b9cd272 IREmitter: ReplaceAllWith can't remove sideeffect nodes
Ran in to this when removing syscall nodes. If an IR op has side-effects
then this generic helper can not remove them.
First time this was encountered and it was confusing
2021-11-15 14:37:25 -08:00
Ryan Houdek 92b29138ba Linux: Describes syscalls that can be passed through without change
~0 is used as an invalid syscall indicator to signify that it can be
passed through
2021-11-15 14:36:33 -08:00
Ryan Houdek bab4f53b59 Linux: Updates syscalls enums
Adds the Arm64 one as well so we can map syscall names directly with
renaming here

Adds some comments about which syscalls have no implementation.
Automatically generated.
2021-11-15 14:30:57 -08:00
Ryan Houdek aae1dd4d81 Adds new script to generate syscall enums
Lets us define our own enums for syscall mapping.
Necessary to get all the syscall IDs in a sane way
2021-11-15 14:29:49 -08:00
Ryan Houdek 977d92b7bc Merge pull request #1370 from lioncash/enum-shift
EnumUtils: Remove shift operators from helper macro
2021-11-15 14:16:01 -08:00
lioncash e1cf60c2ab EnumUtils: Remove shift operators from helper macro
These aren't strictly necessary for enum flags and through discussion in
\#1363, would lead to an awkward to use overload.

If these are ever needed, they can be added back at a later date.
2021-11-15 12:03:43 -05:00
Ryan Houdek cabb58ad29 Merge pull request #1368 from Sonicadvance1/32bit_host_runner
unittests: Enables 32-bit host runner
2021-11-15 09:00:17 -08:00
lioncash 44f27c81d3 OpcodeDispatcher: Mark size retrieval functions as nodiscard
Allows the compiler to warn about obvious bug cases.
2021-11-15 11:56:21 -05:00
lioncash c9c8c38d86 OpcodeDispatcher: Increase BitSize helper return vals to uint32_t
Pointed out by neobrain in #1362 that there are some cases where this
will overflow and not fit into 8 bits.
2021-11-15 11:51:09 -05:00
lioncash 5c8c36c5d2 OpcodeDispatcher: Amend cast in GPROffset
There's no bug here, but it definitely looks weird to not be aligned
with the return type.
2021-11-15 11:48:18 -05:00
Ryan Houdek d413cf8d62 unittests: Enables 32-bit host runner
A few tests messing with segments can't be run on the host. This is
because they don't exactly match expected Linux LDT/GDT setup.
2021-11-13 01:37:09 -08:00
Ryan Houdek 18623d07aa TestHarnessRunner: Adds support for 32-bit host runner
Little bit of work to set up the local descriptor table entries for
32-bit execution.
Also setting up the far call state.
2021-11-13 01:37:09 -08:00
Ryan Houdek e63c832a5a unittests: Fixes 32-bit unit test missing flag
This was being run as 64-bit and happening to work
2021-11-13 01:36:09 -08:00
Ryan Houdek 42161b20a4 Updates xbyak external 2021-11-13 01:36:09 -08:00
Ryan Houdek 081d0169d8 Merge pull request #1360 from Sonicadvance1/map_32bit_emulation
Linux: Emulate MAP_32BIT on mmap
2021-11-12 21:49:35 -08:00
Ryan Houdek 77093a7f4d Merge pull request #1358 from Sonicadvance1/signal_rip_adjust
Dispatcher: Partial support for RIP adjust in signal handler
2021-11-12 21:49:25 -08:00
Ryan Houdek bb7f542813 Linux: Emulate MAP_32BIT on mmap
Previous we were just using an address hint to emulate MAP_32BIT.
Seemingly this behaviour has changed on AArch64 where it now isn't
guaranteed to scan up from the hint provided if exact allocation fails.

Now we pull in the full 32-bit allocator and add support for MAP_32BIT
in it. This limits the allocations there in to the first 2GB which Linux
expects.

Necessary for Mono's trampolines to work since it requires code to be in
the first 2GB on x86-64.
2021-11-12 21:25:15 -08:00
Ryan Houdek 2f28548ce4 Dispatcher: Partial support for RIP adjust in signal handler
When the context structure is adjusted in the signal handler then this
updates the state of the CPU on sigreturn.

We check to see if the RIP was adjusted and in this case we will update
the JIT in a potentially unsafe fashion.
This works well enough with signal handlers that effectively just long
jump and not much else.

Fixes wine initial prefix setup where rundl32 expects to do a try-catch
fault capture but infinite loops without this.
2021-11-12 21:22:07 -08:00
Ryan Houdek 498a845c0b Merge pull request #1365 from lioncash/mulx
OpcodeDispatcher: Implement MULX
2021-11-12 21:17:25 -08:00
Ryan Houdek 59ed91c50a Merge pull request #1366 from lioncash/enumreg
OpcodeDecoder: Shorten up GPR and MM base offset retrieval
2021-11-12 20:56:01 -08:00
Ryan Houdek ad864e0e52 Merge pull request #1367 from lioncash/flag
OpcodeDispatcher: Resolve sign mismatches in loops
2021-11-12 17:35:16 -08:00
lioncash 9849aef5bc Frontend: Handle ignoring of widening modes if not 64-bit mode 2021-11-12 17:00:43 -05:00
lioncash 11c3ae19b4 OpcodeDispatcher: Add helper for retrieving MM base offset
Shortens up some lines to make them easier to read.
2021-11-12 16:52:20 -05:00
lioncash 4351ec9f4e OpcodeDispatcher: Add helper for retrieving GPR offsets
Greatly shortens the length of multiple lines. This makes it nicer to
see which register is being loaded or stored.
2021-11-12 16:52:13 -05:00
lioncash 59469705e4 OpcodeDispatcher: Implement MULX 2021-11-12 16:46:38 -05:00
lioncash 763e7c5e08 OpcodeDispatcher: Resolve sign mismatches in loops
Silences compiler warnings
2021-11-12 13:59:46 -05:00
lioncash 2edb5d3004 X86Enums: Convert constants into enums
Allows for parameters and functions to make use of the enum type for
enforcing type checking.
2021-11-12 13:34:55 -05:00
Ryan Houdek c6662a46b5 Merge pull request #1364 from lioncash/fptr
Frontend: Mark modrm function LUT as static
2021-11-12 06:02:08 -08:00
Ryan Houdek d1fa8bcfba Merge pull request #1363 from lioncash/enum
OpcodeDecoder: Convert enums into enum classes where applicable
2021-11-12 05:56:23 -08:00
Ryan Houdek 3a3cbf2b5f Merge pull request #1362 from lioncash/bit
OpcodeDispatcher: Add helper for getting bit sizes
2021-11-12 05:53:05 -08:00
lioncash a0eb2eb2e2 Frontend: Mark modrm function LUT as static
This doesn't need to be constructed on a by-instance basis.
2021-11-12 08:48:49 -05:00
lioncash d58952d1ef OpcodeDispatcher: Make selection flag enum an enum class
Makes the enum and member variable strongly typed, so that it's harder
to implicitly use incorrect values.
2021-11-12 08:41:07 -05:00
lioncash 05ce65322f OpcodeDispatcher: Make X87Tag an enum class
Makes the tag parameter strongly-typed and prevents implicitly passing
it an incorrect value.
2021-11-12 08:31:50 -05:00
lioncash 3e5037042f OpcodeDispatcher: Make Segment enum an enum class
Avoids polluting the surrounding scope.
2021-11-12 08:24:21 -05:00
lioncash 6d82410f44 OpcodeDispatcher: Convert OpType enum into enum class
Prevents pollution of the surrounding scope a little.
2021-11-12 08:19:14 -05:00
lioncash 49966f2954 Utils: Add header for enum-based utilities
Adds a header with a few utilities that make working with strongly typed
enums a little more convenient, especially when working with enum
classes.
2021-11-12 08:15:08 -05:00
Ryan Houdek 51d21d861d Merge pull request #1361 from lioncash/rorx
OpcodeDispatcher: Handle RORX
2021-11-12 05:06:01 -08:00
lioncash 6a79b5cf56 OpcodeDispatcher: Add helper for getting bit sizes
Allows expressing what we're directly getting instead of needing to
repeat it in several places, given how common this is.
2021-11-12 08:01:43 -05:00
lioncash 1d88b499e0 OpcodeDispatcher: Handle RORX 2021-11-12 07:28:28 -05:00
Ryan Houdek b327daf5f5 Merge pull request #1359 from Sonicadvance1/fix_brk_allocate
ELFCodeLoader: Fixes BRK allocation on some ELF loading
2021-11-12 02:40:11 -08:00
Ryan Houdek 90511ddd19 Merge pull request #1357 from Sonicadvance1/fix_sigsuspend_signalmask
Linux: SignalDelegator update signal mask on return from sigsuspend
2021-11-12 02:39:57 -08:00
Ryan Houdek 4be8e23299 ELFCodeLoader: Fixes BRK allocation on some ELF loading
With a DYN ELF in most cases the first program header will be at offset
0.

This isn't always true though, in the case of `micro` the first program
header is at a 4MB offset.
This was causing us to miscalculate the size of ELF in memory, which was
causing BRK to intersect with the ELF on allocation.
This would cause ELF loading to fail and then us to close in some cases.

Also makes sure to pass -1 as FD in mmap with anonymous to match.
2021-11-11 22:13:59 -08:00
Ryan Houdek d4a7afbf5c Linux: SignalDelegator update signal mask on return from sigsuspend
We were failing to update the emulated signal mask on return from
sigsuspend.
This was breaking Unity titles which was using sigsuspend and then
modifying the signal mask with sigprocaddr
2021-11-11 22:06:51 -08:00
Ryan Houdek d3cd354c48 Merge pull request #1356 from lioncash/adx
OpcodeDecoder: Add support for ADX
2021-11-10 10:06:19 -08:00
lioncash 5d707cf831 OpcodeDecoder: Add support for ADX
Another CPU extension we can cross off the list.

Addresses #1355
2021-11-10 12:11:01 -05:00
Ryan Houdek e37a0bbd34 Merge pull request #1354 from lioncash/bmishift
OpcodeDispatcher: Implement BMI2 SARX/SHLX/SHRX
2021-11-07 00:42:41 -07:00
lioncash 9df8245ef2 OpcodeDispatcher: Implement BMI2 SARX/SHLX/SHRX
Gets the basic BMI2 shifts out of the way.
2021-11-07 02:12:06 -05:00
255 changed files with 7998 additions and 4425 deletions

No files matched your search

+20 -5
View File
@@ -17,6 +17,7 @@ option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_STATIC_PIE "Enables static-pie build" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -77,6 +78,11 @@ if (ENABLE_XRAY)
link_libraries(-fxray-instrument)
endif()
if (ENABLE_COMPILE_TIME_TRACE)
add_compile_options(-ftime-trace)
link_libraries(-ftime-trace)
endif()
set (PTHREAD_LIB pthread)
if (ENABLE_LLD)
set (LD_OVERRIDE "-fuse-ld=lld")
@@ -519,18 +525,27 @@ endif()
# Package creation
set (CPACK_GENERATOR "DEB")
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_PACKAGE_CONTACT "team@fex-emu.org")
if (ENABLE_STATIC_PIE)
set (CPACK_PACKAGE_NAME fex-emu-static)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu")
else()
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu-static")
endif()
set (CPACK_PACKAGE_FILE_NAME "${CPACK_PACKAGE_NAME}-${GIT_DESCRIBE_STRING}_${CMAKE_SYSTEM_PROCESSOR}")
set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.org>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libstdc++6")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA "${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm")
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "qemu-user-static")
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "${CPACK_DEBIAN_PACKAGE_CONFLICTS}, qemu-user-static")
endif()
include (CPack)
+3
View File
@@ -0,0 +1,3 @@
x86 and x86-64 Linux emulator
FEX is very much work in progress, so expect things to change.
+1
View File
@@ -0,0 +1 @@
activate-noawait ldconfig
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"StallProcess": "1"
}
}
+62
View File
@@ -10,6 +10,10 @@
"/usr/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/lib/x86_64-linux-gnu/libGL.so",
"/lib/x86_64-linux-gnu/libGL.so.1",
"/lib/x86_64-linux-gnu/libGL.so.1.2.0",
@@ -25,6 +29,9 @@
"/usr/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/lib/x86_64-linux-gnu/libGLESv2.so",
"/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0"
@@ -36,6 +43,9 @@
"/usr/lib/x86_64-linux-gnu/libX11.so",
"/usr/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libX11.so",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/lib/x86_64-linux-gnu/libX11.so",
"/lib/x86_64-linux-gnu/libX11.so.6",
"/lib/x86_64-linux-gnu/libX11.so.6.4.0"
@@ -48,6 +58,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/lib/x86_64-linux-gnu/libvulkan_radeon.so"
],
"Comment": [
@@ -61,6 +72,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/lib/x86_64-linux-gnu/libvulkan_lvp.so"
]
},
@@ -71,6 +83,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/lib/x86_64-linux-gnu/libvulkan_freedreno.so"
]
},
@@ -81,6 +94,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/lib/x86_64-linux-gnu/libvulkan_intel.so"
]
},
@@ -91,6 +105,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/lib/x86_64-linux-gnu/libvulkan_panfrost.so"
]
},
@@ -101,6 +116,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/usr/local/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/lib/x86_64-linux-gnu/libGLX_nvidia.so.0"
],
"Comment": [
@@ -114,6 +130,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/lib/x86_64-linux-gnu/libvulkan_virtio.so"
]
},
@@ -123,6 +140,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb.so",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/lib/x86_64-linux-gnu/libxcb.so",
"/lib/x86_64-linux-gnu/libxcb.so.1",
"/lib/x86_64-linux-gnu/libxcb.so.1.1.0"
@@ -134,6 +154,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
@@ -145,6 +168,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
@@ -156,6 +182,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
@@ -167,6 +196,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
@@ -178,6 +210,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/lib/x86_64-linux-gnu/libxcb-sync.so",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
@@ -189,6 +224,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
@@ -200,6 +238,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-present.so",
"/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0"
@@ -211,6 +252,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
@@ -222,6 +266,9 @@
"/usr/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/lib/x86_64-linux-gnu/libxshmfence.so",
"/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0"
@@ -233,6 +280,9 @@
"/usr/lib/x86_64-linux-gnu/libdrm.so",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/lib/x86_64-linux-gnu/libdrm.so",
"/lib/x86_64-linux-gnu/libdrm.so.2",
"/lib/x86_64-linux-gnu/libdrm.so.2.4.0"
@@ -244,6 +294,9 @@
"/usr/lib/x86_64-linux-gnu/libasound.so",
"/usr/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libasound.so",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/lib/x86_64-linux-gnu/libasound.so",
"/lib/x86_64-linux-gnu/libasound.so.2",
"/lib/x86_64-linux-gnu/libasound.so.2.0.0"
@@ -255,6 +308,9 @@
"/usr/lib/x86_64-linux-gnu/libXrender.so",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/lib/x86_64-linux-gnu/libXrender.so",
"/lib/x86_64-linux-gnu/libXrender.so.1",
"/lib/x86_64-linux-gnu/libXrender.so.1.3.0"
@@ -266,6 +322,9 @@
"/usr/lib/x86_64-linux-gnu/libXext.so",
"/usr/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libXext.so",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/lib/x86_64-linux-gnu/libXext.so",
"/lib/x86_64-linux-gnu/libXext.so.6",
"/lib/x86_64-linux-gnu/libXext.so.6.4.0"
@@ -277,6 +336,9 @@
"/usr/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/lib/x86_64-linux-gnu/libXfixes.so",
"/lib/x86_64-linux-gnu/libXfixes.so.3",
"/lib/x86_64-linux-gnu/libXfixes.so.3.1.0"
+28 -14
View File
@@ -7,7 +7,7 @@ def print_enums(ops, defines):
output_file.write("enum IROps : uint8_t {\n")
for op_key, op_vals in ops.items():
output_file.write("\t\tOP_%s,\n" % op_key.upper())
output_file.write("\tOP_%s,\n" % op_key.upper())
output_file.write("};\n")
@@ -20,7 +20,10 @@ def print_ir_structs(ops, defines):
# Print out defines here
for op_val in defines:
output_file.write("\t%s;\n" % op_val)
if op_val:
output_file.write("\t%s;\n" % op_val)
else:
output_file.write("\n")
output_file.write("// Default structs\n")
output_file.write("struct __attribute__((packed)) IROp_Header {\n")
@@ -81,11 +84,21 @@ def print_ir_structs(ops, defines):
output_file.write("\tstatic constexpr IROps OPCODE = OP_%s;\n" % op_key.upper())
if (SSAArgs > 0):
# Add helpers for accessing SSA arguments, given how frequently they're accessed
output_file.write("\n")
output_file.write("\t[[nodiscard]] OrderedNodeWrapper& Args(size_t Index) {\n")
output_file.write("\t\treturn Header.Args[Index];\n")
output_file.write("\t}\n")
output_file.write("\t[[nodiscard]] const OrderedNodeWrapper& Args(size_t Index) const {\n")
output_file.write("\t\treturn Header.Args[Index];\n")
output_file.write("\t}\n")
output_file.write("};\n")
# Add a static assert that the IR ops must be pod
output_file.write("static_assert(std::is_trivial<IROp_%s>::value);\n\n" % op_key)
output_file.write("static_assert(std::is_standard_layout<IROp_%s>::value);\n\n" % op_key)
output_file.write("static_assert(std::is_trivial_v<IROp_%s>);\n" % op_key)
output_file.write("static_assert(std::is_standard_layout_v<IROp_%s>);\n\n" % op_key)
output_file.write("#undef IROP_STRUCTS\n")
output_file.write("#endif\n\n")
@@ -106,12 +119,12 @@ def print_ir_sizes(ops, defines):
output_file.write("// Make sure our array maps directly to the IROps enum\n")
output_file.write("static_assert(IRSizes[IROps::OP_LAST] == -1ULL);\n\n")
output_file.write("[[maybe_unused]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
output_file.write("[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) std::string_view const& GetName(IROps Op);\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) uint8_t GetArgs(IROps Op);\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("__attribute__((const)) __attribute__((visibility(\"default\"))) bool HasSideEffects(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] std::string_view const& GetName(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
@@ -270,7 +283,8 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write("\t\t\n")
output_file.write("\t\toperator Wrapper<IROp_Header>() const { return Wrapper<IROp_Header> {reinterpret_cast<IROp_Header*>(first), Node}; }\n")
output_file.write("\t\toperator OrderedNode *() { return Node; }\n")
output_file.write("\t\toperator OpNodeWrapper () { return Node->Header.Value; }\n")
output_file.write("\t\toperator const OrderedNode *() const { return Node; }\n")
output_file.write("\t\toperator OpNodeWrapper () const { return Node->Header.Value; }\n")
output_file.write("\t};\n")
output_file.write("\ttemplate <class T>\n")
@@ -301,18 +315,18 @@ def print_ir_allocator_helpers(ops, defines):
output_file.write("\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpSize(OrderedNode *Op) const {\n")
output_file.write("\tuint8_t GetOpSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Size;\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElements(OrderedNode *Op) const {\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\tLOGMAN_THROW_A(HeaderOp->HasDest, \"Op %s has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(HeaderOp->HasDest, \"Op {} has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\treturn HeaderOp->Size / HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(OrderedNode *Op) const {\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->HasDest;\n")
output_file.write("\t}\n\n")
+16 -4
View File
@@ -3,7 +3,6 @@ set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (SRCS
Common/Paths.cpp
Common/JitSymbols.cpp
Common/NetStream.cpp
Common/SoftFloat-3e/extF80_add.c
Common/SoftFloat-3e/extF80_div.c
Common/SoftFloat-3e/extF80_sub.c
@@ -142,7 +141,9 @@ set (SRCS
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
)
@@ -248,6 +249,7 @@ set(OUTPUT_CONFIG_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigValues.inl")
set(OUTPUT_CONFIG_OPTION_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigOptions.inl")
set(INPUT_CONFIG_NAME "${CMAKE_BINARY_DIR}/generated/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
set(OUTPUT_MAN_NAME_COMPRESS "${CMAKE_BINARY_DIR}/generated/FEX.1.gz")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
@@ -263,6 +265,12 @@ add_custom_command(
"${OUTPUT_CONFIG_OPTION_NAME}"
)
add_custom_command(
OUTPUT "${OUTPUT_MAN_NAME_COMPRESS}"
DEPENDS "${OUTPUT_MAN_NAME}"
COMMAND "gzip" "-kf9n" "${OUTPUT_MAN_NAME}"
)
set_source_files_properties(${OUTPUT_CONFIG_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
@@ -270,15 +278,18 @@ set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
set_source_files_properties(${OUTPUT_MAN_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_MAN_NAME_COMPRESS} PROPERTIES
GENERATED TRUE)
# Create the target
add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_CONFIG_NAME}"
DEPENDS "${OUTPUT_CONFIG_OPTION_NAME}"
DEPENDS "${OUTPUT_MAN_NAME}")
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the man page
install(FILES ${OUTPUT_MAN_NAME} DESTINATION ${MAN_DIR}/man1)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -336,6 +347,7 @@ function(AddLibrary Name Type)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
set_target_properties(${Name} PROPERTIES VERSION ${FEXCore_VERSION} SOVERSION ${FEXCore_VERSION})
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
+14 -10
View File
@@ -1,13 +1,16 @@
#pragma once
#include "Common/MathUtils.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include <stdint.h>
#include <stdlib.h>
#include <type_traits>
namespace FEXCore {
template<typename T>
struct BitSet final {
using ElementType = T;
@@ -17,12 +20,12 @@ struct BitSet final {
ElementType *Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
void Free() {
@@ -61,8 +64,8 @@ struct BitSetView final {
ElementType *Memory;
void GetView(BitSet<T> &Set, uint64_t ElementOffset) {
LOGMAN_THROW_A((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
@@ -87,11 +90,12 @@ struct BitSetView final {
bool operator[](T Element) {
return Get(Element);
}
};
static_assert(sizeof(BitSet<uint32_t>) == sizeof(uintptr_t), "Needs to just be a pointer");
static_assert(std::is_trivially_copyable<BitSet<uint32_t>>::value, "Needs to trivially copyable");
static_assert(std::is_trivially_copyable_v<BitSet<uint32_t>>, "Needs to trivially copyable");
static_assert(sizeof(BitSetView<uint32_t>) == sizeof(uintptr_t), "Needs to just be a pointer");
static_assert(std::is_trivially_copyable<BitSetView<uint32_t>>::value, "Needs to trivially copyable");
static_assert(std::is_trivially_copyable_v<BitSetView<uint32_t>>, "Needs to trivially copyable");
} // namespace FEXCore
+17 -30
View File
@@ -1,66 +1,53 @@
#include "Common/JitSymbols.h"
#include <string>
#include <sstream>
#include <unistd.h>
namespace FEXCore {
JITSymbols::JITSymbols() {
std::stringstream PerfMap;
PerfMap << "/tmp/perf-" << getpid() << ".map";
#include <fmt/format.h>
fp = fopen(PerfMap.str().c_str(), "wb");
namespace FEXCore {
JITSymbols::JITSymbols() : fp{nullptr, std::fclose} {
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
fp.reset(fopen(PerfMap.c_str(), "wb"));
if (fp) {
// Disable buffering on this file
setvbuf(fp, nullptr, _IONBF, 0);
setvbuf(fp.get(), nullptr, _IONBF, 0);
}
}
JITSymbols::~JITSymbols() {
if (fp) {
fclose(fp);
}
}
JITSymbols::~JITSymbols() = default;
void JITSymbols::Register(void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << "JIT_0x" << GuestAddr << "_" << HostAddr << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{:x} {:x} JIT_0x{:x}_{:x}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
}
void JITSymbols::Register(void *HostAddr, uint32_t CodeSize, std::string const &Name) {
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << Name << "_" << HostAddr << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{:x} {:x} {}_{:x}\n", HostAddr, CodeSize, Name, HostAddr);
}
void JITSymbols::RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name) {
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << Name << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{:x} {:x} {}\n", HostAddr, CodeSize, Name);
}
void JITSymbols::RegisterJITSpace(void *HostAddr, uint32_t CodeSize) {
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " FEXJIT" << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{:x} {:x} FEXJIT\n", HostAddr, CodeSize);
}
}
} // namespace FEXCore
+11 -6
View File
@@ -1,19 +1,24 @@
#pragma once
#include <cstdint>
#include <cstdio>
#include <string>
#include <memory>
#include <string_view>
namespace FEXCore {
class JITSymbols final {
public:
JITSymbols();
~JITSymbols();
void Register(void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterJITSpace(void *HostAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
private:
FILE* fp{};
using FILEPtr = std::unique_ptr<FILE, decltype(&std::fclose)>;
FILEPtr fp;
};
}
-13
View File
@@ -1,13 +0,0 @@
#pragma once
#include <stdint.h>
static inline uint64_t AlignUp(uint64_t value, uint64_t size) {
return value + (size - value % size) % size;
};
static inline uint64_t AlignDown(uint64_t value, uint64_t size) {
return value - value % size;
};
+1 -1
View File
@@ -72,7 +72,7 @@ namespace FEXCore::Paths {
// Ensure the folder structure is created for our Data
if (!std::filesystem::exists(*EntryCache, ec) &&
!std::filesystem::create_directories(*EntryCache, ec)) {
LogMan::Msg::D("Couldn't create EntryCache directory: '%s'", EntryCache->c_str());
LogMan::Msg::DFmt("Couldn't create EntryCache directory: '{}'", *EntryCache);
}
}
+7 -7
View File
@@ -111,7 +111,7 @@ struct X80SoftFloat {
}
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FSCALE which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
X80SoftFloat Int = FRNDINT(rhs);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
@@ -121,7 +121,7 @@ struct X80SoftFloat {
}
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used F2XM1 which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
BIGFLOAT Src1_d = lhs;
BIGFLOAT Result = exp2l(Src1_d);
Result -= 1.0;
@@ -129,7 +129,7 @@ struct X80SoftFloat {
}
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FYL2X which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = Src2_d * log2l(Src1_d);
@@ -137,7 +137,7 @@ struct X80SoftFloat {
}
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FATAN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = atan2l(Src1_d, Src2_d);
@@ -145,21 +145,21 @@ struct X80SoftFloat {
}
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FTAN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
BIGFLOAT Src_d = lhs;
Src_d = tanl(Src_d);
return Src_d;
}
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FSIN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
BIGFLOAT Src_d = lhs;
Src_d = sinl(Src_d);
return Src_d;
}
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FCOS which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
BIGFLOAT Src_d = lhs;
Src_d = cosl(Src_d);
return Src_d;
+5 -5
View File
@@ -74,14 +74,14 @@ namespace JSON {
json_t const *json = json_createWithPool(&Data.at(0), &Pool.PoolObject);
if (!json) {
LogMan::Msg::E("Couldn't create json");
LogMan::Msg::EFmt("Couldn't create json");
return;
}
json_t const* ConfigList = json_getProperty(json, "Config");
if (!ConfigList) {
LogMan::Msg::E("Couldn't get config list");
LogMan::Msg::EFmt("Couldn't get config list");
return;
}
@@ -92,12 +92,12 @@ namespace JSON {
const char* ConfigString = json_getValue(ConfigItem);
if (!ConfigName) {
LogMan::Msg::E("Couldn't get config name");
LogMan::Msg::EFmt("Couldn't get config name");
return;
}
if (!ConfigString) {
LogMan::Msg::E("Couldn't get ConfigString for '%s'", ConfigName);
LogMan::Msg::EFmt("Couldn't get ConfigString for '{}'", ConfigName);
return;
}
@@ -173,7 +173,7 @@ namespace JSON {
if (!Global &&
!std::filesystem::exists(ConfigFile, ec) &&
!std::filesystem::create_directories(ConfigFile, ec)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigFile.c_str());
LogMan::Msg::DFmt("Couldn't create config directory: '{}'", ConfigFile);
// Let's go local in this case
return "./" + Filename + ".json";
}
@@ -201,6 +201,15 @@
"Disables logging"
]
},
"OutputSocket": {
"Type": "str",
"Default": "",
"Desc": [
"Socket to connect to",
"eg: localhost:8087",
"If set will override the OutputLog location"
]
},
"OutputLog": {
"Type": "str",
"Default": "stderr",
+11 -17
View File
@@ -1437,9 +1437,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS16<true>(
@@ -1499,9 +1498,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS32<true>(
@@ -1561,9 +1559,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS64<true>(
@@ -1744,8 +1741,8 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
const uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A_FMT(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
#endif
DesiredReg = GetRdReg(NextInstr);
}
@@ -1865,8 +1862,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
const uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A_FMT(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
#endif
uint32_t StatusReg = GetRmReg(NextInstr);
uint32_t StoreResultReg = GetRdReg(NextInstr);
@@ -1885,7 +1882,7 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
break;
}
else {
LogMan::Msg::A("Unknown instruction 0x%08x", NextInstr);
LogMan::Msg::AFmt("Unknown instruction 0x{:08x}", NextInstr);
}
}
@@ -1951,9 +1948,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS16<DoRetry>(
@@ -2028,9 +2024,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS32<DoRetry>(
@@ -2105,9 +2100,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS64<DoRetry>(
@@ -1,7 +1,8 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include "aarch64/cpu-aarch64.h"
#include "cpu-features.h"
@@ -32,7 +33,7 @@ Arm64Emitter::Arm64Emitter(size_t size) : vixl::aarch64::Assembler(size, vixl::a
SetCPUFeatures(Features);
if (!SupportsAtomics) {
WARN_ONCE("Host CPU doesn't support atomics. Expect bad performance");
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
@@ -147,26 +148,52 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
void Arm64Emitter::SpillStaticRegs() {
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t SpillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & SpillMask) &&
((1U << Reg2.GetCode()) & SpillMask)) {
stp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & SpillMask)) {
str(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & SpillMask)) {
str(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
}
}
void Arm64Emitter::FillStaticRegs() {
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t FillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & FillMask) &&
((1U << Reg2.GetCode()) & FillMask)) {
ldp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & FillMask)) {
ldr(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & FillMask)) {
ldr(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
}
}
}
@@ -65,8 +65,8 @@ protected:
bool SupportsRCPC{};
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs();
void FillStaticRegs();
void SpillStaticRegs(bool FPRs = true, uint32_t SpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t FillMask = ~0U);
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
@@ -12,15 +12,15 @@ namespace FEXCore::ArchHelpers::Arm64 {
// Obvously such a configuration can't do the actual arm64-specific stuff
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASPAL Not Implemented");
ERROR_AND_DIE_FMT("HandleCASPAL Not Implemented");
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASAL Not Implemented");
ERROR_AND_DIE_FMT("HandleCASAL Not Implemented");
}
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleAtomicMemOp Not Implemented");
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
#endif
+29 -11
View File
@@ -13,6 +13,11 @@
namespace FEXCore::ArchHelpers::Context {
enum ContextFlags : uint32_t {
CONTEXT_FLAG_INJIT = (1U << 0),
CONTEXT_FLAG_32BIT = (1U << 1),
};
struct X86ContextBackup {
// Host State
// RIP and RSP is stored in GPRs here
@@ -22,8 +27,12 @@ struct X86ContextBackup {
// Guest state
int Signal;
uint32_t Flags;
uint64_t OriginalRIP;
uint64_t FPStateLocation;
uint64_t UContextLocation;
uint64_t SigInfoLocation;
FEXCore::Core::CPUState GuestState;
static constexpr int RedZoneSize = 128;
};
@@ -40,6 +49,11 @@ struct ArmContextBackup {
// Guest state
int Signal;
uint32_t Flags;
uint64_t OriginalRIP;
uint64_t FPStateLocation;
uint64_t UContextLocation;
uint64_t SigInfoLocation;
FEXCore::Core::CPUState GuestState;
// Arm64 doesn't have a red zone
@@ -108,7 +122,7 @@ static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
auto MContext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&MContext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
return HostState->FPRs[id];
}
@@ -127,7 +141,7 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
@@ -135,7 +149,8 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -146,7 +161,7 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Backup->FPRs[0], 32 * sizeof(__uint128_t));
HostState->FPCR = Backup->FPCR;
HostState->FPSR = Backup->FPSR;
@@ -160,7 +175,8 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -193,15 +209,15 @@ static inline void SetState(void* ucontext, uint64_t val) {
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not impelented for x86 host");
ERROR_AND_DIE_FMT("Not impelented for x86 host");
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
ERROR_AND_DIE("Not impelented for x86 host");
ERROR_AND_DIE_FMT("Not impelented for x86 host");
}
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not implemented for x86 host");
ERROR_AND_DIE_FMT("Not implemented for x86 host");
}
using ContextBackup = X86ContextBackup;
@@ -220,7 +236,8 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -238,7 +255,8 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -27,7 +27,7 @@ namespace FEXCore {
<< std::endl;
}
Output.close();
LogMan::Msg::D("Dumped %d blocks of sampling data", SamplingMap.size());
LogMan::Msg::DFmt("Dumped {} blocks of sampling data", SamplingMap.size());
}
BlockSamplingData::BlockData *BlockSamplingData::GetBlockData(uint64_t RIP) {
+22 -22
View File
@@ -397,7 +397,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 16) | // Reserved
(0 << 17) | // Reserved
(0 << 18) | // RDSEED
(0 << 19) | // ADCX and ADOX instructions
(1 << 19) | // ADCX and ADOX instructions
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
(0 << 22) | // Reserved
@@ -888,25 +888,25 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
CTX = ctx;
using namespace std::placeholders;
RegisterFunction(0, std::bind(&CPUIDEmu::Function_0h, this, _1));
RegisterFunction(1, std::bind(&CPUIDEmu::Function_01h, this, _1));
RegisterFunction(2, std::bind(&CPUIDEmu::Function_02h, this, _1));
RegisterFunction(0, &CPUIDEmu::Function_0h);
RegisterFunction(1, &CPUIDEmu::Function_01h);
RegisterFunction(2, &CPUIDEmu::Function_02h);
// 3: Serial Number(previously), now reserved
#ifndef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x4, std::bind(&CPUIDEmu::Function_04h, this, _1));
RegisterFunction(0x4, &CPUIDEmu::Function_04h);
#endif
// 5: Monitor/mwait
// Thermal and power management
RegisterFunction(6, std::bind(&CPUIDEmu::Function_06h, this, _1));
RegisterFunction(6, &CPUIDEmu::Function_06h);
// Extended feature flags
RegisterFunction(7, std::bind(&CPUIDEmu::Function_07h, this, _1));
RegisterFunction(7, &CPUIDEmu::Function_07h);
// 9: Direct Cache Access information
// 0x0A: Architectural performance monitoring
// 0x0B: Extended topology enumeration
// 0x0D: Processor extended state enumeration
RegisterFunction(0x0D, std::bind(&CPUIDEmu::Function_0Dh, this, _1));
RegisterFunction(0x0D, &CPUIDEmu::Function_0Dh);
// 0x0F: Intel RDT monitoring
// 0x10: Intel RDT allocation enumeration
// 0x12: Intel SGX capability enumeration
@@ -915,38 +915,38 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
RegisterFunction(0x15, std::bind(&CPUIDEmu::Function_15h, this, _1));
RegisterFunction(0x15, &CPUIDEmu::Function_15h);
#endif
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
// Largest extended function number
RegisterFunction(0x8000'0000, std::bind(&CPUIDEmu::Function_8000_0000h, this, _1));
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
// Processor vendor
RegisterFunction(0x8000'0001, std::bind(&CPUIDEmu::Function_8000_0001h, this, _1));
RegisterFunction(0x8000'0001, &CPUIDEmu::Function_8000_0001h);
// Processor brand string
RegisterFunction(0x8000'0002, std::bind(&CPUIDEmu::Function_8000_0002h, this, _1));
RegisterFunction(0x8000'0002, &CPUIDEmu::Function_8000_0002h);
// Processor brand string continued
RegisterFunction(0x8000'0003, std::bind(&CPUIDEmu::Function_8000_0003h, this, _1));
RegisterFunction(0x8000'0003, &CPUIDEmu::Function_8000_0003h);
// Processor brand string continued
RegisterFunction(0x8000'0004, std::bind(&CPUIDEmu::Function_8000_0004h, this, _1));
RegisterFunction(0x8000'0004, &CPUIDEmu::Function_8000_0004h);
// 0x8000'0005: L1 Cache and TLB identifiers
#ifdef CPUID_AMD
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_8000_0005h, this, _1));
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_8000_0005h);
#else
// This is full reserved on Intel platforms
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_Reserved, this, _1));
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_Reserved);
#endif
// 0x8000'0006: L2 Cache identifiers
RegisterFunction(0x8000'0006, std::bind(&CPUIDEmu::Function_8000_0006h, this, _1));
RegisterFunction(0x8000'0006, &CPUIDEmu::Function_8000_0006h);
// Advanced power management information
RegisterFunction(0x8000'0007, std::bind(&CPUIDEmu::Function_8000_0007h, this, _1));
RegisterFunction(0x8000'0007, &CPUIDEmu::Function_8000_0007h);
// Virtual and physical address sizes
RegisterFunction(0x8000'0008, std::bind(&CPUIDEmu::Function_8000_0008h, this, _1));
RegisterFunction(0x8000'0008, &CPUIDEmu::Function_8000_0008h);
// 0x8000'000A: SVM Revision
// TLB 1GB page identifiers
RegisterFunction(0x8000'0019, std::bind(&CPUIDEmu::Function_8000_0019h, this, _1));
RegisterFunction(0x8000'0019, &CPUIDEmu::Function_8000_0019h);
// 0x8000'001A: Performance optimization identifiers
// 0x8000'001B: Instruction based sampling identifiers
@@ -954,7 +954,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x8000'001D: Cache properties
#ifdef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x8000'001D, std::bind(&CPUIDEmu::Function_8000_001Dh, this, _1));
RegisterFunction(0x8000'001D, &CPUIDEmu::Function_8000_001Dh);
#endif
// 0x8000'001E: Extended APIC ID
// 0x8000'001F: AMD Secure Encryption
+7 -8
View File
@@ -1,13 +1,12 @@
#pragma once
#include <functional>
#include <cstdint>
#include <unordered_map>
#include <utility>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
#include <cstdint>
#include <utility>
namespace FEXCore {
namespace Context {
struct Context;
@@ -27,22 +26,22 @@ public:
void Init(FEXCore::Context::Context *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) {
auto Handler = FunctionHandlers.find(Function);
const auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end()) {
return Function_Reserved(Leaf);
}
return Handler->second(Leaf);
return (this->*Handler->second)(Leaf);
}
private:
FEXCore::Context::Context *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
using FunctionHandler = std::function<FEXCore::CPUID::FunctionResults(uint32_t Leaf)>;
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf);
void RegisterFunction(uint32_t Function, FunctionHandler Handler) {
FunctionHandlers[Function] = Handler;
FunctionHandlers.insert_or_assign(Function, Handler);
}
std::unordered_map<uint32_t, FunctionHandler> FunctionHandlers;
+27 -43
View File
@@ -35,7 +35,9 @@ namespace FEXCore {
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CompileService::Initialize() {
@@ -58,23 +60,15 @@ namespace FEXCore {
// Grab the work queue and clear it
// We don't need to grab the queue mutex since this thread will no longer receive any work events
// Threads are bounded 1:1
while (WorkQueue.size()) {
WorkItem *Item = WorkQueue.front();
while (!WorkQueue.empty()) {
WorkQueue.pop();
delete Item;
}
// Go through the garbage collection array and clear it
// It's safe to clear things that aren't marked safe since we are clearing cache
if (GCArray.size()) {
// Clean up our GC array
for (auto it = GCArray.begin(); it != GCArray.end();) {
delete *it;
it = GCArray.erase(it);
}
}
GCArray.clear();
LOGMAN_THROW_A(CompileThreadData->LocalIRCache.size() == 0, "Compile service must never have LocalIRCache");
LOGMAN_THROW_A_FMT(CompileThreadData->LocalIRCache.empty(), "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
@@ -85,28 +79,26 @@ namespace FEXCore {
SelectedThread->CPUBackend->ClearCache();
}
CompileService::WorkItem *CompileService::CompileCode(uint64_t RIP) {
// Tell the worker thread to compile code for us
WorkItem *Item = new WorkItem{};
Item->RIP = RIP;
WorkItem* ResultItem = nullptr;
{
// Tell the worker thread to compile code for us
auto Item = std::make_unique<WorkItem>();
Item->RIP = RIP;
// Fill the threads work queue
std::scoped_lock<std::mutex> lk(QueueMutex);
WorkQueue.emplace(Item);
std::scoped_lock lk(QueueMutex);
ResultItem = WorkQueue.emplace(std::move(Item)).get();
}
// Notify the thread that it has more work
StartWork.NotifyAll();
return Item;
return ResultItem;
}
void CompileService::ExecutionThread() {
// Ignore signals coming from the guest
CTX->SignalDelegation->MaskThreadSignals();
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
@@ -118,18 +110,18 @@ namespace FEXCore {
if (ShuttingDown.load()) {
break;
}
std::scoped_lock<std::mutex> lk(CompileMutex);
std::scoped_lock lk(CompileMutex);
size_t WorkItems{};
do {
// Grab a work item
WorkItem *Item{};
std::unique_ptr<WorkItem> Item{};
{
std::scoped_lock<std::mutex> lk(QueueMutex);
std::scoped_lock lk(QueueMutex);
WorkItems = WorkQueue.size();
if (WorkItems) {
Item = WorkQueue.front();
if (WorkItems != 0) {
Item = std::move(WorkQueue.front());
WorkQueue.pop();
}
}
@@ -137,7 +129,7 @@ namespace FEXCore {
// If we had a work item then work on it
if (Item) {
// Make sure it's not in lookup cache by accident
LOGMAN_THROW_A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
LOGMAN_THROW_A_FMT(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
@@ -145,11 +137,11 @@ namespace FEXCore {
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
LOGMAN_THROW_A(Generated == true, "Compile Service doesn't have IR Cache");
LOGMAN_THROW_A_FMT(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
ERROR_AND_DIE("Couldn't compile code for thread at RIP: 0x%lx", Item->RIP);
ERROR_AND_DIE_FMT("Couldn't compile code for thread at RIP: 0x{:x}", Item->RIP);
}
Item->CodePtr = CodePtr;
@@ -159,23 +151,15 @@ namespace FEXCore {
Item->StartAddr = StartAddr;
Item->Length = Length;
GCArray.emplace_back(Item);
Item->ServiceWorkDone.NotifyAll();
auto& GCItem = GCArray.emplace_back(std::move(Item));
GCItem->ServiceWorkDone.NotifyAll();
}
} while (WorkItems != 0);
if (GCArray.size()) {
// Clean up our GC array
for (auto it = GCArray.begin(); it != GCArray.end();) {
if ((*it)->SafeToClear) {
delete *it;
it = GCArray.erase(it);
}
else {
++it;
}
}
}
// Clean up any safe entries in our GC array if we have any.
std::erase_if(GCArray, [](const auto& Entry) {
return Entry->SafeToClear.load(std::memory_order_relaxed);
});
}
}
}
+5 -2
View File
@@ -48,6 +48,9 @@ class CompileService final {
// Public for threading
void ExecutionThread();
bool IsAddressInJITCode(uint64_t Address) const {
return CompileThreadData->CPUBackend->IsAddressInJITCode(Address, false, false);
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
@@ -57,8 +60,8 @@ class CompileService final {
std::mutex QueueMutex{};
std::mutex CompileMutex{};
std::queue<WorkItem*> WorkQueue{};
std::vector<WorkItem*> GCArray{};
std::queue<std::unique_ptr<WorkItem>> WorkQueue{};
std::vector<std::unique_ptr<WorkItem>> GCArray{};
Event StartWork{};
std::atomic_bool ShuttingDown{false};
};
+53 -50
View File
@@ -168,7 +168,7 @@ namespace FEXCore::Context {
}
}
LOGMAN_MSG_A("Must never get here");
LOGMAN_MSG_A_FMT("Must never get here");
}
void Context::AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn) {
@@ -192,12 +192,6 @@ namespace FEXCore::Context {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
if (Config.GdbServer) {
StartGdbServer();
}
else {
StopGdbServer();
}
}
Context::~Context() {
@@ -245,12 +239,18 @@ namespace FEXCore::Context {
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
if (Config.GdbServer) {
StartGdbServer();
}
else {
StopGdbServer();
}
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
LocalLoader = Loader;
using namespace FEXCore::Core;
FEXCore::CPU::InitializeInterpreterOpHandlers();
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
@@ -342,7 +342,6 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Return);
Thread->RunningEvents.WaitingToStart.store(true);
}
for (auto &Thread : Threads) {
@@ -381,7 +380,7 @@ namespace FEXCore::Context {
this->Config.MaxInstPerBlock = 1;
Run();
WaitForThreadsToRun();
WaitForIdleWithTimeout();
WaitForIdle();
this->Config.RunningMode = PreviousRunningMode;
this->Config.MaxInstPerBlock = PreviousMaxIntPerBlock;
}
@@ -408,6 +407,13 @@ namespace FEXCore::Context {
if (Thread->RunningEvents.Running.load()) {
StopThread(Thread);
}
// If the thread is waiting to start but immediately killed then there can be a hang
// This occurs in the case of gdb attach with immediate kill
if (Thread->RunningEvents.WaitingToStart.load()) {
Thread->RunningEvents.EarlyExit = true;
Thread->StartRunning.NotifyAll();
}
}
}
@@ -543,14 +549,16 @@ namespace FEXCore::Context {
#elif (_M_ARM_64 && JIT_ARM64)
State->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, State, CompileThread);
#else
ERROR_AND_DIE("FEXCore has been compiled without a viable JIT core");
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
State->CPUBackend = CustomCPUFactory(this, State);
break;
default: ERROR_AND_DIE("Unknown core configuration");
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
}
}
@@ -583,7 +591,7 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
auto It = std::find(Threads.begin(), Threads.end(), Thread);
LOGMAN_THROW_A(It != Threads.end(), "Thread wasn't in Threads");
LOGMAN_THROW_A_FMT(It != Threads.end(), "Thread wasn't in Threads");
Threads.erase(It);
}
@@ -666,9 +674,7 @@ namespace FEXCore::Context {
uint64_t TotalInstructions {0};
uint64_t TotalInstructionsLength {0};
if (!Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP)) {
return {};
}
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP);
auto CodeBlocks = Thread->FrontendDecoder->GetDecodedBlocks();
@@ -681,7 +687,6 @@ namespace FEXCore::Context {
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
@@ -689,11 +694,6 @@ namespace FEXCore::Context {
uint64_t InstsInBlock = Block.NumInstructions;
if (Block.HasInvalidInstruction) {
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
break;
}
for (size_t i = 0; i < InstsInBlock; ++i) {
FEXCore::X86Tables::X86InstInfo const* TableInfo {nullptr};
FEXCore::X86Tables::DecodedInst const* DecodedInfo {nullptr};
@@ -723,7 +723,7 @@ namespace FEXCore::Context {
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
}
if (TableInfo->OpcodeDispatcher) {
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
Thread->OpDispatcher->HandledLock = false;
Thread->OpDispatcher->ResetDecodeFailure();
@@ -734,7 +734,7 @@ namespace FEXCore::Context {
else {
if (Thread->OpDispatcher->HandledLock != IsLocked) {
HadDispatchError = true;
LogMan::Msg::E("Missing LOCK HANDLER at 0x%lx{'%s'}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
}
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
@@ -742,8 +742,9 @@ namespace FEXCore::Context {
}
}
else {
LogMan::Msg::E("Missing OpDispatcher at 0x%lx{'%s'}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
HadDispatchError = true;
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
}
// If we had a dispatch error then leave early
@@ -770,21 +771,20 @@ namespace FEXCore::Context {
Thread->OpDispatcher->Finalize();
auto IRDumper = [Thread, GuestRIP](IR::RegisterAllocationData* RA) {
const auto IRDumper = [Thread, GuestRIP](IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
if (Thread->CTX->Config.DumpIR() =="stderr") {
if (DumpIRStr =="stderr") {
f = stderr;
}
else if (Thread->CTX->Config.DumpIR() =="stdout") {
else if (DumpIRStr =="stdout") {
f = stdout;
}
else {
std::stringstream fileName;
fileName << Thread->CTX->Config.DumpIR() << "/" << std::hex << GuestRIP << (RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.str().c_str(), "w");
const auto fileName = fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.c_str(), "w");
CloseAfter = true;
}
@@ -792,7 +792,7 @@ namespace FEXCore::Context {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fprintf(f,"IR-%s 0x%lx:\n%s\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str().c_str());
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
fclose(f);
@@ -814,15 +814,15 @@ namespace FEXCore::Context {
out.seekg(0);
auto reparsed = IR::Parse(&out);
if (reparsed == nullptr) {
LOGMAN_MSG_A("Failed to parse ir\n");
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
std::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::I("one:\n %s", out.str().c_str());
LogMan::Msg::I("two:\n %s", out2.str().c_str());
LOGMAN_MSG_A("Parsed ir doesn't match\n");
LogMan::Msg::IFmt("one:\n {}", out.str());
LogMan::Msg::IFmt("two:\n {}", out2.str());
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
}
}
}
@@ -837,7 +837,7 @@ namespace FEXCore::Context {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::I("IR 0x%lx:\n%s\n@@@@@\n", GuestRIP, out.str().c_str());
LogMan::Msg::IFmt("IR 0x{:x}:\n{}\n@@@@@\n", GuestRIP, out.str());
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
@@ -968,7 +968,7 @@ namespace FEXCore::Context {
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
if (hash == AOTEntry->GuestHash) {
IRList = AOTEntry->GetIRData();
//LogMan::Msg::D("using %s + %lx -> %lx\n", file->second.fileid.c_str(), AOTEntry->first, GuestRIP);
//LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
RAData = AOTEntry->GetRAData();;
@@ -978,10 +978,10 @@ namespace FEXCore::Context {
GeneratedIR = true;
} else {
LogMan::Msg::I("AOTIR: hash check failed %lx\n", MappedStart);
LogMan::Msg::IFmt("AOTIR: hash check failed {:x}\n", MappedStart);
}
} else {
//LogMan::Msg::I("AOTIR: Failed to find %lx, %lx, %s\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid.c_str());
//LogMan::Msg::IFmt("AOTIR: Failed to find {:x}, {:x}, {}\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid);
}
}
}
@@ -1076,7 +1076,7 @@ namespace FEXCore::Context {
AOTIRCache.insert({Module, {Array, FilePtr, Size}});
LogMan::Msg::D("AOTIR: Module %s has %ld functions", Module.c_str(), Array->Count);
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
@@ -1136,7 +1136,7 @@ namespace FEXCore::Context {
auto NewBlock = CompileBlock(Frame, GuestRIP);
if (NewBlock == 0) {
LogMan::Msg::E("CompileBlockJit: Failed to compile code %lX - aborting process", GuestRIP);
LogMan::Msg::EFmt("CompileBlockJit: Failed to compile code {:X} - aborting process", GuestRIP);
// Return similar behaviour of SIGILL abort
Frame->Thread->StatusCode = 128 + SIGILL;
Stop(false /* Ignore current thread */);
@@ -1180,7 +1180,7 @@ namespace FEXCore::Context {
Thread->CompileService->Initialize();
}
auto WorkItem = Thread->CompileService->CompileCode(GuestRIP);
auto* WorkItem = Thread->CompileService->CompileCode(GuestRIP);
WorkItem->ServiceWorkDone.Wait();
// Return here with the data in place
CodePtr = WorkItem->CodePtr;
@@ -1304,14 +1304,17 @@ namespace FEXCore::Context {
Thread->StartRunning.Wait();
}
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
if (!Thread->RunningEvents.EarlyExit.load()) {
Thread->RunningEvents.WaitingToStart = false;
Thread->RunningEvents.Running = true;
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = true;
Thread->RunningEvents.WaitingToStart = false;
Thread->RunningEvents.Running = false;
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
// If it is the parent thread that died then just leave
// XXX: This doesn't make sense when the parent thread doesn't outlive its children
@@ -25,6 +25,7 @@
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#include <sys/syscall.h>
#include <unistd.h>
namespace FEXCore::CPU {
@@ -205,11 +206,31 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ret();
}
constexpr bool SignalSafeCompile = true;
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x0, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
}
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
@@ -217,6 +238,24 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ldr(x3, &l_ExitFunctionLink);
blr(x3);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
mov(x4, x0);
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
mov(x0, x4);
}
if (SRAEnabled)
FillStaticRegs();
br(x0);
@@ -226,16 +265,53 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
bind(&NoBlock);
if (SRAEnabled)
SpillStaticRegs();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x2, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Reload x2 to bring back RIP
ldr(x2, MemOperand(sp, 8, Offset));
}
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileBlock);
if (SRAEnabled)
SpillStaticRegs();
// X2 contains our guest RIP
blr(x3); // { CTX, Frame, RIP}
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
}
if (SRAEnabled)
FillStaticRegs();
@@ -250,6 +326,17 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
hlt(0);
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
hlt(0);
}
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
@@ -349,8 +436,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext) {
void Arm64Dispatcher::SpillSRA(void *ucontext, uint32_t IgnoreMask) {
for(int i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
// Skip this one, it's already spilled
continue;
}
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
@@ -18,7 +18,7 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
protected:
void SpillSRA(void *ucontext) override;
void SpillSRA(void *ucontext, uint32_t IgnoreMask) override;
};
}
@@ -1,8 +1,8 @@
#include "Common/MathUtils.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/CompileService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
@@ -12,12 +12,13 @@
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <atomic>
#include <condition_variable>
#include <bits/types/siginfo_t.h>
#include <signal.h>
#include <string.h>
#include <csignal>
#include <cstring>
namespace FEXCore::CPU {
@@ -27,15 +28,19 @@ void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuS
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
Thread->RunningEvents.ThreadSleeping = true;
// Go to sleep
Thread->StartRunning.Wait();
Thread->RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
Thread->RunningEvents.ThreadSleeping = false;
ctx->IdleWaitCV.notify_all();
}
void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -65,13 +70,31 @@ void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
SignalFrames.push(NewSP);
// Signal frames are only used on the interpreter
// The JITS require the stack to be setup correctly on rt_sigreturn
if (CTX->Config.Core() == FEXCore::Config::CONFIG_INTERPRETER) {
SignalFrames.push(NewSP);
}
Context->Flags = 0;
Context->FPStateLocation = 0;
Context->UContextLocation = 0;
Context->SigInfoLocation = 0;
return Context;
}
void Dispatcher::RestoreThreadState(void *ucontext) {
LOGMAN_THROW_A(!SignalFrames.empty(), "Trying to restore a signal frame when we don't have any");
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uint64_t OldSP{};
if (CTX->Config.Core() == FEXCore::Config::CONFIG_IRJIT) {
OldSP = ArchHelpers::Context::GetSp(ucontext);
}
else {
LOGMAN_THROW_A_FMT(!SignalFrames.empty(), "Trying to restore a signal frame when we don't have any");
OldSP = SignalFrames.top();
SignalFrames.pop();
}
uintptr_t NewSP = OldSP;
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
@@ -80,6 +103,77 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
if (Context->UContextLocation) {
auto Frame = ThreadState->CurrentFrame;
if (Context->Flags &ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT) {
// XXX: Unsupported since it needs state reconstruction
// If we are in the JIT then SRA might need to be restored to values from the context
// We can't currently support this since it might result in tearing without real state reconstruction
}
if (!(Context->Flags & ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_32BIT)) {
auto *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(Context->UContextLocation);
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP]) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP];
// XXX: Full context setting
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x];
COPY_REG(R8);
COPY_REG(R9);
COPY_REG(R10);
COPY_REG(R11);
COPY_REG(R12);
COPY_REG(R13);
COPY_REG(R14);
COPY_REG(R15);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
}
}
else {
auto *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(Context->UContextLocation);
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP]) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
// XXX: Full context setting
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_##x];
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
}
}
}
}
static uint32_t ConvertSignalToTrapNo(int Signal, siginfo_t *HostSigInfo) {
@@ -115,7 +209,8 @@ static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
StoreThreadState(Signal, ucontext);
auto ContextBackup = StoreThreadState(Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
// Ref count our faults
@@ -140,15 +235,39 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// Otherwise we might load garbage
if (SRAEnabled) {
if (IsAddressInJITCode(OldPC, false)) {
uint32_t IgnoreMask{};
#ifdef _M_ARM_64
if (Frame->InSyscallInfo != 0) {
// We are in a syscall, this means we are in a weird register state
// We need to spill SRA but only some of it, since some values have already been spilled
// Lower 16 bits tells us which registers are already spilled to the context
// So we ignore spilling those ones
uint16_t NumRegisters = std::popcount(Frame->InSyscallInfo & 0xFFFF);
if (NumRegisters >= 4) {
// Unhandled case
IgnoreMask = 0;
}
else {
IgnoreMask = Frame->InSyscallInfo & 0xFFFF;
}
}
else {
// We must spill everything
IgnoreMask = 0;
}
#endif
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
SpillSRA(ucontext, IgnoreMask);
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT;
} else {
if (!IsAddressInJITCode(OldPC, true)) {
// This is likely to cause issues but in some cases it isn't fatal
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
// Only throw a log message in this case
LogMan::Msg::E("Signals in dispatcher have unsynchronized context");
LogMan::Msg::EFmt("Signals in dispatcher have unsynchronized context");
}
}
}
@@ -180,8 +299,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// siginfo_t
siginfo_t *HostSigInfo = reinterpret_cast<siginfo_t*>(info);
if (GuestAction->sa_flags & SA_SIGINFO) {
// Backup where we think the RIP currently is
ContextBackup->OriginalRIP = Frame->State.rip;
if (GuestAction->sa_flags & SA_SIGINFO) {
// Setup ucontext a bit
if (Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::_libc_fpstate);
@@ -196,6 +317,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
NewGuestSP = AlignDown(NewGuestSP, alignof(siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
ContextBackup->FPStateLocation = FPStateLocation;
ContextBackup->UContextLocation = UContextLocation;
ContextBackup->SigInfoLocation = SigInfoLocation;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
siginfo_t *guest_siginfo = reinterpret_cast<siginfo_t*>(SigInfoLocation);
@@ -264,6 +389,8 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_32BIT;
NewGuestSP -= sizeof(FEXCore::x86::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::_libc_fpstate));
uint64_t FPStateLocation = NewGuestSP;
@@ -276,6 +403,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
ContextBackup->FPStateLocation = FPStateLocation;
ContextBackup->UContextLocation = UContextLocation;
ContextBackup->SigInfoLocation = SigInfoLocation;
FEXCore::x86::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(UContextLocation);
FEXCore::x86::siginfo_t *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(SigInfoLocation);
@@ -368,7 +499,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_siginfo->_sifields._timer.sigval.sival_int = HostSigInfo->si_int;
break;
default:
LogMan::Msg::E("Unhandled siginfo_t for signal: %d\n", Signal);
LogMan::Msg::EFmt("Unhandled siginfo_t for signal: {}\n", Signal);
break;
}
@@ -402,7 +533,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SignalReturn;
LOGMAN_THROW_A(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
LOGMAN_THROW_A_FMT(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[FEXCore::X86State::REG_RSP] = NewGuestSP;
}
@@ -454,7 +585,8 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
}
@@ -486,11 +618,21 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
}
// We need to be a little bit careful here
// If we were already paused (due to GDB) and we are immediately stopping (due to gdb kill)
// Then we need to ensure we don't double decrement our idle thread counter
if (ThreadState->RunningEvents.ThreadSleeping) {
// If the thread was sleeping then its idle counter was decremented
// Reincrement it here to not break logic
++ThreadState->CTX->IdleWaitRefCount;
}
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -531,15 +673,19 @@ void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) const {
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher, bool IncludeCompileService) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher) {
return IsAddressInDispatcher(Address);
if (IncludeDispatcher && IsAddressInDispatcher(Address)) {
return true;
}
if (IncludeCompileService && ThreadState->CompileService && ThreadState->CompileService->IsAddressInJITCode(Address)) {
return true;
}
return false;
}
@@ -3,6 +3,7 @@
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <bits/types/stack_t.h>
#include <cstdint>
@@ -47,6 +48,8 @@ public:
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
@@ -67,7 +70,7 @@ public:
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
@@ -77,12 +80,12 @@ protected:
: CTX {ctx}
, ThreadState {Thread} {}
void StoreThreadState(int Signal, void *ucontext);
ArchHelpers::Context::ContextBackup* StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext) {}
virtual void SpillSRA(void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
@@ -141,7 +141,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
je(NoBlock);
// Update L1
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
@@ -274,10 +274,15 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
{
// Signal return handler
SignalHandlerReturnAddress = getCurr<uint64_t>();
ud2();
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = getCurr<uint64_t>();
ud2();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
+59 -43
View File
@@ -192,7 +192,7 @@ Decoder::~Decoder() {
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_A(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
@@ -209,7 +209,7 @@ uint64_t Decoder::ReadData(uint8_t Size) {
}
if (Size > sizeof(uint64_t)) {
LOGMAN_MSG_A("Unknown data size to read");
LOGMAN_MSG_A_FMT("Unknown data size to read");
return 0;
}
@@ -351,10 +351,9 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
Operand->Data.SIB.Index = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X ? 1 : 0, SIB.index, false, false, false, false, 0b100);
Operand->Data.SIB.Base = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
uint64_t Literal {0};
LOGMAN_THROW_A(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
LOGMAN_THROW_A_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
Literal = ReadData(Displacement);
uint64_t Literal = ReadData(Displacement);
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -364,8 +363,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
// Explained in Table 1-14. "Operand Addressing Using ModRM and SIB Bytes"
if (ModRM.rm == 0b101) {
// 32bit Displacement
uint32_t Literal;
Literal = ReadData(4);
const uint32_t Literal = ReadData(4);
Operand->Type = DecodedOperand::OpType::RIPRelative;
Operand->Data.RIPLiteral.Value.u = Literal;
@@ -378,8 +376,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
else {
uint8_t DisplacementSize = ModRM.mod == 1 ? 1 : 4;
uint32_t Literal{};
Literal = ReadData(DisplacementSize);
uint32_t Literal = ReadData(DisplacementSize);
if (DisplacementSize == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -397,26 +394,26 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
return false;
}
LOGMAN_THROW_A(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
uint8_t DestSize{};
const bool HasWideningDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST) != 0 ||
Options.w;
(Options.w && CTX->Config.Is64BitMode);
const bool HasNarrowingDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
bool HasXMMSrc = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
@@ -535,7 +532,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LOGMAN_THROW_A(!HasMODRM, "This instruction shouldn't have ModRM!");
LOGMAN_THROW_A_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
// This also means that the destination is always a GPR on these ones
@@ -641,7 +638,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (Bytes != 0) {
LOGMAN_THROW_A(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
@@ -666,7 +663,8 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LOGMAN_THROW_A(Bytes == 0, "Inst at 0x%lx: 0x%04x '%s' Had an instruction of size %d with %d remaining", DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining",
DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
}
@@ -677,21 +675,22 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
return false;
}
LOGMAN_THROW_A(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX,
"REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
@@ -747,7 +746,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
3,
};
uint8_t Field = RegToField[ModRM.reg];
LOGMAN_THROW_A(Field != 255, "Invalid field selected!");
LOGMAN_THROW_A_FMT(Field != 255, "Invalid field selected!");
LocalOp = (Field << 3) | ModRM.rm;
return NormalOp(&SecondModRMTableOps[LocalOp], LocalOp);
@@ -775,6 +774,11 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
const uint8_t Byte1 = ReadByte();
DecodedHeader options{};
if ((Byte1 & 0b10000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.R shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
}
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
@@ -785,8 +789,15 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
options.w = (Byte2 & 0b10000000) != 0;
if ((Byte1 & 0b01000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.X shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::E("We don't understand a map_select of: %d", map_select);
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
return false;
}
}
@@ -871,12 +882,19 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->LastEscapePrefix == 0xF2) // REPNE
if (DecodeInst->LastEscapePrefix == 0xF2) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F2;
else if (DecodeInst->LastEscapePrefix == 0x66) // Operand Size
} else if (DecodeInst->LastEscapePrefix == 0xF3) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F3;
} else if (DecodeInst->LastEscapePrefix == 0x66) {
// Operand size
Prefix = PF_38_66;
}
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
@@ -991,7 +1009,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
LOGMAN_THROW_A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
@@ -1044,13 +1062,13 @@ void Decoder::BranchTargetInMultiblockRange() {
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
break;
}
case 0xE9:
case 0xEB: // Both are unconditional JMP instructions
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
Conditional = false;
break;
@@ -1116,7 +1134,7 @@ const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
@@ -1130,7 +1148,6 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
EntryPoint = PC;
InstStream = _InstStream;
bool ErrorDuringDecoding = false;
uint64_t TotalInstructions{};
// If we don't have symbols available then we become a bit optimistic about multiblock ranges
@@ -1162,20 +1179,15 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
while (1) {
ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
LogMan::Msg::D("Couldn't Decode something at 0x%lx, Started at 0x%lx", PC + PCOffset, PC);
if (Blocks.size() == 1) {
return false;
}
LOGMAN_THROW_A(Blocks.size() != 1, "Decode Error in entry block");
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", PC + PCOffset, PC);
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
if (ErrorDuringDecoding && Blocks.size() != 1) {
ErrorDuringDecoding = false;
}
break;
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
}
DecodedMinAddress = std::min(DecodedMinAddress, RIPToDecode + PCOffset);
@@ -1184,6 +1196,11 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
++BlockNumberOfInstructions;
++DecodedSize;
// Can not continue this block at all on invalid instruction
if (CurrentBlockDecoding.HasInvalidInstruction) {
break;
}
bool CanContinue = false;
if (!(DecodeInst->TableInfo->Flags &
(FEXCore::X86Tables::InstFlags::FLAGS_BLOCK_END | FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP))) {
@@ -1228,7 +1245,6 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
std::sort(Blocks.begin(), Blocks.end(), [](const FEXCore::Frontend::Decoder::DecodedBlocks& a, const FEXCore::Frontend::Decoder::DecodedBlocks& b) {
return a.Entry < b.Entry;
});
return !ErrorDuringDecoding;
}
}
+2 -2
View File
@@ -27,7 +27,7 @@ public:
Decoder(FEXCore::Context::Context *ctx);
~Decoder();
bool DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
@@ -91,7 +91,7 @@ private:
void DecodeModRM_16(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
const std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp {
static constexpr std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp{
&FEXCore::Frontend::Decoder::DecodeModRM_64,
&FEXCore::Frontend::Decoder::DecodeModRM_16,
};
+238 -81
View File
@@ -12,7 +12,7 @@ $end_info$
#include <string>
#include <memory>
#include <optional>
#include "Common/NetStream.h"
#include <vector>
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
@@ -24,6 +24,7 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/NetStream.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
@@ -473,6 +474,48 @@ std::string buildTargetXML() {
return xml.str();
}
std::string buildMemoryMap() {
std::ostringstream xml;
xml << "<?xml version='1.0'?>\n";
xml << "<!DOCTYPE memory-map>\n";
xml << "<memory-map>";
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
while (std::getline(fs, Line)) {
if (fs.eof()) break;
uint64_t Begin, End;
char r,w,x,p;
if (sscanf(Line.c_str(), "%lx-%lx %c%c%c%c", &Begin, &End, &r, &w, &x, &p) == 6) {
xml << "<memory type=\"ram\" start=\"0x" << std::hex << Begin << "\" length=\"0x" << (End - Begin) << "\"/>\n";
}
}
xml << "</memory-map>";
xml << std::flush;
return xml.str();
}
std::string buildOSData() {
std::ostringstream xml;
xml << "<?xml version='1.0'?>\n";
xml << "<!DOCTYPE target SYSTEM \"osdata.dtd\">\n";
xml << "<osdata type=\"processes\">";
// XXX
xml << "</osdata>";
xml << std::flush;
return xml.str();
}
GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
std::string object;
std::string rw;
@@ -539,11 +582,11 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
ThreadString.clear();
std::ostringstream ss;
ss << "<?xml version=\"1.0\?>\n";
ss << "<?xml version=\"1.0\"?>\n";
ss << "<threads>\n";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << "\t<thread id=\"" << std::hex << Thread->ThreadManager.GetTID() << "\" core=\"" << i << "\" name=\"" << getThreadName(Thread->ThreadManager.GetTID()) << "\">\n";
for (auto &Thread : *Threads) {
// Thread id is in hex without 0x prefix
ss << "\t<thread id=\"" << std::hex << Thread->ThreadManager.GetTID() << "\" name=\"" << getThreadName(Thread->ThreadManager.GetTID()) << "\">\n";
ss << "\t</thread>\n";
}
@@ -554,6 +597,20 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
return {encode(ThreadString.substr(offset, length)), HandledPacketType::TYPE_ACK};
}
if (object == "memory-map") {
if (offset == 0) {
MemoryMapString = buildMemoryMap();
}
return {encode(MemoryMapString.substr(offset, length)), HandledPacketType::TYPE_ACK};
}
if (object == "osdata") {
if (offset == 0) {
OSDataString = buildOSData();
}
return {encode(OSDataString.substr(offset, length)), HandledPacketType::TYPE_ACK};
}
return {"", HandledPacketType::TYPE_UNKNOWN};
}
@@ -643,12 +700,82 @@ GdbServer::HandledPacketType GdbServer::handleMemory(const std::string &packet)
GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
const auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
const auto MatchStr = [](const std::string &Str, const char *str) -> bool { return Str.rfind(str, 0) == 0; };
if (match("qSupported")) {
return {"PacketSize=5000;xmlRegisters=i386;qXfer:exec-file:read+;qXfer:features:read+;", HandledPacketType::TYPE_ACK};
const auto split = [](const std::string &Str, char deliminator) -> std::vector<std::string> {
std::vector<std::string> Elements;
std::istringstream Input(Str);
for (std::string line;
std::getline(Input, line);
Elements.emplace_back(line));
return Elements;
};
if (match("QNonStop:")) {
auto ss = std::istringstream(packet);
ss.seekg(std::string("QNonStop:").size());
ss.get(); // discard colon
ss >> NonStopMode;
return {"OK", HandledPacketType::TYPE_ACK};
}
if (match("qSupported:")) {
// eg: qSupported:multiprocess+;swbreak+;hwbreak+;qRelocInsn+;fork-events+;vfork-events+;exec-events+;vContSupported+;QThreadEvents+;no-resumed+;memory-tagging+;xmlRegisters=i386
auto Features = split(packet.substr(strlen("qSupported:")), ';');
// For feature documentation
// https://sourceware.org/gdb/current/onlinedocs/gdb/General-Query-Packets.html#qSupported
std::string SupportedFeatures{};
// Required features
SupportedFeatures += "PacketSize=5000;";
SupportedFeatures += "xmlRegisters=i386;";
// XXX: Not yet supported, would be easy
// SupportedFeatures += "qXfer:auxv-file:read+";
SupportedFeatures += "qXfer:exec-file:read+;";
SupportedFeatures += "qXfer:features:read+;";
// XXX: Requires parsing the ELF and watching the library list
// SupportedFeatures += "qXfer:libraries:read+;";
SupportedFeatures += "qXfer:memory-map:read+;";
SupportedFeatures += "qXfer:siginfo:read+;";
SupportedFeatures += "qXfer:siginfo:write+;";
// XXX: Allowing this causes GDB to crash
SupportedFeatures += "qXfer:threads:read+;";
// QCatchSignals
// QPassSignals
SupportedFeatures += "QNonStop+;";
SupportedFeatures += "qXfer:osdata:read+;";
// Causes GDB to crash?
// SupportedFeatures += "QStartNoAckMode+;";
for (auto &Feature : Features) {
if (MatchStr(Feature, "swbreak+")) {
SupportedFeatures += "swbreak+;";
}
if (MatchStr(Feature, "hwbreak+")) {
SupportedFeatures += "hwbreak+;";
}
if (MatchStr(Feature, "vContSupported+")) {
SupportedFeatures += "vContSupported+;";
}
// Unsupported:
// multiprocess
// qRelocInsn
// fork-events
// vfork-events
// exec-events
// QThreadEvents
// no-resumed
// memory-tagging
}
return {SupportedFeatures, HandledPacketType::TYPE_ACK};
}
if (match("qAttached")) {
return {"1", HandledPacketType::TYPE_ACK}; // We don't currently support launching executables from gdb.
return {"tnotrun:0", HandledPacketType::TYPE_ACK}; // We don't currently support launching executables from gdb.
}
if (match("qXfer")) {
return handleXfer(packet);
@@ -667,7 +794,10 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
ss << "m";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
ss << std::hex << Thread->ThreadManager.TID << ",";
ss << std::hex << Thread->ThreadManager.TID;
if (i != (Threads->size() - 1)) {
ss << ",";
}
}
return {ss.str(), HandledPacketType::TYPE_ACK};
}
@@ -700,6 +830,30 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
return {"", HandledPacketType::TYPE_UNKNOWN};
}
GdbServer::HandledPacketType GdbServer::ThreadAction(char action, uint32_t tid) {
switch (action) {
case 'c': {
CTX->Run();
CTX->WaitForThreadsToRun();
return {"", HandledPacketType::TYPE_ONLYACK};
}
case 's': {
CTX->Step();
SendPacketPair({"OK", HandledPacketType::TYPE_ACK});
auto str = fmt::format("T05thread:{:02x};core:2c;", getpid());
SendPacketPair({std::move(str), HandledPacketType::TYPE_ACK});
return {"OK", HandledPacketType::TYPE_ACK};
}
case 't':
// This thread isn't part of the thread pool
CTX->Stop(false /* Ignore current thread */);
return {"OK", HandledPacketType::TYPE_ACK};
default:
return {"E00", HandledPacketType::TYPE_ACK};
}
}
GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
const auto match = [&](const std::string& str) -> std::optional<std::istringstream> {
if (packet.rfind(str, 0) == 0) {
@@ -762,44 +916,25 @@ GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
return {F_data(ret, data), HandledPacketType::TYPE_ACK};
}
if ((ss = match("vCont?"))) {
return {"vCont;c;C;t;s;S;r", HandledPacketType::TYPE_ACK}; // We support continue, step and terminate
// FIXME: We also claim to support continue with signal... because it's compulsory
return {"vCont;c;t;s;r", HandledPacketType::TYPE_ACK}; // We support continue, step and terminate
// FIXME: We also claim to support continue with signal... because it's compulsory
}
if ((ss = match("vCont;"))) {
char action;
int thread;
char action{};
int thread{};
action = ss->get();
action = ss->get();
if (ss->peek() == ':') {
ss->get();
*ss >> std::hex >> thread;
}
if (ss->peek() == ':') {
ss->get();
*ss >> std::hex >> thread;
}
if (ss->fail()) {
return {"E00", HandledPacketType::TYPE_ACK};
}
switch (action) {
case 'c': {
CTX->Run();
return {"", HandledPacketType::TYPE_ONLYACK};
}
case 's': {
CTX->Step();
SendPacketPair({"OK", HandledPacketType::TYPE_ACK});
auto str = fmt::format("T05thread:{:02x};core:2c;", getpid());
SendPacketPair({std::move(str), HandledPacketType::TYPE_ACK});
return {"OK", HandledPacketType::TYPE_ACK};
}
case 't':
// This thread isn't part of the thread pool
CTX->Stop(false /* Ignore current thread */);
return {"OK", HandledPacketType::TYPE_ACK};
default:
return {"E00", HandledPacketType::TYPE_ACK};
}
if (ss->fail()) {
return {"E00", HandledPacketType::TYPE_ACK};
}
return ThreadAction(action, thread);
}
return {"", HandledPacketType::TYPE_ACK};
}
@@ -854,11 +989,24 @@ GdbServer::HandledPacketType GdbServer::ProcessPacket(const std::string &packet)
// Indicates the reason that the thread has stopped
// Behaviour changes if the target is in non-stop mode
// Binja doesn't support S response here
//return {"S00", HandledPacketType::TYPE_ACK};
auto str = fmt::format("T00thread:{:02x};core:2c;", getpid());
auto str = fmt::format("T00thread:{:x};", getpid());
return {std::move(str), HandledPacketType::TYPE_ACK};
}
case 'c':
// Continue
CTX->Run();
CTX->WaitForThreadsToRun();
return {"OK", HandledPacketType::TYPE_ACK};
case 'D':
// Detach
// Ensure the threads are back in running state on detach
CTX->Run();
CTX->WaitForThreadsToRun();
return {"OK", HandledPacketType::TYPE_ACK};
case 'g':
// We might be running while we try reading
// Pause up front
CTX->Pause();
return {readRegs(), HandledPacketType::TYPE_ACK};
case 'p':
return readReg(packet);
@@ -875,6 +1023,8 @@ GdbServer::HandledPacketType GdbServer::ProcessPacket(const std::string &packet)
case '!': // Enable extended mode
case 'T': // Is a thread alive?
return {"OK", HandledPacketType::TYPE_ACK};
case 's': // Step
return ThreadAction('s', 0);
case 'Z': // Inserts breakpoint or watchpoint
return handleBreakpoint(packet);
case 'k': // Kill the process
@@ -908,6 +1058,8 @@ void GdbServer::SendPacketPair(const HandledPacketType& response) {
}
void GdbServer::GdbServerLoop() {
OpenListenSocket();
while (!CTX->CoreShuttingDown.load()) {
CommsStream = OpenSocket();
@@ -953,6 +1105,8 @@ void GdbServer::GdbServerLoop() {
CommsStream.reset();
}
}
close(ListenSocket);
}
static void* ThreadHandler(void *Arg) {
FEXCore::GdbServer *This = reinterpret_cast<FEXCore::GdbServer*>(Arg);
@@ -961,49 +1115,52 @@ static void* ThreadHandler(void *Arg) {
}
void GdbServer::StartThread() {
gdbServerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
gdbServerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void GdbServer::OpenListenSocket() {
// open socket
struct addrinfo hints, *res;
memset(&hints, 0, sizeof(hints));
hints.ai_family = AF_UNSPEC;
hints.ai_socktype = SOCK_STREAM;
hints.ai_flags = AI_PASSIVE;
if(getaddrinfo(NULL, "8086", &hints, &res) < 0) {
perror("getaddrinfo");
}
int on = 1;
ListenSocket = socket(res->ai_family, res->ai_socktype, res->ai_protocol);
if (ListenSocket < 0) {
perror("socket");
}
if(setsockopt(ListenSocket, SOL_SOCKET, SO_REUSEADDR, (char*)&on, sizeof(on)) < 0) {
perror("setsockopt");
close(ListenSocket);
}
if (bind(ListenSocket, res->ai_addr, res->ai_addrlen) < 0) {
perror("bind");
close(ListenSocket);
}
listen(ListenSocket, 1);
}
std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
// open socket
int sockfd, new_fd;
// Block until a connection arrives
struct sockaddr_storage their_addr;
socklen_t addr_size;
struct addrinfo hints, *res;
struct sockaddr_storage their_addr;
socklen_t addr_size;
LogMan::Msg::IFmt("GdbServer, waiting for connection on localhost:8086");
int new_fd = accept(ListenSocket, (struct sockaddr *)&their_addr, &addr_size);
memset(&hints, 0, sizeof(hints));
hints.ai_family = AF_UNSPEC;
hints.ai_socktype = SOCK_STREAM;
hints.ai_flags = AI_PASSIVE;
if(getaddrinfo(NULL, "8086", &hints, &res) < 0) {
perror("getaddrinfo");
}
int on = 1;
sockfd = socket(res->ai_family, res->ai_socktype, res->ai_protocol);
if (sockfd < 0) {
perror("socket");
}
if(setsockopt(sockfd, SOL_SOCKET, SO_REUSEADDR, (char*)&on, sizeof(on)) < 0) {
perror("setsockopt");
}
if (bind(sockfd, res->ai_addr, res->ai_addrlen) < 0) {
perror("bind");
}
// Block until a connection arrives
LogMan::Msg::IFmt("GdbServer, waiting for connection on localhost:8086");
listen(sockfd, 1);
new_fd = accept(sockfd, (struct sockaddr *)&their_addr, &addr_size);
return std::make_unique<NetStream>(new_fd);
return std::make_unique<FEXCore::Utils::NetStream>(new_fd);
}
+8
View File
@@ -30,6 +30,7 @@ public:
private:
void Break(int signal);
void OpenListenSocket();
std::unique_ptr<std::iostream> OpenSocket();
void StartThread();
std::string ReadPacket(std::iostream &stream);
@@ -60,6 +61,8 @@ private:
HandledPacketType handleBreakpoint(const std::string &packet);
HandledPacketType handleProgramOffsets();
HandledPacketType ThreadAction(char action, uint32_t tid);
std::string readRegs();
HandledPacketType readReg(const std::string& packet);
@@ -69,8 +72,13 @@ private:
std::mutex sendMutex;
bool SettingNoAckMode{false};
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string MemoryMapString{};
std::string OSDataString{};
uint32_t CurrentDebuggingThread{};
int ListenSocket{};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
};
@@ -13,7 +13,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -974,55 +974,4 @@ DEF_OP(FCmp) {
#undef DEF_OP
void InterpreterOps::RegisterALUHandlers() {
#define REGISTER_OP(op, x) FEXCore::CPU::InterpreterOps::OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
REGISTER_OP(NEG, Neg);
REGISTER_OP(MUL, Mul);
REGISTER_OP(UMUL, UMul);
REGISTER_OP(DIV, Div);
REGISTER_OP(UDIV, UDiv);
REGISTER_OP(REM, Rem);
REGISTER_OP(UREM, URem);
REGISTER_OP(MULH, MulH);
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
REGISTER_OP(LUREM, LURem);
REGISTER_OP(NOT, Not);
REGISTER_OP(POPCOUNT, Popcount);
REGISTER_OP(FINDLSB, FindLSB);
REGISTER_OP(FINDMSB, FindMSB);
REGISTER_OP(FINDTRAILINGZEROS, FindTrailingZeros);
REGISTER_OP(COUNTLEADINGZEROES, CountLeadingZeroes);
REGISTER_OP(REV, Rev);
REGISTER_OP(BFI, Bfi);
REGISTER_OP(BFE, Bfe);
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -310,7 +310,7 @@ uint64_t AtomicCompareAndSwap(uint64_t expected, uint64_t desired, uint64_t *add
#endif
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
@@ -774,23 +774,5 @@ DEF_OP(AtomicFetchNeg) {
}
#undef DEF_OP
void InterpreterOps::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
REGISTER_OP(ATOMICSUB, AtomicSub);
REGISTER_OP(ATOMICAND, AtomicAnd);
REGISTER_OP(ATOMICOR, AtomicOr);
REGISTER_OP(ATOMICXOR, AtomicXor);
REGISTER_OP(ATOMICSWAP, AtomicSwap);
REGISTER_OP(ATOMICFETCHADD, AtomicFetchAdd);
REGISTER_OP(ATOMICFETCHSUB, AtomicFetchSub);
REGISTER_OP(ATOMICFETCHAND, AtomicFetchAnd);
REGISTER_OP(ATOMICFETCHOR, AtomicFetchOr);
REGISTER_OP(ATOMICFETCHXOR, AtomicFetchXor);
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -13,6 +13,7 @@ $end_info$
#include <FEXCore/HLE/SyscallHandler.h>
#include <cstdint>
#include <unistd.h>
namespace FEXCore::CPU {
[[noreturn]]
@@ -23,7 +24,7 @@ static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -104,6 +105,34 @@ DEF_OP(Syscall) {
GD = Res;
}
DEF_OP(InlineSyscall) {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
FEXCore::HLE::SyscallArguments Args;
for (size_t j = 0; j < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++j) {
if (Op->Header.Args[j].IsInvalid()) break;
Args.Argument[j] = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[j]);
}
// We don't want the errno handling but I also don't want to write inline ASM atm
uint64_t Res = syscall(
Op->HostSyscallNumber,
Args.Argument[0],
Args.Argument[1],
Args.Argument[2],
Args.Argument[3],
Args.Argument[4],
Args.Argument[5],
Args.Argument[6]
);
if (Res == -1) {
Res = -errno;
}
GD = Res;
}
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
@@ -137,21 +166,5 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void InterpreterOps::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,7 +11,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
uint8_t OpSize = IROp->Size;
@@ -220,18 +220,5 @@ DEF_OP(Vector_FToI) {
}
#undef DEF_OP
void InterpreterOps::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -299,7 +299,7 @@ namespace AES {
}
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -430,14 +430,5 @@ DEF_OP(AESKeyGenAssist) {
}
#undef DEF_OP
void InterpreterOps::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -13,7 +13,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(F80LOADFCW) {
FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle(*GetSrc<uint16_t*>(Data->SSAData, IROp->Args[0]));
}
@@ -357,33 +357,5 @@ DEF_OP(F80BCDSTORE) {
}
#undef DEF_OP
void InterpreterOps::RegisterF80Handlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(F80LOADFCW, F80LOADFCW);
REGISTER_OP(F80ADD, F80ADD);
REGISTER_OP(F80SUB, F80SUB);
REGISTER_OP(F80MUL, F80MUL);
REGISTER_OP(F80DIV, F80DIV);
REGISTER_OP(F80FYL2X, F80FYL2X);
REGISTER_OP(F80ATAN, F80ATAN);
REGISTER_OP(F80FPREM1, F80FPREM1);
REGISTER_OP(F80FPREM, F80FPREM);
REGISTER_OP(F80SCALE, F80SCALE);
REGISTER_OP(F80CVT, F80CVT);
REGISTER_OP(F80CVTINT, F80CVTINT);
REGISTER_OP(F80CVTTO, F80CVTTO);
REGISTER_OP(F80CVTTOINT, F80CVTTOINT);
REGISTER_OP(F80ROUND, F80ROUND);
REGISTER_OP(F80F2XM1, F80F2XM1);
REGISTER_OP(F80TAN, F80TAN);
REGISTER_OP(F80SQRT, F80SQRT);
REGISTER_OP(F80SIN, F80SIN);
REGISTER_OP(F80COS, F80COS);
REGISTER_OP(F80XTRACT_EXP, F80XTRACT_EXP);
REGISTER_OP(F80XTRACT_SIG, F80XTRACT_SIG);
REGISTER_OP(F80CMP, F80CMP);
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,17 +11,11 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
GD = (*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]) >> Op->Flag) & 1;
}
#undef DEF_OP
void InterpreterOps::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -37,8 +37,6 @@ public:
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeInterpreterOpHandlers();
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
@@ -35,24 +35,6 @@ static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
void InitializeInterpreterOpHandlers() {
for (uint32_t i = 0; i <= FEXCore::IR::IROps::OP_LAST; ++i) {
InterpreterOps::OpHandlers[i] = &InterpreterOps::Op_Unhandled;
}
InterpreterOps::RegisterALUHandlers();
InterpreterOps::RegisterAtomicHandlers();
InterpreterOps::RegisterBranchHandlers();
InterpreterOps::RegisterConversionHandlers();
InterpreterOps::RegisterFlagHandlers();
InterpreterOps::RegisterMemoryHandlers();
InterpreterOps::RegisterMiscHandlers();
InterpreterOps::RegisterMoveHandlers();
InterpreterOps::RegisterVectorHandlers();
InterpreterOps::RegisterEncryptionHandlers();
InterpreterOps::RegisterF80Handlers();
}
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: CTX {ctx}
, State {Thread} {
@@ -67,7 +49,6 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
@@ -13,8 +13,6 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
void InitializeInterpreterOpHandlers();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
@@ -161,19 +161,19 @@
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID()];
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, uint32_t Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op];
Res GetDest(void* SSAData, FEXCore::IR::NodeID Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID()];
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
@@ -21,29 +21,304 @@
#include <alloca.h>
#include <algorithm>
#include <array>
#include <atomic>
#include <bit>
#include <cmath>
#include <cstddef>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include <ctime>
#include <limits>
#include <memory>
#include <stddef.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
namespace FEXCore::CPU {
std::array<InterpreterOps::OpHandler, FEXCore::IR::IROps::OP_LAST + 1> InterpreterOps::OpHandlers;
void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node) {
using OpHandler = void (*)(IR::IROp_Header *IROp, InterpreterOps::IROpData *Data, IR::NodeID Node);
using OpHandlerArray = std::array<OpHandler, IR::IROps::OP_LAST + 1>;
constexpr OpHandlerArray InterpreterOpHandlers = [] {
OpHandlerArray Handlers{};
for (auto& Entry : Handlers) {
Entry = &InterpreterOps::Op_Unhandled;
}
#define REGISTER_OP(op, x) Handlers[IR::IROps::OP_##op] = &InterpreterOps::Op_##x
// ALU ops
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
REGISTER_OP(NEG, Neg);
REGISTER_OP(MUL, Mul);
REGISTER_OP(UMUL, UMul);
REGISTER_OP(DIV, Div);
REGISTER_OP(UDIV, UDiv);
REGISTER_OP(REM, Rem);
REGISTER_OP(UREM, URem);
REGISTER_OP(MULH, MulH);
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
REGISTER_OP(LUREM, LURem);
REGISTER_OP(NOT, Not);
REGISTER_OP(POPCOUNT, Popcount);
REGISTER_OP(FINDLSB, FindLSB);
REGISTER_OP(FINDMSB, FindMSB);
REGISTER_OP(FINDTRAILINGZEROS, FindTrailingZeros);
REGISTER_OP(COUNTLEADINGZEROES, CountLeadingZeroes);
REGISTER_OP(REV, Rev);
REGISTER_OP(BFI, Bfi);
REGISTER_OP(BFE, Bfe);
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
// Atomic ops
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
REGISTER_OP(ATOMICSUB, AtomicSub);
REGISTER_OP(ATOMICAND, AtomicAnd);
REGISTER_OP(ATOMICOR, AtomicOr);
REGISTER_OP(ATOMICXOR, AtomicXor);
REGISTER_OP(ATOMICSWAP, AtomicSwap);
REGISTER_OP(ATOMICFETCHADD, AtomicFetchAdd);
REGISTER_OP(ATOMICFETCHSUB, AtomicFetchSub);
REGISTER_OP(ATOMICFETCHAND, AtomicFetchAnd);
REGISTER_OP(ATOMICFETCHOR, AtomicFetchOr);
REGISTER_OP(ATOMICFETCHXOR, AtomicFetchXor);
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
// Flag ops
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
// Memory ops
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
REGISTER_OP(FILLREGISTER, FillRegister);
REGISTER_OP(LOADFLAG, LoadFlag);
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
// Misc ops
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
// Vector ops
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
REGISTER_OP(VAND, VAnd);
REGISTER_OP(VBIC, VBic);
REGISTER_OP(VOR, VOr);
REGISTER_OP(VXOR, VXor);
REGISTER_OP(VADD, VAdd);
REGISTER_OP(VSUB, VSub);
REGISTER_OP(VUQADD, VUQAdd);
REGISTER_OP(VUQSUB, VUQSub);
REGISTER_OP(VSQADD, VSQAdd);
REGISTER_OP(VSQSUB, VSQSub);
REGISTER_OP(VADDP, VAddP);
REGISTER_OP(VADDV, VAddV);
REGISTER_OP(VUMINV, VUMinV);
REGISTER_OP(VURAVG, VURAvg);
REGISTER_OP(VABS, VAbs);
REGISTER_OP(VPOPCOUNT, VPopcount);
REGISTER_OP(VFADD, VFAdd);
REGISTER_OP(VFADDP, VFAddP);
REGISTER_OP(VFSUB, VFSub);
REGISTER_OP(VFMUL, VFMul);
REGISTER_OP(VFDIV, VFDiv);
REGISTER_OP(VFMIN, VFMin);
REGISTER_OP(VFMAX, VFMax);
REGISTER_OP(VFRECP, VFRecp);
REGISTER_OP(VFSQRT, VFSqrt);
REGISTER_OP(VFRSQRT, VFRSqrt);
REGISTER_OP(VNEG, VNeg);
REGISTER_OP(VFNEG, VFNeg);
REGISTER_OP(VNOT, VNot);
REGISTER_OP(VUMIN, VUMin);
REGISTER_OP(VSMIN, VSMin);
REGISTER_OP(VUMAX, VUMax);
REGISTER_OP(VSMAX, VSMax);
REGISTER_OP(VZIP, VZip);
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
REGISTER_OP(VCMPGT, VCMPGT);
REGISTER_OP(VCMPGTZ, VCMPGTZ);
REGISTER_OP(VCMPLTZ, VCMPLTZ);
REGISTER_OP(VFCMPEQ, VFCMPEQ);
REGISTER_OP(VFCMPNEQ, VFCMPNEQ);
REGISTER_OP(VFCMPLT, VFCMPLT);
REGISTER_OP(VFCMPGT, VFCMPGT);
REGISTER_OP(VFCMPLE, VFCMPLE);
REGISTER_OP(VFCMPORD, VFCMPORD);
REGISTER_OP(VFCMPUNO, VFCMPUNO);
REGISTER_OP(VUSHL, VUShl);
REGISTER_OP(VUSHR, VUShr);
REGISTER_OP(VSSHR, VSShr);
REGISTER_OP(VUSHLS, VUShlS);
REGISTER_OP(VUSHRS, VUShrS);
REGISTER_OP(VSSHRS, VSShrS);
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VINSSCALARELEMENT, VInsScalarElement);
REGISTER_OP(VEXTRACTELEMENT, VExtractElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
REGISTER_OP(VUSHRNI, VUShrNI);
REGISTER_OP(VUSHRNI2, VUShrNI2);
REGISTER_OP(VBITCAST, VBitcast);
REGISTER_OP(VSXTL, VSXTL);
REGISTER_OP(VSXTL2, VSXTL2);
REGISTER_OP(VUXTL, VUXTL);
REGISTER_OP(VUXTL2, VUXTL2);
REGISTER_OP(VSQXTN, VSQXTN);
REGISTER_OP(VSQXTN2, VSQXTN2);
REGISTER_OP(VSQXTUN, VSQXTUN);
REGISTER_OP(VSQXTUN2, VSQXTUN2);
REGISTER_OP(VUMUL, VUMul);
REGISTER_OP(VSMUL, VSMul);
REGISTER_OP(VUMULL, VUMull);
REGISTER_OP(VSMULL, VSMull);
REGISTER_OP(VUMULL2, VUMull2);
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
// Encryption ops
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
REGISTER_OP(F80ADD, F80ADD);
REGISTER_OP(F80SUB, F80SUB);
REGISTER_OP(F80MUL, F80MUL);
REGISTER_OP(F80DIV, F80DIV);
REGISTER_OP(F80FYL2X, F80FYL2X);
REGISTER_OP(F80ATAN, F80ATAN);
REGISTER_OP(F80FPREM1, F80FPREM1);
REGISTER_OP(F80FPREM, F80FPREM);
REGISTER_OP(F80SCALE, F80SCALE);
REGISTER_OP(F80CVT, F80CVT);
REGISTER_OP(F80CVTINT, F80CVTINT);
REGISTER_OP(F80CVTTO, F80CVTTO);
REGISTER_OP(F80CVTTOINT, F80CVTTOINT);
REGISTER_OP(F80ROUND, F80ROUND);
REGISTER_OP(F80F2XM1, F80F2XM1);
REGISTER_OP(F80TAN, F80TAN);
REGISTER_OP(F80SQRT, F80SQRT);
REGISTER_OP(F80SIN, F80SIN);
REGISTER_OP(F80COS, F80COS);
REGISTER_OP(F80XTRACT_EXP, F80XTRACT_EXP);
REGISTER_OP(F80XTRACT_SIG, F80XTRACT_SIG);
REGISTER_OP(F80CMP, F80CMP);
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
return Handlers;
}();
void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
LOGMAN_MSG_A_FMT("Unhandled IR Op: {}", FEXCore::IR::GetName(IROp->Op));
}
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node) {
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
template<typename R, typename... Args>
FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
static FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
return {FABI_UNKNOWN, (void*)fn};
}
@@ -173,7 +448,7 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
decltype(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>) handlers[] = {
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
@@ -181,7 +456,8 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Inf
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7> };
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags]);
return true;
@@ -280,11 +556,11 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
auto CodeLast = CurrentIR->at(BlockIROp->Last);
for (auto [CodeNode, IROp] : CurrentIR->GetCode(BlockNode)) {
uint32_t ID = CurrentIR->GetID(CodeNode);
uint32_t Op = IROp->Op;
const auto ID = CurrentIR->GetID(CodeNode);
const uint32_t Op = IROp->Op;
// Execute handler
OpHandler Handler = InterpreterOps::OpHandlers[Op];
OpHandler Handler = InterpreterOpHandlers[Op];
Handler(IROp, &OpData, ID);
@@ -46,18 +46,6 @@ namespace FEXCore::CPU {
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
static void RegisterALUHandlers();
static void RegisterAtomicHandlers();
static void RegisterBranchHandlers();
static void RegisterConversionHandlers();
static void RegisterFlagHandlers();
static void RegisterMemoryHandlers();
static void RegisterMiscHandlers();
static void RegisterMoveHandlers();
static void RegisterVectorHandlers();
static void RegisterEncryptionHandlers();
static void RegisterF80Handlers();
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
uint64_t CurrentEntry{};
@@ -72,10 +60,7 @@ namespace FEXCore::CPU {
IR::NodeIterator BlockIterator{0, 0};
};
using OpHandler = std::function<void(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)>;
static std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers;
#define DEF_OP(x) static void Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) static void Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -159,6 +144,7 @@ namespace FEXCore::CPU {
DEF_OP(Jump);
DEF_OP(CondJump);
DEF_OP(Syscall);
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
@@ -407,4 +393,4 @@ namespace FEXCore::CPU {
}
};
};
} // namespace FEXCore::CPU
@@ -22,7 +22,7 @@ static inline void CacheLineFlush(char *Addr) {
#endif
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
@@ -265,25 +265,4 @@ DEF_OP(CacheLineClear) {
}
#undef DEF_OP
void InterpreterOps::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
REGISTER_OP(FILLREGISTER, FillRegister);
REGISTER_OP(LOADFLAG, LoadFlag);
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -22,7 +22,7 @@ static void StopThread(FEXCore::Core::InternalThreadState *Thread) {
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -42,9 +42,12 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 4: // HLT
case FEXCore::IR::Break_Halt: // HLT
StopThread(Data->State);
break;
case FEXCore::IR::Break_InvalidInstruction:
tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
break;
default: LOGMAN_MSG_A_FMT("Unknown Break Reason: {}", Op->Reason); break;
}
}
@@ -138,21 +141,5 @@ DEF_OP(Print) {
}
#undef DEF_OP
void InterpreterOps::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,7 +11,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
uintptr_t Src = GetSrc<uintptr_t>(Data->SSAData, Op->Header.Args[0]);
@@ -38,13 +38,5 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void InterpreterOps::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,7 +11,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(VectorZero) {
uint8_t OpSize = IROp->Size;
memset(GDP, 0, OpSize);
@@ -1931,101 +1931,5 @@ DEF_OP(VTBL1) {
}
#undef DEF_OP
void InterpreterOps::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
REGISTER_OP(VAND, VAnd);
REGISTER_OP(VBIC, VBic);
REGISTER_OP(VOR, VOr);
REGISTER_OP(VXOR, VXor);
REGISTER_OP(VADD, VAdd);
REGISTER_OP(VSUB, VSub);
REGISTER_OP(VUQADD, VUQAdd);
REGISTER_OP(VUQSUB, VUQSub);
REGISTER_OP(VSQADD, VSQAdd);
REGISTER_OP(VSQSUB, VSQSub);
REGISTER_OP(VADDP, VAddP);
REGISTER_OP(VADDV, VAddV);
REGISTER_OP(VUMINV, VUMinV);
REGISTER_OP(VURAVG, VURAvg);
REGISTER_OP(VABS, VAbs);
REGISTER_OP(VPOPCOUNT, VPopcount);
REGISTER_OP(VFADD, VFAdd);
REGISTER_OP(VFADDP, VFAddP);
REGISTER_OP(VFSUB, VFSub);
REGISTER_OP(VFMUL, VFMul);
REGISTER_OP(VFDIV, VFDiv);
REGISTER_OP(VFMIN, VFMin);
REGISTER_OP(VFMAX, VFMax);
REGISTER_OP(VFRECP, VFRecp);
REGISTER_OP(VFSQRT, VFSqrt);
REGISTER_OP(VFRSQRT, VFRSqrt);
REGISTER_OP(VNEG, VNeg);
REGISTER_OP(VFNEG, VFNeg);
REGISTER_OP(VNOT, VNot);
REGISTER_OP(VUMIN, VUMin);
REGISTER_OP(VSMIN, VSMin);
REGISTER_OP(VUMAX, VUMax);
REGISTER_OP(VSMAX, VSMax);
REGISTER_OP(VZIP, VZip);
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
REGISTER_OP(VCMPGT, VCMPGT);
REGISTER_OP(VCMPGTZ, VCMPGTZ);
REGISTER_OP(VCMPLTZ, VCMPLTZ);
REGISTER_OP(VFCMPEQ, VFCMPEQ);
REGISTER_OP(VFCMPNEQ, VFCMPNEQ);
REGISTER_OP(VFCMPLT, VFCMPLT);
REGISTER_OP(VFCMPGT, VFCMPGT);
REGISTER_OP(VFCMPLE, VFCMPLE);
REGISTER_OP(VFCMPORD, VFCMPORD);
REGISTER_OP(VFCMPUNO, VFCMPUNO);
REGISTER_OP(VUSHL, VUShl);
REGISTER_OP(VUSHR, VUShr);
REGISTER_OP(VSSHR, VSShr);
REGISTER_OP(VUSHLS, VUShlS);
REGISTER_OP(VUSHRS, VUShrS);
REGISTER_OP(VSSHRS, VSShrS);
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VINSSCALARELEMENT, VInsScalarElement);
REGISTER_OP(VEXTRACTELEMENT, VExtractElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
REGISTER_OP(VUSHRNI, VUShrNI);
REGISTER_OP(VUSHRNI2, VUShrNI2);
REGISTER_OP(VBITCAST, VBitcast);
REGISTER_OP(VSXTL, VSXTL);
REGISTER_OP(VSXTL2, VSXTL2);
REGISTER_OP(VUXTL, VUXTL);
REGISTER_OP(VUXTL2, VUXTL2);
REGISTER_OP(VSQXTN, VSQXTN);
REGISTER_OP(VSQXTN2, VSQXTN2);
REGISTER_OP(VSQXTUN, VSQXTUN);
REGISTER_OP(VSQXTUN2, VSQXTUN2);
REGISTER_OP(VUMUL, VUMul);
REGISTER_OP(VSMUL, VSMul);
REGISTER_OP(VUMULL, VUMull);
REGISTER_OP(VSMULL, VSMull);
REGISTER_OP(VUMULL2, VUMull2);
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -34,7 +34,7 @@ static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -9,7 +9,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
+160 -37
View File
@@ -11,12 +11,13 @@ $end_info$
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/MathUtils.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -105,23 +106,16 @@ DEF_OP(ExitFunction) {
}
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
Label *TargetLabel;
auto IsTarget = JumpTargets.find(Op->Header.Args[0].ID());
if (IsTarget == JumpTargets.end()) {
TargetLabel = &JumpTargets.try_emplace(Op->Header.Args[0].ID()).first->second;
}
else {
TargetLabel = &IsTarget->second;
}
PendingTargetLabel = TargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
#define GRFCMP(Node) (Op->CompareSize == 4 ? GetDst(Node).S() : GetDst(Node).D())
Condition MapBranchCC(IR::CondClassType Cond) {
static Condition MapBranchCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return Condition::eq;
case FEXCore::IR::COND_NEQ: return Condition::ne;
@@ -136,7 +130,7 @@ Condition MapBranchCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FLU: return Condition::lt;
case FEXCore::IR::COND_FGE: return Condition::ge;
case FEXCore::IR::COND_FLEU:return Condition::le;
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FGT: return Condition::gt;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:
@@ -153,22 +147,10 @@ Condition MapBranchCC(IR::CondClassType Cond) {
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
Label *TrueTargetLabel;
Label *FalseTargetLabel;
auto TrueIter = JumpTargets.find(Op->TrueBlock.ID());
auto FalseIter = JumpTargets.find(Op->FalseBlock.ID());
if (TrueIter == JumpTargets.end()) {
TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
}
else {
TrueTargetLabel = &TrueIter->second;
}
Label *TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
uint64_t Const;
bool isConst = IsInlineConstant(Op->Cmp2, &Const);
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_EQ) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
@@ -178,10 +160,11 @@ DEF_OP(CondJump) {
cbnz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
} else {
if (IsGPR(Op->Cmp1.ID())) {
if (isConst)
if (isConst) {
cmp(GRCMP(Op->Cmp1.ID()), Const);
else
} else {
cmp(GRCMP(Op->Cmp1.ID()), GRCMP(Op->Cmp2.ID()));
}
} else if (IsFPR(Op->Cmp1.ID())) {
fcmp(GRFCMP(Op->Cmp1.ID()), GRFCMP(Op->Cmp2.ID()));
} else {
@@ -191,13 +174,7 @@ DEF_OP(CondJump) {
b(TrueTargetLabel, MapBranchCC(Op->Cond));
}
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
else {
FalseTargetLabel = &FalseIter->second;
}
PendingTargetLabel = FalseTargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
DEF_OP(Syscall) {
@@ -235,6 +212,152 @@ DEF_OP(Syscall) {
mov(GetReg<RA_64>(Node), x0);
}
DEF_OP(InlineSyscall) {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
// Arguments are passed as follows:
// X8: SyscallNumber - RA INTERSECT
// X0: Arg0 & Return
// X1: Arg1
// X2: Arg2
// X3: Arg3
// X4: Arg4 - RA INTERSECT
// X5: Arg5 - RA INTERSECT
// X6: Arg6 - Doesn't exist in x86-64 land. RA INTERSECT
// One argument is removed from the SyscallArguments::MAX_ARGS since the first argument was syscall number
const static std::array<vixl::aarch64::Register, FEXCore::HLE::SyscallArguments::MAX_ARGS-1> RegArgs = {{
x0, x1, x2, x3, x4, x5
}};
bool Intersects{};
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
std::vector<vixl::aarch64::Register> IntersectRegs(FEXCore::HLE::SyscallArguments::MAX_ARGS);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
if (Reg.GetCode() == x8.GetCode() ||
Reg.GetCode() == x4.GetCode() ||
Reg.GetCode() == x5.GetCode()) {
SpillMask |= (1U << Reg.GetCode());
Intersects = true;
}
}
// XXX: For some reason spilling only the x4, x5, and x8 registers was causing issues
// Come back to this once investigation reveals why it fails the gvisor ioctl test
// For now override to all GPRs
SpillMask = ~0U;
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Now that we have claimed to be a syscall we can set up the arguments
if (Intersects) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
}
}
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
mov(RegArgs[i], Reg);
}
}
else {
auto Reg = GetReg<RA_32>(Op->Header.Args[i].ID());
if (Reg.GetCode() == w8.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == w4.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == w5.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
uxtw(RegArgs[i], Reg);
}
}
}
}
else {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
mov(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
else {
uxtw(RegArgs[i], GetReg<RA_32>(Op->Header.Args[i].ID()));
}
}
}
LoadConstant(x8, Op->HostSyscallNumber);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Result is now in x0
// Move result to its destination register
if (CTX->Config.Is64BitMode()) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), w0);
}
}
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
// Arguments are passed as follows:
@@ -256,7 +379,6 @@ DEF_OP(Thunk) {
FillStaticRegs(); // load from ctx after ra64 refill
}
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
const auto *OldCode = (const uint8_t *)&Op->CodeOriginalLow;
@@ -370,6 +492,7 @@ void Arm64JITCore::RegisterBranchHandlers() {
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
+18 -19
View File
@@ -42,7 +42,7 @@ void Arm64JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
using namespace vixl;
using namespace vixl::aarch64;
void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -307,7 +307,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
}
void Arm64JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
Arm64JITCore::CodeBuffer Arm64JITCore::AllocateNewCodeBuffer(size_t Size) {
@@ -396,6 +396,8 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
if (!CompileThread) {
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalReturnInstruction = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
@@ -482,7 +484,7 @@ Arm64JITCore::~Arm64JITCore() {
FreeCodeBuffer(InitialCodeBuffer);
}
IR::PhysicalRegister Arm64JITCore::GetPhys(uint32_t Node) const {
IR::PhysicalRegister Arm64JITCore::GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
@@ -491,7 +493,7 @@ IR::PhysicalRegister Arm64JITCore::GetPhys(uint32_t Node) const {
}
template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(uint32_t Node) const {
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
@@ -506,7 +508,7 @@ aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(uint32_t Node) const
}
template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(uint32_t Node) const {
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
@@ -521,18 +523,18 @@ aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(uint32_t Node) const
}
template<>
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_32>(uint32_t Node) const {
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_32>(IR::NodeID Node) const {
uint32_t Reg = GetPhys(Node).Reg;
return RA32Pair[Reg];
}
template<>
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_64>(uint32_t Node) const {
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_64>(IR::NodeID Node) const {
uint32_t Reg = GetPhys(Node).Reg;
return RA64Pair[Reg];
}
aarch64::VRegister Arm64JITCore::GetSrc(uint32_t Node) const {
aarch64::VRegister Arm64JITCore::GetSrc(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
@@ -546,7 +548,7 @@ aarch64::VRegister Arm64JITCore::GetSrc(uint32_t Node) const {
FEX_UNREACHABLE;
}
aarch64::VRegister Arm64JITCore::GetDst(uint32_t Node) const {
aarch64::VRegister Arm64JITCore::GetDst(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
@@ -588,18 +590,18 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
}
}
FEXCore::IR::RegisterClassType Arm64JITCore::GetRegClass(uint32_t Node) const {
FEXCore::IR::RegisterClassType Arm64JITCore::GetRegClass(IR::NodeID Node) const {
return FEXCore::IR::RegisterClassType {GetPhys(Node).Class};
}
bool Arm64JITCore::IsFPR(uint32_t Node) const {
bool Arm64JITCore::IsFPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
}
bool Arm64JITCore::IsGPR(uint32_t Node) const {
bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
@@ -672,7 +674,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
bind(&RunBlock);
}
//LOGMAN_THROW_A(RAData->HasFullRA(), "Arm64 JIT only works with RA");
//LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
SpillSlots = RAData->SpillSlots();
@@ -695,11 +697,8 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
#endif
{
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
@@ -716,7 +715,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
uint32_t ID = IR->GetID(CodeNode);
const auto ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
+21 -16
View File
@@ -64,6 +64,9 @@ public:
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
@@ -75,7 +78,7 @@ private:
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
std::map<IR::OrderedNodeWrapper::NodeOffsetType, aarch64::Label> JumpTargets;
std::map<IR::NodeID, aarch64::Label> JumpTargets;
/**
* @name Register Allocation
@@ -99,30 +102,30 @@ private:
constexpr static uint8_t RA_FPR = 2;
template<uint8_t RAType>
[[nodiscard]] aarch64::Register GetReg(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg(IR::NodeID Node) const;
template<>
[[nodiscard]] aarch64::Register GetReg<RA_32>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] aarch64::Register GetReg<RA_64>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_64>(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(IR::NodeID Node) const;
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(IR::NodeID Node) const;
[[nodiscard]] aarch64::VRegister GetSrc(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetDst(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetSrc(IR::NodeID Node) const;
[[nodiscard]] aarch64::VRegister GetDst(IR::NodeID Node) const;
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node) const;
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
[[nodiscard]] MemOperand GenerateMemOperand(uint8_t AccessSize,
aarch64::Register Base,
@@ -172,6 +175,7 @@ private:
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
uint64_t UnimplementedInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
@@ -181,8 +185,8 @@ private:
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
using OpHandler = void (Arm64JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
void RegisterBranchHandlers();
@@ -193,7 +197,7 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -277,6 +281,7 @@ private:
DEF_OP(Jump);
DEF_OP(CondJump);
DEF_OP(Syscall);
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
@@ -11,7 +11,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
+15 -14
View File
@@ -17,7 +17,7 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -38,20 +38,12 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 0: // Hard fault
case 5: // Guest ud2
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
case FEXCore::IR::Break_Overflow: // overflow
hlt(4);
break;
case 1: // Int <imm8>
hlt(4);
break;
case 2: // overflow
hlt(4);
break;
case 3: // int 1
hlt(4);
break;
case 4: { // HLT
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
@@ -62,13 +54,22 @@ DEF_OP(Break) {
br(TMP1);
break;
}
case 6: { // INT3
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
br(TMP1);
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->UnimplementedInstructionAddress);
br(TMP1);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
}
}
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VectorZero) {
uint8_t OpSize = IROp->Size;
switch (OpSize) {
@@ -20,7 +20,7 @@ namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -15,7 +15,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
@@ -28,7 +28,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -128,18 +128,10 @@ DEF_OP(ExitFunction) {
}
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
Label *TargetLabel;
auto IsTarget = JumpTargets.find(Op->Header.Args[0].ID());
if (IsTarget == JumpTargets.end()) {
TargetLabel = &JumpTargets.try_emplace(Op->Header.Args[0].ID()).first->second;
}
else {
TargetLabel = &IsTarget->second;
}
PendingTargetLabel = TargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
@@ -147,18 +139,7 @@ DEF_OP(Jump) {
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
Label *TrueTargetLabel;
Label *FalseTargetLabel;
auto TrueIter = JumpTargets.find(Op->TrueBlock.ID());
auto FalseIter = JumpTargets.find(Op->FalseBlock.ID());
if (TrueIter == JumpTargets.end()) {
TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
}
else {
TrueTargetLabel = &TrueIter->second;
}
Label *TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
if (IsGPR(Op->Cmp1.ID())) {
uint64_t Const;
@@ -168,24 +149,18 @@ DEF_OP(CondJump) {
cmp(GRCMP(Op->Cmp1.ID()), GRCMP(Op->Cmp2.ID()));
}
} else if (IsFPR(Op->Cmp1.ID())) {
if (Op->CompareSize == 4)
if (Op->CompareSize == 4) {
ucomiss(GetSrc(Op->Cmp1.ID()), GetSrc(Op->Cmp2.ID()));
else
} else {
ucomisd(GetSrc(Op->Cmp1.ID()), GetSrc(Op->Cmp2.ID()));
}
}
auto [_, __, JCC] = GetCC(Op->Cond);
(this->*JCC)(*TrueTargetLabel, T_NEAR);
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
else {
FalseTargetLabel = &FalseIter->second;
}
PendingTargetLabel = FalseTargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
DEF_OP(Syscall) {
@@ -15,7 +15,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
@@ -13,7 +13,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -14,7 +14,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
+29 -31
View File
@@ -98,7 +98,7 @@ void X86JITCore::PopRegs() {
add(rsp, 16 * RAXMM_x.size());
}
void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -308,7 +308,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
}
void X86JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
@@ -357,6 +357,8 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
@@ -431,7 +433,7 @@ void X86JITCore::ClearCache() {
}
}
IR::PhysicalRegister X86JITCore::GetPhys(uint32_t Node) const {
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
@@ -439,16 +441,16 @@ IR::PhysicalRegister X86JITCore::GetPhys(uint32_t Node) const {
return PhyReg;
}
bool X86JITCore::IsFPR(uint32_t Node) const {
bool X86JITCore::IsFPR(IR::NodeID Node) const {
return RAData->GetNodeRegister(Node).Class == IR::FPRClass.Val;
}
bool X86JITCore::IsGPR(uint32_t Node) const {
bool X86JITCore::IsGPR(IR::NodeID Node) const {
return RAData->GetNodeRegister(Node).Class == IR::GPRClass.Val;
}
template<uint8_t RAType>
Xbyak::Reg X86JITCore::GetSrc(uint32_t Node) const {
Xbyak::Reg X86JITCore::GetSrc(IR::NodeID Node) const {
// rax, rcx, rdx, rsi, r8, r9,
// r10
// Callee Saved
@@ -467,24 +469,24 @@ Xbyak::Reg X86JITCore::GetSrc(uint32_t Node) const {
}
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_64>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_64>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_32>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_32>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_16>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_16>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_8>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_8>(IR::NodeID Node) const;
Xbyak::Xmm X86JITCore::GetSrc(uint32_t Node) const {
Xbyak::Xmm X86JITCore::GetSrc(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
template<uint8_t RAType>
Xbyak::Reg X86JITCore::GetDst(uint32_t Node) const {
Xbyak::Reg X86JITCore::GetDst(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
if constexpr (RAType == RA_64)
return RA64[PhyReg.Reg].cvt64();
@@ -499,19 +501,19 @@ Xbyak::Reg X86JITCore::GetDst(uint32_t Node) const {
}
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_64>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_64>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_32>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_32>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_16>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_16>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_8>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_8>(IR::NodeID Node) const;
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(uint32_t Node) const {
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
if constexpr (RAType == RA_64)
return RA64Pair[PhyReg.Reg];
@@ -520,12 +522,12 @@ std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(uint32_t Node) const {
}
template
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_64>(uint32_t Node) const;
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_64>(IR::NodeID Node) const;
template
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_32>(uint32_t Node) const;
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_32>(IR::NodeID Node) const;
Xbyak::Xmm X86JITCore::GetDst(uint32_t Node) const {
Xbyak::Xmm X86JITCore::GetDst(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
@@ -687,15 +689,11 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
// if there is a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
@@ -724,10 +722,10 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
Inst << "\t" << Name << " ";
}
uint8_t NumArgs = IR::GetArgs(IROp->Op);
const uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t ArgNode = IROp->Args[i].ID();
uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
const auto ArgNode = IROp->Args[i].ID();
const uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
if (PhysReg >= GPRPairBase)
Inst << "Pair" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else if (PhysReg >= XMMBase)
@@ -739,7 +737,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
LogMan::Msg::DFmt("{}", Inst.str());
}
#endif
uint32_t ID = IR->GetID(CodeNode);
const auto ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
+19 -14
View File
@@ -9,8 +9,6 @@ $end_info$
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Common/MathUtils.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
@@ -20,6 +18,7 @@ using namespace Xbyak;
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/MathUtils.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <tuple>
@@ -80,6 +79,10 @@ public:
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
@@ -88,7 +91,7 @@ private:
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
uint64_t Entry;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, Label> JumpTargets;
std::unordered_map<IR::NodeID, Label> JumpTargets;
Xbyak::util::Cpu Features{};
bool MemoryDebug = false;
@@ -114,21 +117,21 @@ private:
constexpr static uint8_t RA_64 = 3;
constexpr static uint8_t RA_XMM = 4;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] Xbyak::Reg GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetSrc(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] Xbyak::Reg GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetDst(IR::NodeID Node) const;
[[nodiscard]] Xbyak::Xmm GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetSrc(IR::NodeID Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(IR::NodeID Node) const;
[[nodiscard]] Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
@@ -163,6 +166,8 @@ private:
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
@@ -177,8 +182,8 @@ private:
std::tuple<SetCC, CMovCC, JCC> GetCC(IR::CondClassType cond);
using OpHandler = void (X86JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
using OpHandler = void (X86JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
void RegisterBranchHandlers();
@@ -192,7 +197,7 @@ private:
void PushRegs();
void PopRegs();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -17,7 +17,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
+18 -14
View File
@@ -26,7 +26,7 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -47,20 +47,12 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 0: // Hard fault
case 5: // Guest ud2
ud2();
break;
case 1: // Int <imm8>
ud2();
break;
case 2: // overflow
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
case FEXCore::IR::Break_Overflow: // overflow
ud2();
break;
case 3: // int 1
ud2();
break;
case 4: { // HLT
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
@@ -70,7 +62,7 @@ DEF_OP(Break) {
jmp(TMP1);
break;
}
case 6: // INT3
case FEXCore::IR::Break_Interrupt3: // INT3
{
if (CTX->GetGdbServerStatus()) {
// Adjust the stack first for a regular return
@@ -93,6 +85,18 @@ DEF_OP(Break) {
}
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
mov(TMP1, ThreadSharedData.Dispatcher->UnimplementedInstructionAddress);
jmp(TMP1);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
}
}
@@ -15,7 +15,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -16,7 +16,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VectorZero) {
auto Dst = GetDst(Node);
vpxor(Dst, Dst, Dst);
+2 -2
View File
@@ -37,11 +37,11 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(PageMemory != -1ULL, "Failed to allocate page memory");
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
+1 -1
View File
@@ -50,7 +50,7 @@ public:
auto InsertPoint =
#endif
BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A(InsertPoint.second == true, "Dupplicate block mapping added");
LOGMAN_THROW_A_FMT(InsertPoint.second == true, "Dupplicate block mapping added");
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
CodePages[CurrentPage].push_back(Address);
File diff suppressed because it is too large. Load diff
+50 -25
View File
@@ -5,6 +5,7 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
@@ -26,18 +27,20 @@ class OpDispatchBuilder final : public IREmitter {
friend class FEXCore::IR::Pass;
friend class FEXCore::IR::PassManager;
enum {
FLAGS_OP_NONE, // must rely on x86 flags
FLAGS_OP_CMP, // flags were set by a CMP between flagsOpDest/flagsOpDestSigned and flagsOpSrc/flagsOpSrcSigned with flagsOpSize size
FLAGS_OP_AND, // flags were set by an AND/TEST, flagsOpDest contains the resulting value of flagsOpSize size
FLAGS_OP_FCMP, // flags were set by a ucomis* / comis*
enum class SelectionFlag {
Nothing, // must rely on x86 flags
CMP, // flags were set by a CMP between flagsOpDest/flagsOpDestSigned and flagsOpSrc/flagsOpSrcSigned with flagsOpSize size
AND, // flags were set by an AND/TEST, flagsOpDest contains the resulting value of flagsOpSize size
FCMP, // flags were set by a ucomis* / comis*
};
public:
int flagsOp;
uint8_t flagsOpSize;
OrderedNode* flagsOpDest, *flagsOpSrc;
OrderedNode* flagsOpDestSigned, *flagsOpSrcSigned;
SelectionFlag flagsOp{};
uint8_t flagsOpSize{};
OrderedNode* flagsOpDest{};
OrderedNode* flagsOpSrc{};
OrderedNode* flagsOpDestSigned{};
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
bool ShouldDump {false};
@@ -51,7 +54,7 @@ public:
OrderedNode* GetNewJumpBlock(uint64_t RIP) {
auto it = JumpTargets.find(RIP);
LOGMAN_THROW_A(it != JumpTargets.end(), "Couldn't find block generated for 0x%lx", RIP);
LOGMAN_THROW_A_FMT(it != JumpTargets.end(), "Couldn't find block generated for 0x{:x}", RIP);
return it->second.BlockEntry;
}
@@ -69,7 +72,7 @@ public:
}
void StartNewBlock() {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
}
bool FinishOp(uint64_t NextRIP, bool LastOp) {
@@ -231,9 +234,9 @@ public:
void PopcountOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
enum Segment {
Segment_FS,
Segment_GS,
enum class Segment {
FS,
GS,
};
template<Segment Seg>
void ReadSegmentReg(OpcodeArgs);
@@ -324,13 +327,22 @@ public:
template<size_t ElementSize>
void PSIGN(OpcodeArgs);
// BMI Ops
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
void BEXTRBMIOp(OpcodeArgs);
void BLSIBMIOp(OpcodeArgs);
void BLSMSKBMIOp(OpcodeArgs);
void BLSRBMIOp(OpcodeArgs);
// BMI2 Ops
void BMI2Shift(OpcodeArgs);
void BZHI(OpcodeArgs);
void MULX(OpcodeArgs);
void RORX(OpcodeArgs);
// ADX Ops
void ADXOp(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -486,6 +498,8 @@ public:
void UnimplementedOp(OpcodeArgs);
void InvalidOp(OpcodeArgs);
#undef OpcodeArgs
void SetPackedRFLAG(bool Lower8, OrderedNode *Src);
@@ -508,17 +522,28 @@ private:
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align);
uint8_t GetDstSize(FEXCore::X86Tables::DecodedOp Op) const;
uint8_t GetSrcSize(FEXCore::X86Tables::DecodedOp Op) const;
[[nodiscard]] static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
}
[[nodiscard]] static uint32_t MMBaseOffset() {
return static_cast<uint32_t>(offsetof(Core::CPUState, mm[0][0]));
}
[[nodiscard]] uint8_t GetDstSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint8_t GetSrcSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint32_t GetDstBitSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint32_t GetSrcBitSize(X86Tables::DecodedOp Op) const;
template<unsigned BitOffset>
void SetRFLAG(OrderedNode *Value) {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
_StoreFlag(_Bfe(1, 0, Value), BitOffset);
}
void SetRFLAG(OrderedNode *Value, unsigned BitOffset) {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
_StoreFlag(_Bfe(1, 0, Value), BitOffset);
}
@@ -547,13 +572,13 @@ private:
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
OrderedNode * GetX87Top();
enum X87Tag {
TAG_VALID = 0b00,
TAG_ZERO = 0b01,
TAG_SPECIAL = 0b10,
TAG_EMPTY = 0b11
enum class X87Tag {
Valid = 0b00,
Zero = 0b01,
Special = 0b10,
Empty = 0b11
};
void SetX87TopTag(OrderedNode *Value, uint32_t Tag);
void SetX87TopTag(OrderedNode *Value, X87Tag Tag);
OrderedNode *GetX87FTW(OrderedNode *Value);
void SetX87Top(OrderedNode *Value);
@@ -53,7 +53,7 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t RCON = Op->Src[1].Data.Literal.Value;
auto Res = _VAESKeyGenAssist(Src, RCON);
@@ -39,28 +39,30 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
};
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
uint8_t NumFlags = FlagOffsets.size();
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
NumFlags = 5;
}
auto OneConst = _Constant(1);
for (int i = 0; i < NumFlags; ++i) {
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffsets[i])), OneConst);
SetRFLAG(Tmp, FlagOffsets[i]);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffset)), OneConst);
SetRFLAG(Tmp, FlagOffset);
}
}
OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
OrderedNode *Original = _Constant(2);
uint8_t NumFlags = FlagOffsets.size();
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
NumFlags = 5;
}
for (int i = 0; i < NumFlags; ++i) {
OrderedNode *Flag = _LoadFlag(FlagOffsets[i]);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
OrderedNode *Flag = _LoadFlag(FlagOffset);
Flag = _Bfe(4, 32, 0, Flag);
Flag = _Lshl(Flag, _Constant(FlagOffsets[i]));
Flag = _Lshl(Flag, _Constant(FlagOffset));
Original = _Or(Original, Flag);
}
return Original;
@@ -130,13 +132,17 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
case 64:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", Size); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", Size);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
const auto SrcSize = GetSrcSize(Op);
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -146,7 +152,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -185,7 +191,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (GetSrcSize(Op)) {
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
@@ -198,7 +204,9 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", GetSrcSize(Op)); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
@@ -262,6 +270,8 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
}
void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
const auto SrcSize = GetSrcSize(Op);
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -271,7 +281,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -308,7 +318,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (GetSrcSize(Op)) {
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
@@ -321,7 +331,9 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", GetSrcSize(Op)); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
@@ -515,8 +515,8 @@ void OpDispatchBuilder::PSHUFBOp(OpcodeArgs) {
template<size_t ElementSize, bool HalfSize, bool Low>
void OpDispatchBuilder::PSHUFDOp(OpcodeArgs) {
LOGMAN_THROW_A(ElementSize != 0, "What. No element size?");
auto Size = GetSrcSize(Op);
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
@@ -552,8 +552,8 @@ void OpDispatchBuilder::PSHUFDOp<4, false, true>(OpcodeArgs);
template<size_t ElementSize>
void OpDispatchBuilder::SHUFOp(OpcodeArgs) {
LOGMAN_THROW_A(ElementSize != 0, "What. No element size?");
auto Size = GetSrcSize(Op);
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
@@ -620,7 +620,7 @@ void OpDispatchBuilder::PINSROp(OpcodeArgs) {
Src = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ElementSize, Op->Flags, -1);
}
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Index = Op->Src[1].Data.Literal.Value;
uint8_t NumElements = Size / ElementSize;
@@ -641,7 +641,7 @@ template
void OpDispatchBuilder::PINSROp<8>(OpcodeArgs);
void OpDispatchBuilder::InsertPSOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Imm = Op->Src[1].Data.Literal.Value;
uint8_t CountS = (Imm >> 6);
uint8_t CountD = (Imm >> 4) & 0b11;
@@ -689,7 +689,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Index = Op->Src[1].Data.Literal.Value;
const uint8_t NumElements = Size / ElementSize;
@@ -784,7 +784,7 @@ template<size_t ElementSize>
void OpDispatchBuilder::PSRLI(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t ShiftConstant = Op->Src[1].Data.Literal.Value;
auto Size = GetSrcSize(Op);
@@ -804,7 +804,7 @@ template<size_t ElementSize>
void OpDispatchBuilder::PSLLI(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t ShiftConstant = Op->Src[1].Data.Literal.Value;
auto Size = GetSrcSize(Op);
@@ -867,7 +867,7 @@ template
void OpDispatchBuilder::PSRAOp<4>(OpcodeArgs);
void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -879,7 +879,7 @@ void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
}
void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -892,7 +892,7 @@ void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
template<size_t ElementSize>
void OpDispatchBuilder::PSRAIOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1133,7 +1133,7 @@ void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *MemDest = _LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RDI]), GPRClass);
OrderedNode *MemDest = _LoadContext(GPRSize, GPROffset(X86State::REG_RDI), GPRClass);
const size_t NumElements = Size / 64;
for (size_t Element = 0; Element < NumElements; ++Element) {
@@ -1226,7 +1226,9 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
case 0x07: case 0x0F: case 0x17: case 0x1F: // Ordered
Result = _VFCMPORD(Size, ElementSize, Src2, Src);
break;
default: LOGMAN_MSG_A("Unknown Comparison type: %d", CompType);
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
break;
}
if constexpr (Scalar) {
@@ -1443,7 +1445,7 @@ void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
flagsOp = FLAGS_OP_FCMP;
flagsOp = SelectionFlag::FCMP;
flagsOpDest = Src1;
flagsOpSrc = Src2;
flagsOpSize = GetSrcSize(Op);
@@ -2110,7 +2112,7 @@ void OpDispatchBuilder::ExtendVectorElements<4, 8, true>(OpcodeArgs);
template<size_t ElementSize, bool Scalar>
void OpDispatchBuilder::VectorRound(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Mode = Op->Src[1].Data.Literal.Value;
uint64_t RoundControlSource = (Mode >> 2) & 1;
uint64_t RoundControl = Mode & 0b11;
@@ -2155,7 +2157,7 @@ void OpDispatchBuilder::VectorRound<8, true>(OpcodeArgs);
template<size_t ElementSize>
void OpDispatchBuilder::VectorBlend(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Select = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -2275,7 +2277,7 @@ void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
template<size_t ElementSize>
void OpDispatchBuilder::DPPOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Mask = Op->Src[1].Data.Literal.Value;
uint8_t SrcMask = Mask >> 4;
uint8_t DstMask = Mask & 0xF;
@@ -2323,7 +2325,7 @@ template
void OpDispatchBuilder::DPPOp<8>(OpcodeArgs);
void OpDispatchBuilder::MPSADBWOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Select = Op->Src[1].Data.Literal.Value;
// Src1 needs to be in byte offset
@@ -10,6 +10,7 @@ $end_info$
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
@@ -27,15 +28,15 @@ OrderedNode *OpDispatchBuilder::GetX87Top() {
return _LoadContext(1, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC, GPRClass);
}
void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, uint32_t Tag) {
void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, X87Tag Tag) {
// if we are popping then we must first mark this location as empty
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
Mask = _Lshl(Mask, TopOffset);
OrderedNode *NewFTW = _Andn(FTW, Mask);
if (Tag != 0) {
auto TagVal = _Lshl(_Constant(Tag), TopOffset);
if (Tag != X87Tag::Valid) {
auto TagVal = _Lshl(_Constant(ToUnderlying(Tag)), TopOffset);
NewFTW = _Or(NewFTW, TagVal);
}
@@ -72,7 +73,7 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
// Implicit arg
auto offset = _Constant(Op->OP & 7);
data = _And(_Add(orig_top, offset), mask);
data = _LoadContextIndexed(data, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
data = _LoadContextIndexed(data, 16, MMBaseOffset(), 16, FPRClass);
}
OrderedNode *converted = data;
@@ -82,10 +83,10 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
}
auto top = _And(_Sub(orig_top, _Constant(1)), mask);
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
// Write to ST[TOP]
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
//_StoreContext(converted, 16, offsetof(FEXCore::Core::CPUState, mm[7][0]));
}
@@ -101,25 +102,25 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
auto orig_top = GetX87Top();
auto mask = _Constant(7);
auto top = _And(_Sub(orig_top, _Constant(1)), mask);
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
// Read from memory
OrderedNode *data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags, -1);
OrderedNode *converted = _F80BCDLoad(data);
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *converted = _F80BCDStore(data);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
}
@@ -129,7 +130,7 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
// Update TOP
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto low = _Constant(Lower);
@@ -137,7 +138,7 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
// Write to ST[TOP]
_StoreContextIndexed(data, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -159,7 +160,7 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
// Update TOP
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
size_t read_width = GetSrcSize(Op);
@@ -190,13 +191,13 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
converted = _VInsElement(16, 8, 1, 0, converted, _VCastFromGPR(16, 8, upper));
// Write to ST[TOP]
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
}
template<size_t width>
void OpDispatchBuilder::FST(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
if constexpr (width == 80) {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, data, 10, 1);
}
@@ -207,7 +208,7 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
// Set the new top now
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
@@ -226,14 +227,14 @@ void OpDispatchBuilder::FIST(OpcodeArgs) {
auto Size = GetSrcSize(Op);
auto orig_top = GetX87Top();
OrderedNode *data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
data = _F80CVTInt(data, Truncate, Size);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, 1);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
// Set the new top now
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
@@ -274,22 +275,22 @@ void OpDispatchBuilder::FADD(OpcodeArgs) {
if constexpr (ResInST0 == OpResult::RES_STI) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Add(a, b);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -336,23 +337,23 @@ void OpDispatchBuilder::FMUL(OpcodeArgs) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Mul(a, b);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -399,10 +400,10 @@ void OpDispatchBuilder::FDIV(OpcodeArgs) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
if constexpr (reverse) {
@@ -414,14 +415,14 @@ void OpDispatchBuilder::FDIV(OpcodeArgs) {
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -483,10 +484,10 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
if constexpr (ResInST0 == OpResult::RES_STI) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
if constexpr (reverse) {
@@ -498,7 +499,7 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
@@ -506,7 +507,7 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -541,7 +542,7 @@ void OpDispatchBuilder::FSUB<32, true, true, OpDispatchBuilder::OpResult::RES_ST
void OpDispatchBuilder::FCHS(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(0);
auto high = _Constant(0b1'000'0000'0000'0000ULL);
@@ -551,12 +552,12 @@ void OpDispatchBuilder::FCHS(OpcodeArgs) {
auto result = _VXor(a, data, 16, 1);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FABS(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(~0ULL);
auto high = _Constant(0b0'111'1111'1111'1111ULL);
@@ -566,12 +567,12 @@ void OpDispatchBuilder::FABS(OpcodeArgs) {
auto result = _VAnd(a, data, 16, 1);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FTST(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(0);
OrderedNode *data = _VCastFromGPR(16, 8, low);
@@ -595,28 +596,28 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
void OpDispatchBuilder::FRNDINT(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Round(a);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto exp = _F80XTRACT_EXP(a);
auto sig = _F80XTRACT_SIG(a);
// Write to ST[TOP]
_StoreContextIndexed(exp, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(sig, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(exp, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(sig, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
@@ -661,10 +662,10 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
// Implicit arg
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *Res = _F80Cmp(a, b,
(1 << FCMP_FLAG_EQ) |
@@ -692,16 +693,16 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
if constexpr (poptwice) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
top = _And(_Add(top, _Constant(1)), mask);
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
else if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
@@ -738,12 +739,12 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
// Write to ST[TOP]
_StoreContextIndexed(b, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(b, top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FST(OpcodeArgs) {
@@ -756,14 +757,14 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
// Write to ST[TOP]
_StoreContextIndexed(a, arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
top = _And(_Add(top, _Constant(1)), _Constant(7));
SetX87Top(top);
}
@@ -772,14 +773,14 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
template<FEXCore::IR::IROps IROp>
void OpDispatchBuilder::X87UnaryOp(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Round(a);
// Overwrite the op
result.first->Header.Op = IROp;
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -798,8 +799,8 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
auto mask = _Constant(7);
OrderedNode *st1 = _And(_Add(top, _Constant(1)), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
st1 = _LoadContextIndexed(st1, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
st1 = _LoadContextIndexed(st1, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Add(a, st1);
// Overwrite the op
@@ -811,7 +812,7 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
}
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -842,17 +843,17 @@ void OpDispatchBuilder::X87ModifySTP<true>(OpcodeArgs);
void OpDispatchBuilder::X87SinCos(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto sin = _F80SIN(a);
auto cos = _F80COS(a);
// Write to ST[TOP]
_StoreContextIndexed(sin, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(cos, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(sin, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(cos, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
@@ -860,12 +861,12 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
auto orig_top = GetX87Top();
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
OrderedNode *st0 = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st0 = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
if (Plus1) {
auto low = _Constant(0x8000'0000'0000'0000ULL);
@@ -878,16 +879,16 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
auto result = _F80FYL2X(st0, st1);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87TAN(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80TAN(a);
@@ -897,24 +898,24 @@ void OpDispatchBuilder::X87TAN(OpcodeArgs) {
data = _VInsGPR(16, 8, data, high, 1);
// Write to ST[TOP]
_StoreContextIndexed(result, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(data, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87ATAN(OpcodeArgs) {
auto orig_top = GetX87Top();
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80ATAN(st1, a);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
@@ -1168,14 +1169,14 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
auto SevenConst = _Constant(7);
auto TenConst = _Constant(10);
for (int i = 0; i < 7; ++i) {
auto data = _LoadContextIndexed(Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
_StoreMem(FPRClass, 16, ST0Location, data, 1);
ST0Location = _Add(ST0Location, TenConst);
Top = _And(_Add(Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
auto data = _LoadContextIndexed(Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
@@ -1237,7 +1238,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// Mask off the top bits
Reg = _VAnd(16, 16, Reg, Mask);
_StoreContextIndexed(Reg, Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
ST0Location = _Add(ST0Location, TenConst);
Top = _And(_Add(Top, OneConst), SevenConst);
@@ -1252,12 +1253,12 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
ST0Location = _Add(ST0Location, _Constant(8));
OrderedNode *RegHigh = _LoadMem(FPRClass, 2, ST0Location, 1);
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
_StoreContextIndexed(Reg, Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *Result = _VExtractToGPR(16, 8, a, 1);
// Extract the sign bit
@@ -1333,7 +1334,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
Type = COMPARE_ZERO;
break;
default:
LOGMAN_MSG_A("Unhandled FCMOV op: 0x%x", Opcode);
LOGMAN_MSG_A_FMT("Unhandled FCMOV op: 0x{:x}", Opcode);
break;
}
@@ -1369,12 +1370,12 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
auto Result = _VBSL(VecCond, b, a);
// Write to ST[TOP]
_StoreContextIndexed(Result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87EMMS(OpcodeArgs) {
@@ -1391,7 +1392,7 @@ void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
top = _And(_Add(top, offset), _Constant(7));
// Set this argument's tag as empty now
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
}
}
+2 -2
View File
@@ -53,7 +53,7 @@ namespace FEXCore {
// Be warned, a thread will inherit the signal mask if created from this thread
int Result = pthread_sigmask(how, &SignalSet, nullptr);
if (Result != 0) {
LogMan::Msg::E("Couldn't register thread to mask signals");
LogMan::Msg::EFmt("Couldn't register thread to mask signals");
}
}
@@ -91,7 +91,7 @@ namespace FEXCore {
HostSignalHandler &Handler = HostHandlers[Signal];
if (!Thread) {
LogMan::Msg::E("[%d] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", ::gettid());
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", ::gettid());
}
else {
if (Handler.Handler &&
@@ -18,6 +18,7 @@ void InitializeH0F38Tables() {
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
static constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -99,6 +100,9 @@ void InitializeH0F38Tables() {
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_66, 0xF6), 1, X86InstInfo{"ADCX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F3, 0xF6), 1, X86InstInfo{"ADOX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
};
#undef OPD
@@ -393,16 +393,16 @@ void InitializeVEXTables() {
{OPD(2, 0b10, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b11, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b00, 0xF5), 1, X86InstInfo{"BZHI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b00, 0xF5), 1, X86InstInfo{"BZHI", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF5), 1, X86InstInfo{"PEXT", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF5), 1, X86InstInfo{"PDEP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b00, 0xF7), 1, X86InstInfo{"BEXTR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF7), 1, X86InstInfo{"SHLX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b10, 0xF7), 1, X86InstInfo{"SARX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xF7), 1, X86InstInfo{"SHLX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b10, 0xF7), 1, X86InstInfo{"SARX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
// VEX Map 3
{OPD(3, 0b01, 0x00), 1, X86InstInfo{"VPERMQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -479,7 +479,7 @@ void InitializeVEXTables() {
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b11, 0xF0), 1, X86InstInfo{"RORX", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b11, 0xF0), 1, X86InstInfo{"RORX", TYPE_INST, FLAGS_MODRM, 1, nullptr}},
// VEX Map 4 - 31 (Reserved)
};
@@ -34,7 +34,7 @@ static inline void GenerateTable(X86InstInfo *FinalTable, U8U8InfoStruct const *
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
@@ -51,7 +51,7 @@ static inline void GenerateTable(X86InstInfo *FinalTable, U16U8InfoStruct const
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
@@ -68,7 +68,7 @@ static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, U8U8InfoStruct
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if (Info.Type == TYPE_COPY_OTHER) {
FinalTable[OpNum + i] = OtherLocal[OpNum + i];
}
@@ -90,7 +90,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, U16U8InfoStruct con
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry %s->%s", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if ((OpNum & 0b11'000'000) == 0b11'000'000) {
// If the mod field is 0b11 then it is a regular op
FinalTable[OpNum + i] = Info;
@@ -98,7 +98,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, U16U8InfoStruct con
else {
// If the mod field is !0b11 then this instruction is duplicated through the whole mod [0b00, 0b10] range
// and the modrm.rm space because that is used part of the instruction encoding
LOGMAN_THROW_A((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
LOGMAN_THROW_A_FMT((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
for (uint16_t mod = 0b00'000'000; mod < 0b11'000'000; mod += 0b01'000'000) {
for (uint16_t rm = 0b000; rm < 0b1'000; ++rm) {
FinalTable[(OpNum | mod | rm) + i] = Info;
+11 -11
View File
@@ -66,27 +66,27 @@ namespace FEXCore {
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
LogMan::Msg::D("Load lib: %s -> %s", Name, SOName.c_str());
LogMan::Msg::DFmt("LoadLib: {} -> {}", Name, SOName);
auto Handle = dlopen(SOName.c_str(), RTLD_LOCAL | RTLD_NOW);
if (!Handle) {
LogMan::Msg::E("Load lib: failed to dlopen %s: %s", SOName.c_str(), dlerror());
return;
ERROR_AND_DIE_FMT("LoadLib: Failed to dlopen thunk library {}: {}", SOName, dlerror());
}
const auto InitSym = std::string("fexthunks_exports_") + Name;
ExportEntry* (*InitFN)(void *, uintptr_t);
auto InitSym = std::string("fexthunks_exports_") + (const char*)Name;
(void*&)InitFN = dlsym(Handle, InitSym.c_str());
if (!InitFN) {
LogMan::Msg::E("Load lib: failed to find export %s", InitSym.c_str());
return;
ERROR_AND_DIE_FMT("LoadLib: Failed to find export {}", InitSym);
}
auto Exports = InitFN((void*)&CallCallback, CallbackThunks);
if (!Exports) {
ERROR_AND_DIE_FMT("LoadLib: Failed to initialize thunk library {}. "
"Check if the corresponding host library is installed "
"or disable thunking of this library.", Name);
}
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
@@ -98,7 +98,7 @@ namespace FEXCore {
That->Thunks[*reinterpret_cast<IR::SHA256Sum*>(Exports[i].sha256)] = Exports[i].Fn;
}
LogMan::Msg::D("Loaded %d syms", i);
LogMan::Msg::DFmt("Loaded {} syms", i);
}
}
+96 -60
View File
@@ -1,74 +1,81 @@
{
"Defines": [
"constexpr static uint8_t COND_EQ = 0",
"constexpr static uint8_t COND_NEQ = 1",
"constexpr static uint8_t COND_UGE = 2",
"constexpr static uint8_t COND_ULT = 3",
"constexpr static uint8_t COND_MI = 4",
"constexpr static uint8_t COND_PL = 5",
"constexpr static uint8_t COND_VS = 6",
"constexpr static uint8_t COND_VC = 7",
"constexpr static uint8_t COND_UGT = 8",
"constexpr static uint8_t COND_ULE = 9",
"constexpr static uint8_t COND_SGE = 10",
"constexpr static uint8_t COND_SLT = 11",
"constexpr static uint8_t COND_SGT = 12",
"constexpr static uint8_t COND_SLE = 13",
"constexpr uint8_t COND_EQ = 0",
"constexpr uint8_t COND_NEQ = 1",
"constexpr uint8_t COND_UGE = 2",
"constexpr uint8_t COND_ULT = 3",
"constexpr uint8_t COND_MI = 4",
"constexpr uint8_t COND_PL = 5",
"constexpr uint8_t COND_VS = 6",
"constexpr uint8_t COND_VC = 7",
"constexpr uint8_t COND_UGT = 8",
"constexpr uint8_t COND_ULE = 9",
"constexpr uint8_t COND_SGE = 10",
"constexpr uint8_t COND_SLT = 11",
"constexpr uint8_t COND_SGT = 12",
"constexpr uint8_t COND_SLE = 13",
"constexpr static uint8_t COND_FLU = 16 /* float less or unordred */",
"constexpr static uint8_t COND_FGE = 17 /* float greater or equal */",
"constexpr static uint8_t COND_FLEU = 18 /* float less or equal or unordred */",
"constexpr static uint8_t COND_FGT = 19 /* float greater */",
"constexpr static uint8_t COND_FU = 20 /* float unordred */",
"constexpr static uint8_t COND_FNU = 21 /* float not unordred */",
"constexpr uint8_t COND_FLU = 16 /* float less or unordred */",
"constexpr uint8_t COND_FGE = 17 /* float greater or equal */",
"constexpr uint8_t COND_FLEU = 18 /* float less or equal or unordred */",
"constexpr uint8_t COND_FGT = 19 /* float greater */",
"constexpr uint8_t COND_FU = 20 /* float unordred */",
"constexpr uint8_t COND_FNU = 21 /* float not unordred */",
"static constexpr FEXCore::IR::RegisterClassType GPRClass {0}",
"static constexpr FEXCore::IR::RegisterClassType GPRFixedClass {1}",
"static constexpr FEXCore::IR::RegisterClassType FPRClass {2}",
"static constexpr FEXCore::IR::RegisterClassType FPRFixedClass {3}",
"static constexpr FEXCore::IR::RegisterClassType GPRPairClass {4}",
"static constexpr FEXCore::IR::RegisterClassType ComplexClass {5}",
"static constexpr FEXCore::IR::RegisterClassType InvalidClass {7}",
"constexpr FEXCore::IR::RegisterClassType GPRClass {0}",
"constexpr FEXCore::IR::RegisterClassType GPRFixedClass {1}",
"constexpr FEXCore::IR::RegisterClassType FPRClass {2}",
"constexpr FEXCore::IR::RegisterClassType FPRFixedClass {3}",
"constexpr FEXCore::IR::RegisterClassType GPRPairClass {4}",
"constexpr FEXCore::IR::RegisterClassType ComplexClass {5}",
"constexpr FEXCore::IR::RegisterClassType InvalidClass {7}",
"",
"static constexpr uint8_t InvalidReg {31}",
"constexpr uint8_t InvalidReg {31}",
"",
"static const FEXCore::IR::TypeDefinition i8 {TypeDefinition::Create(1, 0)}",
"static const FEXCore::IR::TypeDefinition i16 {TypeDefinition::Create(2, 0)}",
"static const FEXCore::IR::TypeDefinition i32 {TypeDefinition::Create(4, 0)}",
"static const FEXCore::IR::TypeDefinition i64 {TypeDefinition::Create(8, 0)}",
"static const FEXCore::IR::TypeDefinition i128 {TypeDefinition::Create(16, 0)}",
"constexpr FEXCore::IR::TypeDefinition i8 {TypeDefinition::Create(1, 0)}",
"constexpr FEXCore::IR::TypeDefinition i16 {TypeDefinition::Create(2, 0)}",
"constexpr FEXCore::IR::TypeDefinition i32 {TypeDefinition::Create(4, 0)}",
"constexpr FEXCore::IR::TypeDefinition i64 {TypeDefinition::Create(8, 0)}",
"constexpr FEXCore::IR::TypeDefinition i128 {TypeDefinition::Create(16, 0)}",
"",
"static const FEXCore::IR::TypeDefinition i8v8 {TypeDefinition::Create(1, 8)}",
"static const FEXCore::IR::TypeDefinition i8v16 {TypeDefinition::Create(1, 16)}",
"static const FEXCore::IR::TypeDefinition i16v4 {TypeDefinition::Create(2, 4)}",
"static const FEXCore::IR::TypeDefinition i16v8 {TypeDefinition::Create(2, 8)}",
"static const FEXCore::IR::TypeDefinition i32v2 {TypeDefinition::Create(4, 2)}",
"static const FEXCore::IR::TypeDefinition i32v4 {TypeDefinition::Create(4, 4)}",
"static const FEXCore::IR::TypeDefinition i64v2 {TypeDefinition::Create(8, 2)}",
"constexpr FEXCore::IR::TypeDefinition i8v8 {TypeDefinition::Create(1, 8)}",
"constexpr FEXCore::IR::TypeDefinition i8v16 {TypeDefinition::Create(1, 16)}",
"constexpr FEXCore::IR::TypeDefinition i16v4 {TypeDefinition::Create(2, 4)}",
"constexpr FEXCore::IR::TypeDefinition i16v8 {TypeDefinition::Create(2, 8)}",
"constexpr FEXCore::IR::TypeDefinition i32v2 {TypeDefinition::Create(4, 2)}",
"constexpr FEXCore::IR::TypeDefinition i32v4 {TypeDefinition::Create(4, 4)}",
"constexpr FEXCore::IR::TypeDefinition i64v2 {TypeDefinition::Create(8, 2)}",
"",
"constexpr static uint8_t FCMP_FLAG_EQ = 0",
"constexpr static uint8_t FCMP_FLAG_LT = 1",
"constexpr static uint8_t FCMP_FLAG_UNORDERED = 2",
"constexpr uint8_t FCMP_FLAG_EQ = 0",
"constexpr uint8_t FCMP_FLAG_LT = 1",
"constexpr uint8_t FCMP_FLAG_UNORDERED = 2",
"static constexpr FEXCore::IR::FenceType Fence_Load {0}",
"static constexpr FEXCore::IR::FenceType Fence_Store {1}",
"static constexpr FEXCore::IR::FenceType Fence_LoadStore {2}",
"constexpr FEXCore::IR::FenceType Fence_Load {0}",
"constexpr FEXCore::IR::FenceType Fence_Store {1}",
"constexpr FEXCore::IR::FenceType Fence_LoadStore {2}",
"constexpr static uint8_t ROUND_MODE_NEAREST = 0",
"constexpr static uint8_t ROUND_MODE_NEGATIVE_INFINITY = 1",
"constexpr static uint8_t ROUND_MODE_POSITIVE_INFINITY = 2",
"constexpr static uint8_t ROUND_MODE_TOWARDS_ZERO = 3",
"constexpr static uint8_t ROUND_MODE_FLUSH_TO_ZERO = 1 << 2",
"constexpr uint8_t ROUND_MODE_NEAREST = 0",
"constexpr uint8_t ROUND_MODE_NEGATIVE_INFINITY = 1",
"constexpr uint8_t ROUND_MODE_POSITIVE_INFINITY = 2",
"constexpr uint8_t ROUND_MODE_TOWARDS_ZERO = 3",
"constexpr uint8_t ROUND_MODE_FLUSH_TO_ZERO = 1 << 2",
"static constexpr FEXCore::IR::RoundType Round_Nearest {ROUND_MODE_NEAREST}",
"static constexpr FEXCore::IR::RoundType Round_Negative_Infinity {ROUND_MODE_NEGATIVE_INFINITY}",
"static constexpr FEXCore::IR::RoundType Round_Positive_Infinity {ROUND_MODE_POSITIVE_INFINITY}",
"static constexpr FEXCore::IR::RoundType Round_Towards_Zero {ROUND_MODE_TOWARDS_ZERO} /* Truncate */",
"static constexpr FEXCore::IR::RoundType Round_Host {ROUND_MODE_TOWARDS_ZERO + 1}",
"constexpr FEXCore::IR::RoundType Round_Nearest {ROUND_MODE_NEAREST}",
"constexpr FEXCore::IR::RoundType Round_Negative_Infinity {ROUND_MODE_NEGATIVE_INFINITY}",
"constexpr FEXCore::IR::RoundType Round_Positive_Infinity {ROUND_MODE_POSITIVE_INFINITY}",
"constexpr FEXCore::IR::RoundType Round_Towards_Zero {ROUND_MODE_TOWARDS_ZERO} /* Truncate */",
"constexpr FEXCore::IR::RoundType Round_Host {ROUND_MODE_TOWARDS_ZERO + 1}",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_SXTX {0};",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_UXTW {1};",
"constexpr static FEXCore::IR::MemOffsetType MEM_OFFSET_SXTW {2};"
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_SXTX {0}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_UXTW {1}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_SXTW {2}",
"constexpr FEXCore::IR::BreakReason Break_Unimplemented {0}",
"constexpr FEXCore::IR::BreakReason Break_Interrupt {1}",
"constexpr FEXCore::IR::BreakReason Break_Interrupt3 {2}",
"constexpr FEXCore::IR::BreakReason Break_Halt {3}",
"constexpr FEXCore::IR::BreakReason Break_Overflow {4}",
"constexpr FEXCore::IR::BreakReason Break_InvalidInstruction {5}"
],
"Ops": {
@@ -364,7 +371,7 @@
"HasSideEffects": true,
"OpClass": "Misc",
"Args": [
"uint8_t", "Reason",
"FEXCore::IR::BreakReason", "Reason",
"uint8_t", "Literal"
]
},
@@ -660,6 +667,8 @@
"Syscall": {
"HasSideEffects": true,
"Desc": ["Dispatches a guest syscall through to the SyscallHandler class"
],
"OpClass": "Branch",
"HasDest": true,
"DestClass": "GPR",
@@ -676,6 +685,33 @@
]
},
"InlineSyscall": {
"HasSideEffects": true,
"Desc": ["Dispatches a guest syscall directly to the host syscall interface,",
"bypassing the SyscallHandler class used by Syscall.",
"This has significantly less overhead than Syscall, which needs to save JIT state first.",
"Can only be used for syscalls that match across architecture,",
"such as gettid (matches on x86/x86-64/Arm64)."
],
"OpClass": "Branch",
"HasDest": true,
"DestClass": "GPR",
"FixedDestSize": "8",
"SSAArgs": "6",
"SSANames": [
"Arg0",
"Arg1",
"Arg2",
"Arg3",
"Arg4",
"Arg5"
],
"Args": [
"int32_t", "HostSyscallNumber"
]
},
"Thunk": {
"HasSideEffects": true,
"OpClass": "Branch",
+15 -13
View File
@@ -97,13 +97,14 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const*
static void PrintArg(std::stringstream *out, IRListView const* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationData *RAData) {
auto [CodeNode, IROp] = IR->at(Arg)();
const auto ArgID = Arg.ID();
if (Arg.ID() == 0) {
if (ArgID.IsInvalid()) {
*out << "%Invalid";
} else {
*out << "%ssa" << std::to_string(Arg.ID());
*out << "%ssa" << ArgID;
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(Arg.ID());
auto PhyReg = RAData->GetNodeRegister(ArgID);
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
@@ -191,17 +192,17 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%ssa" << std::to_string(IR->GetID(BlockNode)) << ") " << "CodeBlock ";
*out << "(%ssa" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "%ssa" << std::to_string(BlockIROp->Begin.ID()) << ", ";
*out << "%ssa" << std::to_string(BlockIROp->Last.ID()) << std::endl;
*out << "%ssa" << BlockIROp->Begin.ID() << ", ";
*out << "%ssa" << BlockIROp->Last.ID() << std::endl;
}
++CurrentIndent;
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
uint32_t ID = IR->GetID(CodeNode);
const auto ID = IR->GetID(CodeNode);
const auto Name = FEXCore::IR::GetName(IROp->Op);
auto Name = FEXCore::IR::GetName(IROp->Op);
bool Skip{};
switch (IROp->Op) {
case IR::OP_PHIVALUE:
@@ -224,7 +225,7 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
NumElements /= ElementSize;
}
*out << "%ssa" << std::to_string(ID);
*out << "%ssa" << ID;
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(ID);
@@ -264,10 +265,10 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
NumElements = IROp->Size / ElementSize;
}
*out << "(%ssa" << std::to_string(ID) << " ";
*out << "i" << std::dec << (ElementSize * 8);
*out << "(%ssa" << ID << ' ';
*out << 'i' << std::dec << (ElementSize * 8);
if (NumElements > 1) {
*out << "v" << std::dec << NumElements;
*out << 'v' << std::dec << NumElements;
}
*out << ") ";
}
@@ -289,8 +290,9 @@ void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationDa
PrintArg(out, IR, PhiOp->Block, RAData);
*out << " ]";
if (PhiOp->Next.ID())
if (PhiOp->Next.ID().IsValid()) {
*out << ", ";
}
NodeBegin = IR->at(PhiOp->Next);
}
+22
View File
@@ -38,6 +38,8 @@ enum class DecodeFailure {
DECODE_INVALID_CONDFLAG,
DECODE_INVALID_MEMOFFSETTYPE,
DECODE_INVALID_FENCETYPE,
DECODE_INVALID_BREAKTYPE,
};
std::string ltrim(std::string String) {
@@ -74,6 +76,7 @@ std::string DecodeErrorToString(DecodeFailure Failure) {
case DecodeFailure::DECODE_INVALID_CONDFLAG: return "Invalid Conditional name";
case DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE: return "Invalid Memory Offset Type";
case DecodeFailure::DECODE_INVALID_FENCETYPE: return "Invalid Fence Type";
case DecodeFailure::DECODE_INVALID_BREAKTYPE: return "Invalid Break Reason Type";
}
return "Unknown Error";
}
@@ -268,6 +271,25 @@ class IRParser: public FEXCore::IR::IREmitter {
return {DecodeFailure::DECODE_INVALID_FENCETYPE, {}};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::BreakReason> DecodeValue(const std::string &Arg) {
static constexpr std::array<std::string_view, 6> Names = {
"Unimplemented",
"Interrupt",
"Interrupt3",
"Halt",
"Overfloat",
"InvalidInstruction",
};
for (size_t i = 0; i < Names.size(); ++i) {
if (Names[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, BreakReason{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_BREAKTYPE, {}};
}
template<>
std::pair<DecodeFailure, OrderedNode*> DecodeValue(const std::string &Arg) {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
+1 -1
View File
@@ -287,7 +287,7 @@ void ConstProp::FCMPOptimization(IREmitter *IREmit, const IRListView& CurrentIR)
auto ghf = IROp->CW<IR::IROp_GetHostFlag>();
auto fcmp = IREmit->GetOpHeader(ghf->GPR)->CW<IR::IROp_FCmp>();
LOGMAN_THROW_A(fcmp->Header.Op == OP_FCMP || fcmp->Header.Op == OP_F80CMP, "Unexpected OP_GETHOSTFLAG source");
LOGMAN_THROW_A_FMT(fcmp->Header.Op == OP_FCMP || fcmp->Header.Op == OP_F80CMP, "Unexpected OP_GETHOSTFLAG source");
if(fcmp->Header.Op == OP_FCMP) {
fcmp->Flags |= 1 << ghf->Flag;
}
@@ -337,7 +337,7 @@ private:
std::unique_ptr<FEXCore::IR::Pass> DCE;
ContextInfo ClassifiedStruct;
std::unordered_map<FEXCore::IR::OrderedNodeWrapper::NodeOffsetType, BlockInfo> OffsetToBlockMap;
std::unordered_map<FEXCore::IR::NodeID, BlockInfo> OffsetToBlockMap;
ContextMemberInfo *FindMemberInfo(ContextInfo *ClassifiedInfo, uint32_t Offset, uint8_t Size);
ContextMemberInfo *RecordAccess(ContextMemberInfo *Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size, LastAccessType AccessType, FEXCore::IR::OrderedNode *Node, FEXCore::IR::OrderedNode *StoreNode = nullptr);
@@ -353,15 +353,15 @@ ContextMemberInfo *RCLSE::FindMemberInfo(ContextInfo *ContextClassificationInfo,
}
ContextMemberInfo *RCLSE::RecordAccess(ContextMemberInfo *Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size, LastAccessType AccessType, FEXCore::IR::OrderedNode *Node, FEXCore::IR::OrderedNode *StoreNode) {
LOGMAN_THROW_A((Offset + Size) <= (Info->Class.Offset + Info->Class.Size), "Access to context item went over member size");
LOGMAN_THROW_A(Info->Accessed != ACCESS_INVALID, "Tried to access invalid member");
LOGMAN_THROW_A_FMT((Offset + Size) <= (Info->Class.Offset + Info->Class.Size), "Access to context item went over member size");
LOGMAN_THROW_A_FMT(Info->Accessed != ACCESS_INVALID, "Tried to access invalid member");
// If we aren't fully overwriting the member then it is a partial write that we need to track
if (Size < Info->Class.Size) {
AccessType = AccessType == ACCESS_WRITE ? ACCESS_PARTIAL_WRITE : ACCESS_PARTIAL_READ;
}
if (Size > Info->Class.Size) {
LOGMAN_MSG_A("Can't handle this");
LOGMAN_MSG_A_FMT("Can't handle this");
}
Info->Accessed = AccessType;
@@ -503,7 +503,7 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
IREmit->Remove(LastStoreNode);
if (LastSize < IROp->Size) {
//printf("RCLSE: Eliminated partial write\n");
//fmt::print("RCLSE: Eliminated partial write\n");
}
Changed = true;
}
@@ -589,7 +589,7 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
RecordAccess(Info, Op->Class, Op->Offset, IROp->Size, ACCESS_READ, LastNode);
Changed = true;
} else {
//printf("RCLSE: Not GPR class, missed, %d, lastS: %d, S: %d, Node S: %d\n", LastClass, LastSize, IROp->Size, IREmit->GetOpSize(LastNode));
//fmt::print("RCLSE: Not GPR class, missed, {}, lastS: {}, S: {}, Node S: {}\n", LastClass, LastSize, IROp->Size, IREmit->GetOpSize(LastNode));
}
}
}
@@ -661,7 +661,9 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
}
else if (IROp->Op == OP_STORECONTEXTINDEXED ||
IROp->Op == OP_LOADCONTEXTINDEXED ||
IROp->Op == OP_SYSCALL) {
IROp->Op == OP_SYSCALL ||
IROp->Op == OP_INLINESYSCALL ||
IROp->Op == OP_BREAK) {
// We can't track through these
ResetClassificationAccesses(&LocalInfo);
}
@@ -116,7 +116,7 @@ uint64_t FPRBit(uint32_t Offset, uint32_t Size) {
else if (Size == 4)
return 1UL << (bitn);
else
LOGMAN_MSG_A("Unexpected FPR size %d", Size);
LOGMAN_MSG_A_FMT("Unexpected FPR size {}", Size);
return 7UL << (bitn); // Return maximum on failure case
}
+25 -21
View File
@@ -5,7 +5,6 @@ desc: Sorts the ssa storage in memory, needed for RA and others
$end_info$
*/
#include "Common/MathUtils.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -13,18 +12,19 @@ $end_info$
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <algorithm>
#include <cstdint>
#include <cstring>
#include <memory>
#include <stdint.h>
#include <string.h>
#include <vector>
namespace FEXCore::IR {
// struct to avoid zero-initialization
struct RemapNode {
IR::OrderedNodeWrapper::NodeOffsetType NodeID;
IR::NodeID NodeID;
};
static_assert(sizeof(RemapNode) == 4);
@@ -75,7 +75,7 @@ bool IRCompaction::Run(IREmitter *IREmit) {
auto HeaderNode = CurrentIR.GetHeaderNode();
auto HeaderOp = CurrentIR.GetHeader();
LOGMAN_THROW_A(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
LOGMAN_THROW_A_FMT(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
// This compaction pass is something that we need to ensure correct ordering and distances between IROps
// Later on we assume that an IROp's SSA value live range is its Node locations
@@ -92,17 +92,17 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// Then create all the ops inside the code blocks
// Zero is always zero(invalid)
OldToNewRemap[0].NodeID = 0;
OldToNewRemap[0].NodeID.Invalidate();
auto LocalHeaderOp = LocalBuilder._IRHeader(OrderedNodeWrapper::WrapOffset(0).GetNode(ListBegin), HeaderOp->BlockCount);
OldToNewRemap[CurrentIR.GetID(HeaderNode)].NodeID = LocalIR.GetID(LocalHeaderOp.Node);
OldToNewRemap[CurrentIR.GetID(HeaderNode).Value].NodeID = LocalIR.GetID(LocalHeaderOp.Node);
{
// Generate our codeblocks and link them together
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
LOGMAN_THROW_A(BlockHeader->Op == OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_A_FMT(BlockHeader->Op == OP_CODEBLOCK, "IR type failed to be a code block");
auto LocalBlockIRNode = LocalBuilder._CodeBlock(LocalHeaderOp, LocalHeaderOp); // Use LocalHeaderOp as a dummy arg for now
OldToNewRemap[CurrentIR.GetID(BlockNode)].NodeID = LocalIR.GetID(LocalBlockIRNode.Node);
OldToNewRemap[CurrentIR.GetID(BlockNode).Value].NodeID = LocalIR.GetID(LocalBlockIRNode.Node);
GeneratedCodeBlocks.emplace_back(CodeBlockData{BlockNode, LocalBlockIRNode});
}
@@ -110,8 +110,6 @@ bool IRCompaction::Run(IREmitter *IREmit) {
LocalHeaderOp.first->Blocks = GeneratedCodeBlocks[0].NewNode->Wrapped(LocalListBegin);
}
{
// Copy all of our IR ops over to the new location
for (auto &Block : GeneratedCodeBlocks) {
@@ -123,7 +121,8 @@ bool IRCompaction::Run(IREmitter *IREmit) {
CodeBlockData LastNode{};
uint32_t i {};
for (auto [CodeNode, IROp] : CurrentIR.GetCode(Block.OldNode)) {
size_t OpSize = FEXCore::IR::GetSize(IROp->Op);
const size_t OpSize = FEXCore::IR::GetSize(IROp->Op);
const auto CodeID = CurrentIR.GetID(CodeNode);
// Allocate the ops locally for our local dispatch
auto LocalPair = LocalBuilder.AllocateRawOp(OpSize);
@@ -137,7 +136,7 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// Set our map remapper to map the new location
// Even nodes that don't have a destination need to be in this map
// Need to be able to remap branch targets any other bits
OldToNewRemap[CurrentIR.GetID(CodeNode)].NodeID = LocalIR.GetID(LocalPair.Node);
OldToNewRemap[CodeID.Value].NodeID = LocalIR.GetID(LocalPair.Node);
if (i == 0) {
FirstNode.OldNode = CodeNode;
@@ -164,7 +163,7 @@ bool IRCompaction::Run(IREmitter *IREmit) {
for (auto &Block : GeneratedCodeBlocks) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = LocalIR.GetOp<FEXCore::IR::IROp_CodeBlock>(Block.NewNode);
LOGMAN_THROW_A(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
#endif
for (auto [LocalNode, LocalIROp] : LocalIR.GetCode(Block.NewNode)) {
@@ -172,13 +171,17 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// Now that we have the op copied over, we need to modify SSA values to point to the new correct locations
// This doesn't use IR::GetArgs(Op) because we need to remap all SSA nodes
// Including ones that we don't RA
uint8_t NumArgs = LocalIROp->NumArgs;
const uint8_t NumArgs = LocalIROp->NumArgs;
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t OldArg = LocalIROp->Args[i].ID();
const auto OldArg = LocalIROp->Args[i].ID();
const auto NewArg = OldToNewRemap[OldArg.Value].NodeID;
#ifndef NDEBUG
LOGMAN_THROW_A(OldToNewRemap[OldArg].NodeID != ~0U, "Tried remapping unfound node %%ssa%d", OldArg);
LOGMAN_THROW_A_FMT(NewArg.Value != UINT32_MAX,
"Tried remapping unfound node %ssa{}", OldArg);
#endif
LocalIROp->Args[i].NodeOffset = OldToNewRemap[OldArg].NodeID * sizeof(OrderedNode);
LocalIROp->Args[i].NodeOffset = NewArg.Value * sizeof(OrderedNode);
}
}
}
@@ -193,16 +196,17 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// if (NewListSize < OldListSize ||
// NewDataSize < OldDataSize) {
// if (NewListSize < OldListSize) {
// LogMan::Msg::D("Shaved %ld bytes off the list size", OldListSize - NewListSize);
// LogMan::Msg::DFmt("Shaved {} bytes off the list size", OldListSize - NewListSize);
// }
// if (NewDataSize < OldDataSize) {
// LogMan::Msg::D("Shaved %ld bytes off the data size", OldDataSize - NewDataSize);
// LogMan::Msg::DFmt("Shaved {} bytes off the data size", OldDataSize - NewDataSize);
// }
// }
// if (NewListSize > OldListSize ||
// NewDataSize > OldDataSize) {
// LOGMAN_MSG_A("Whoa. Compaction made the IR a different size when it shouldn't have. 0x%lx > 0x%lx or 0x%lx > 0x%lx",NewListSize, OldListSize, NewDataSize, OldDataSize);
// LOGMAN_MSG_A_FMT("Whoa. Compaction made the IR a different size when it shouldn't have. 0x{:x} > 0x{:x} or 0x{:x} > 0x{:x}",
// NewListSize, OldListSize, NewDataSize, OldDataSize);
// }
IREmit->CopyData(LocalBuilder);
+16 -17
View File
@@ -69,15 +69,12 @@ bool IRValidation::Run(IREmitter *IREmit) {
EntryBlock = BlockNode;
}
uint32_t BlockID = CurrentIR.GetID(BlockNode);
const auto BlockID = CurrentIR.GetID(BlockNode);
BlockInfo *CurrentBlock = &OffsetToBlockMap.try_emplace(BlockID).first->second;
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
uint32_t ID = CurrentIR.GetID(CodeNode);
uint8_t OpSize = IROp->Size;
const auto ID = CurrentIR.GetID(CodeNode);
const uint8_t OpSize = IROp->Size;
if (IROp->HasDest) {
HadError |= OpSize == 0;
@@ -130,7 +127,7 @@ bool IRValidation::Run(IREmitter *IREmit) {
case OP_PHIVALUE:
case OP_CONDJUMP:
case OP_JUMP:
// These override the nubmer of args for RA, so ignore them.
// These override the number of args for RA, so ignore them.
break;
default:
HadError |= true;
@@ -139,23 +136,25 @@ bool IRValidation::Run(IREmitter *IREmit) {
}
for (uint32_t i = 0; i < NumArgs; ++i) {
OrderedNodeWrapper Arg = IROp->Args[i];
const auto ArgID = Arg.ID();
// Was an argument defined after this node?
if (Arg.ID() >= ID) {
if (ArgID >= ID) {
HadError |= true;
Errors << "%ssa" << ID << ": Arg[" << i << "] has definition after use at %ssa" << Arg.ID() << std::endl;
Errors << "%ssa" << ID << ": Arg[" << i << "] has definition after use at %ssa" << ArgID << std::endl;
}
if (Arg.ID() != 0 && !NodeIsLive.Get(Arg.ID())) {
if (ArgID.IsValid() && !NodeIsLive.Get(ArgID.Value)) {
HadError |= true;
Errors << "%ssa" << ID << ": Arg[" << i << "] refrences dead %ssa" << Arg.ID() << std::endl;
Errors << "%ssa" << ID << ": Arg[" << i << "] references dead %ssa" << ArgID << std::endl;
}
if (Arg.ID() != 0) {
Uses[Arg.ID()]++;
if (ArgID.IsValid()) {
Uses[ArgID.Value]++;
}
}
NodeIsLive.Set(ID);
NodeIsLive.Set(ID.Value);
switch (IROp->Op) {
case IR::OP_EXITFUNCTION: {
@@ -259,8 +258,8 @@ bool IRValidation::Run(IREmitter *IREmit) {
}
}
for (int i = 0; i < CurrentIR.GetSSACount(); i++) {
auto [Node, IROp] = CurrentIR.at(i)();
for (uint32_t i = 0; i < CurrentIR.GetSSACount(); i++) {
auto [Node, IROp] = CurrentIR.at(IR::NodeID{i})();
if (Node->NumUses != Uses[i] && IROp->Op != OP_CODEBLOCK && IROp->Op != OP_IRHEADER) {
HadError |= true;
Errors << "%ssa" << i << " Has " << Uses[i] << " Uses, but reports " << Node->NumUses << std::endl;
@@ -282,7 +281,7 @@ bool IRValidation::Run(IREmitter *IREmit) {
LogMan::Msg::EFmt("{}", Out.str());
LOGMAN_MSG_A("Encountered IR validation Error");
LOGMAN_MSG_A_FMT("Encountered IR validation Error");
Errors.clear();
Warnings.clear();
+1 -1
View File
@@ -24,7 +24,7 @@ private:
BitSet<uint64_t> NodeIsLive;
OrderedNode *EntryBlock;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, BlockInfo> OffsetToBlockMap;
std::unordered_map<IR::NodeID, BlockInfo> OffsetToBlockMap;
size_t MaxNodes{};
friend class RAValidation;
+43 -43
View File
@@ -16,11 +16,11 @@ namespace FEXCore::IR::Validation {
// Hold the mapping of physical registers to the SSA id it holds at any given point in the IR
struct RegState {
static constexpr uint32_t UninitializedValue = 0;
static constexpr uint32_t InvalidReg = 0xffff'ffff;
static constexpr uint32_t CorruptedPair = 0xffff'fffe;
static constexpr uint32_t ClobberedValue = 0xffff'fffd;
static constexpr uint32_t StaticAssigned = 0xffff'ff00;
static constexpr IR::NodeID UninitializedValue{0};
static constexpr IR::NodeID InvalidReg {0xffff'ffff};
static constexpr IR::NodeID CorruptedPair {0xffff'fffe};
static constexpr IR::NodeID ClobberedValue {0xffff'fffd};
static constexpr IR::NodeID StaticAssigned {0xffff'ff00};
// This class makes some assumptions about how the host registers are arranged and mapped to virtual registers:
// 1. There will be less than 32 GPRs and 32 FPRs
@@ -31,10 +31,10 @@ struct RegState {
// These assumptions were all true for the state of the arm64 and x86 jits at the time this was written
// Mark a physical register as containing a SSA id
bool Set(PhysicalRegister Reg, uint32_t ssa) {
LOGMAN_THROW_A(ssa != 0, "RegState assumes ssa0 will be the block header and never assigned to a register");
bool Set(PhysicalRegister Reg, IR::NodeID ssa) {
LOGMAN_THROW_A_FMT(ssa.IsValid(), "RegState assumes ssa0 will be the block header and never assigned to a register");
// PhyscialRegisters aren't fully mapped until assembly emission
// PhysicalRegisters aren't fully mapped until assembly emission
// We need to apply a generic mapping here to catch any aliasing
switch (Reg.Class) {
case GPRClass:
@@ -65,7 +65,7 @@ struct RegState {
// Get the current SSA id
// Or an error value there isn't a (sane) SSA id
uint32_t Get(PhysicalRegister Reg) {
IR::NodeID Get(PhysicalRegister Reg) const {
switch (Reg.Class) {
case GPRClass:
return GPRs[Reg.Reg];
@@ -96,14 +96,14 @@ struct RegState {
// Mark a spill slot as containing a SSA id
void Spill(uint32_t SpillSlot, uint32_t ssa) {
void Spill(uint32_t SpillSlot, IR::NodeID ssa) {
Spills[SpillSlot] = ssa;
}
// Consume (and return) the SSA id currently in a spill slot
uint32_t Unspill(uint32_t SpillSlot) {
IR::NodeID Unspill(uint32_t SpillSlot) {
if (Spills.contains(SpillSlot)) {
uint32_t Value = Spills[SpillSlot];
const auto Value = Spills[SpillSlot];
Spills.erase(SpillSlot);
return Value;
}
@@ -143,7 +143,7 @@ struct RegState {
// Filter out all registers/slots containing an SSA id larger than MaxSSA
// Mark them as Clobbered.
// Useful for backwards edges, where using an SSA from before the
void Filter(uint32_t MaxSSA) {
void Filter(IR::NodeID MaxSSA) {
for (auto &gpr : GPRs) {
if (gpr > MaxSSA) {
gpr = ClobberedValue;
@@ -165,10 +165,10 @@ struct RegState {
}
private:
std::array<uint32_t, 32> GPRs = {};
std::array<uint32_t, 32> FPRs = {};
std::array<IR::NodeID, 32> GPRs = {};
std::array<IR::NodeID, 32> FPRs = {};
std::unordered_map<uint32_t, uint32_t> Spills;
std::unordered_map<uint32_t, IR::NodeID> Spills;
public:
uint32_t Version{}; // Used to force regeneration of RegStates after following backward edges
@@ -181,7 +181,7 @@ public:
private:
// Holds the calculated RegState at the exit of each block
std::unordered_map<uint32_t, RegState> BlockExitState;
std::unordered_map<IR::NodeID, RegState> BlockExitState;
// A queue of blocks we need to visit (or revisit)
std::deque<OrderedNode*> BlocksToVisit;
@@ -197,11 +197,11 @@ bool RAValidation::Run(IREmitter *IREmit) {
// Get the control flow graph from the validation pass
auto ValidationPass = Manager->GetPass<IRValidation>("IRValidation");
LOGMAN_THROW_A(ValidationPass != nullptr, "Couldn't find IRValidation pass");
LOGMAN_THROW_A_FMT(ValidationPass != nullptr, "Couldn't find IRValidation pass");
auto& OffsetToBlockMap = ValidationPass->OffsetToBlockMap;
LOGMAN_THROW_A(ValidationPass->EntryBlock != nullptr, "No entry point");
LOGMAN_THROW_A_FMT(ValidationPass->EntryBlock != nullptr, "No entry point");
BlocksToVisit.push_front(ValidationPass->EntryBlock); // Currently only a single entry point
bool HadError = false;
@@ -213,10 +213,10 @@ bool RAValidation::Run(IREmitter *IREmit) {
while (!BlocksToVisit.empty())
{
auto BlockNode = BlocksToVisit.front();
uint32_t BlockID = CurrentIR.GetID(BlockNode);
const auto BlockID = CurrentIR.GetID(BlockNode);
auto& BlockInfo = OffsetToBlockMap[BlockID];
auto IsFowardsEdge = [&] (uint32_t PredecessorID) {
const auto IsFowardsEdge = [&](IR::NodeID PredecessorID) {
// Blocks are sorted in FEXes IR, so backwards edges always go to a lower (or equal) Block ID
return PredecessorID < BlockID;
};
@@ -226,8 +226,8 @@ bool RAValidation::Run(IREmitter *IREmit) {
bool MissingPredecessor = false;
for (auto Predecessor : BlockInfo.Predecessors) {
auto PredecessorID = CurrentIR.GetID(Predecessor);
bool HaveState = BlockExitState.contains(PredecessorID) && BlockExitState[PredecessorID].Version == CurrentVersion;
const auto PredecessorID = CurrentIR.GetID(Predecessor);
const bool HaveState = BlockExitState.contains(PredecessorID) && BlockExitState[PredecessorID].Version == CurrentVersion;
if (IsFowardsEdge(PredecessorID) && !HaveState) {
// We are probably about to visit this node anyway, remove it
@@ -252,7 +252,7 @@ bool RAValidation::Run(IREmitter *IREmit) {
// Second, we need to determine the register status as of Block entry
auto BlockOp = CurrentIR.GetOp<IROp_CodeBlock>(BlockNode);
uint32_t FirstSSA = BlockOp->Begin.ID();
const auto FirstSSA = BlockOp->Begin.ID();
auto& BlockRegState = BlockExitState.try_emplace(BlockID).first->second;
bool EmptyRegState = true;
@@ -278,13 +278,14 @@ bool RAValidation::Run(IREmitter *IREmit) {
}
}
// Thrid, we need to iterate over all IR ops in the block
// Third, we need to iterate over all IR ops in the block
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
uint32_t ID = CurrentIR.GetID(CodeNode);
const auto ID = CurrentIR.GetID(CodeNode);
auto CheckArg = [&] (uint32_t i, OrderedNodeWrapper Arg) {
const auto PhyReg = RAData->GetNodeRegister(Arg.ID());
const auto CheckArg = [&](uint32_t i, OrderedNodeWrapper Arg) {
const auto ArgID = Arg.ID();
const auto PhyReg = RAData->GetNodeRegister(ArgID);
if (PhyReg.IsInvalid())
return;
@@ -300,21 +301,21 @@ bool RAValidation::Run(IREmitter *IREmit) {
auto Upper = BlockRegState.Get(PhysicalRegister(GPRClass, PhyReg.Reg*2 + 1));
Errors << fmt::format("%ssa{}: Arg[{}] expects paired reg{} to contain %ssa{}, but it actually contains {{%ssa{}, %ssa{}}}\n",
ID, i, PhyReg.Reg, Arg.ID(), Lower, Upper);
ID, i, PhyReg.Reg, ArgID, Lower, Upper);
} else if (CurrentSSAAtReg == RegState::UninitializedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but it is uninitialized\n",
ID, i, PhyReg.Reg, Arg.ID());
ID, i, PhyReg.Reg, ArgID);
} else if (CurrentSSAAtReg == RegState::ClobberedValue) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but contents vary depending on control flow\n",
ID, i, PhyReg.Reg, Arg.ID());
} else if (CurrentSSAAtReg != Arg.ID()) {
ID, i, PhyReg.Reg, ArgID);
} else if (CurrentSSAAtReg != ArgID) {
HadError |= true;
Errors << fmt::format("%ssa{}: Arg[{}] expects reg{} to contain %ssa{}, but it actually contains %ssa{}\n",
ID, i, PhyReg.Reg, Arg.ID(), CurrentSSAAtReg);
ID, i, PhyReg.Reg, ArgID, CurrentSSAAtReg);
}
};
@@ -330,8 +331,8 @@ bool RAValidation::Run(IREmitter *IREmit) {
case OP_FILLREGISTER: {
auto FillRegister = IROp->C<IROp_FillRegister>();
uint32_t ExpectedValue = FillRegister->OriginalValue.ID();
uint32_t Value = BlockRegState.Unspill(FillRegister->Slot);
const auto ExpectedValue = FillRegister->OriginalValue.ID();
const auto Value = BlockRegState.Unspill(FillRegister->Slot);
// TODO: This only proves that the Spill has a consistent SSA value
// In the future we need to prove it contains the correct SSA value
@@ -400,14 +401,14 @@ bool RAValidation::Run(IREmitter *IREmit) {
HadError |= true;
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
uint32_t BlockID = CurrentIR.GetID(BlockNode);
auto& BlockInfo = OffsetToBlockMap[BlockID];
const auto BlockID = CurrentIR.GetID(BlockNode);
const auto& BlockInfo = OffsetToBlockMap[BlockID];
Errors << fmt::format("Block {}\n\tPredecessors: ", BlockID);
for (auto Predecessor : BlockInfo.Predecessors) {
auto PredecessorID = CurrentIR.GetID(Predecessor);
bool FowardsEdge = PredecessorID < BlockID;
const auto PredecessorID = CurrentIR.GetID(Predecessor);
const bool FowardsEdge = PredecessorID < BlockID;
if (!FowardsEdge) {
Errors << "(Backwards): ";
}
@@ -417,8 +418,8 @@ bool RAValidation::Run(IREmitter *IREmit) {
Errors << "\n\tSuccessors: ";
for (auto Successor : BlockInfo.Successors) {
auto SuccessorID = CurrentIR.GetID(Successor);
bool FowardsEdge = SuccessorID > BlockID;
const auto SuccessorID = CurrentIR.GetID(Successor);
const bool FowardsEdge = SuccessorID > BlockID;
if (!FowardsEdge) {
Errors << "(Backwards): ";
@@ -440,8 +441,7 @@ bool RAValidation::Run(IREmitter *IREmit) {
FEXCore::IR::Dump(&IrDump, &CurrentIR, RAData);
LogMan::Msg::EFmt("RA Validation Error\n{}\nErrors:\n{}\n", IrDump.str(), Errors.str());
LOGMAN_MSG_A("Encountered RA validation Error");
LOGMAN_MSG_A_FMT("Encountered RA validation Error");
Errors.clear();
}
File diff suppressed because it is too large. Load diff
Loaded 100 of 255 files, more files were not shown because too many files have changed in this diff. Show more