Compare commits

..
576 Commits
Author SHA1 Message Date
Ryan Houdek c094dc238e Docs: Update for release FEX-2509.1 2025-09-15 18:33:36 -07:00
Billy Laws ceaf38e996 Dispatcher: Fix FABI_F32_I16_F80_PTR argument size
This takes an f80 as input and returns an f32. A copy-paste error had
this truncating the input float if !TMP_ABIARGS.
2025-09-15 18:31:55 -07:00
Billy Laws 85e9e255a5 unittests: Add test for x87 mode switches wrongly flushing NZCV 2025-09-15 18:31:50 -07:00
Billy Laws a545865ab7 OpcodeDispatcher: Only flush MMX registers on MMX -> x87 transitions
Flushing other regs is not necessary, and breaks any ConvertNZCVToX87 use
which relies previously saved NZCV values as the flag-setting NZCV op after
the save could trigger a flush of NZCV.
2025-09-15 18:31:44 -07:00
Billy Laws aa8e8f2cb0 OpcodeDispatcher: Don't assert on invalid ALU op encoding 2025-09-15 18:31:37 -07:00
Billy Laws d3a8701e1a WOW64: Fix CsSeg initialization 2025-09-15 18:31:31 -07:00
Ryan Houdek a6bd293ae8 Docs: Update for release FEX-2509 2025-09-09 09:21:22 -07:00
Ryan Houdek decd112a84 Merge pull request #4854 from lioncash/fend
Frontend/OpcodeDispatcher: Remove unused headers
2025-09-08 21:40:30 -07:00
Lioncache f9c15fa0a5 Frontend/OpcodeDispatcher: Remove unused headers
Removes assorted unused headers and forward declarations.
2025-09-09 00:28:05 -04:00
Ryan Houdek 1121f2a1fb Merge pull request #4853 from lioncash/allocator
FEXCore/Allocator: Remove unused headers
2025-09-08 20:25:25 -07:00
Lioncache a7989eb79f FEXCore/Allocator: Remove unused headers
Reveals some more indirect inclusions.
2025-09-08 22:53:09 -04:00
Ryan Houdek 80c199cefb Merge pull request #4852 from lioncash/loading
FileLoading: Add missing self-header
2025-09-08 19:48:01 -07:00
Ryan Houdek 317f92f4c8 Merge pull request #4851 from lioncash/ir
IR: Remove unnecessary forward declarations/headers
2025-09-08 19:47:51 -07:00
Ryan Houdek dd344710d4 Merge pull request #4850 from lioncash/vdso
VDSO_Emulator: Cull unnecessary includes
2025-09-08 19:16:56 -07:00
Lioncache e4e8683c7b FileLoading: Add missing self-header
Ensures that the exposed interface is visible to the implementation (i.e. gets rid of some -Wmissing-declaration warnings).
2025-09-08 22:13:42 -04:00
Lioncache b2407352a9 IR: Remove unnecessary forward declarations/headers
Pares the base IR header down to a modest size and also reveals a few indirect
reliances on the allocator facilities.
2025-09-08 22:03:25 -04:00
Lioncache e1b5e3e112 VDSO_Emulator: Cull unnecessary includes
Reveals a bunch of indirect includes inside the SignalDelegator.
2025-09-08 21:43:31 -04:00
Ryan Houdek 09793d295f Merge pull request #4849 from lioncash/config
Config: Remove unused Context.h include
2025-09-08 16:59:38 -07:00
Ryan Houdek e9f8c01e04 Merge pull request #4848 from lioncash/thread
InternalThreadState: Remove unused includes
2025-09-08 16:59:15 -07:00
Lioncache f69821d7db Config: Remove unused Context.h include
Removes quite a heavy include from the config system and specifies any
indirect inclusion that were relied on because of it.
2025-09-08 15:23:30 -04:00
Lioncache 1e03219ca3 Syscalls/Thread: Add missing forward declarations
Currently we were lucky that these were being seen already when being included.
2025-09-08 14:58:30 -04:00
Lioncache de8b4438e3 ThreadManager: Eliminate indirect inclusions
We had quite a bit of dependencies that were currently being indirectly relied upon.
We can forward declare the relevant ones and add the includes for the ones that are
explicit.
2025-09-08 14:58:27 -04:00
Ryan Houdek 2cfd3330f0 Merge pull request #4847 from lioncash/loader
CodeLoader: Tidy up interface
2025-09-08 11:50:30 -07:00
Ryan Houdek 1988b432ab Merge pull request #4846 from pmatos/fix/VSCond
Fix mapping for VS and VC condition codes
2025-09-08 11:29:35 -07:00
Lioncache 6b2363d3f6 InternalThreadState: Remove unused headers
Removes some includes in the core header and resolves indirect includes elsewhere.
2025-09-08 14:13:01 -04:00
Ryan Houdek e76af56c86 Merge pull request #4845 from lioncash/thunk
Core/Thunks: Remove unused IR header/forward declarations
2025-09-08 10:03:34 -07:00
Ryan Houdek b7173a3ae2 Merge pull request #4844 from lioncash/builtin
General: std::alignment_of -> alignof
2025-09-08 10:03:24 -07:00
Ryan Houdek 3f71517e1f Merge pull request #4843 from lioncash/util
Common: Remove duplicated StringUtil functionality
2025-09-08 10:03:14 -07:00
Lioncache 17fbe52cf6 CodeLoader: Mark GetStackPointer() as const
This isn't intended to modify anything.
2025-09-08 12:19:15 -04:00
Lioncache ecc526a9d8 CodeLoader: Make GetApplicationArguments() return by reference
We don't really need to make this potentially return a null pointer when we could always return a valid instance
that would just happen to be empty if unused.
2025-09-08 12:07:19 -04:00
Lioncache 436f4aa953 CodeLoader: Make GetExecveArguments() return by value
Same behavior, but more straightforward.
2025-09-08 12:07:19 -04:00
Lioncache fd10ce0600 CodeLoader: Return auxv results by value
Same behavior, but without the need to care about out values.
2025-09-08 12:07:14 -04:00
Lioncache d9cd40eb48 CodeLoader: Remove unused AddIR() member function
This isn't used in any implementations anymore.
2025-09-08 10:34:13 -04:00
Lioncache 9b9ab16b55 Common: Remove duplicated StringUtil functionality
We already have an equivalent header within FEXCore that's header only,
so we can adapt it to conform for both cases, allowing for removal of
one of them.
2025-09-08 09:44:25 -04:00
Paulo Matos 14256e7a90 Fix mapping for VS and VC condition codes
This doesn't seem to be generated in-code therefore it's not fixing any
existing bug, but fixes the mapping for future uses of the condition code.
2025-09-08 11:41:15 +02:00
Lioncache d36c387c1f Core/Thunks: Remove unused IR/forward declarations
Moves the IR include into one of the more specific headers, which avoids dumping the IR header into any core bits that use the interface.

Also uncovered a missing header guard.
2025-09-06 22:28:27 -04:00
Lioncache 7e61b172b2 General: std::alignment_of -> alignof
We can just use the compiler built-in in these cases. std::alignment_of
mainly has utility in metaprogramming.
2025-09-06 21:23:10 -04:00
Ryan Houdek bd0cae9298 Merge pull request #4825 from bylaws/unityomg
Frontend: Force acq/rel semantics for known Unity ringbuffer offsets
2025-09-06 17:09:13 -07:00
Ryan Houdek 6ae94581bb Merge pull request #4821 from bylaws/monof
Frontend: Fix tailcall handling when mono hacks are enabled
2025-09-06 16:51:08 -07:00
Ryan Houdek ea8ae2bf2c Merge pull request #4839 from lioncash/utils
Common Utils: Remove unused/unnecessary inclusions from headers
2025-09-06 15:50:00 -07:00
Ryan Houdek 694bc3312e Merge pull request #4842 from lioncash/python
config_generator: Remove superfluous format arguments
2025-09-06 15:49:50 -07:00
Ryan Houdek f565939437 Merge pull request #4841 from lioncash/env
Common: Remove EnvironmentLoader.cpp/.h
2025-09-06 15:49:42 -07:00
Lioncache 272d2747cd config_generator: Remove excess formatting arguments for EnumParser handling
Each instance had one more argument than necessary.
2025-09-06 11:18:38 -04:00
Lioncache 675956c92d config_generator: Indent ENVLOADER conditional bodies
Makes the indentation consistent with the rest of the generated file.
2025-09-06 10:57:31 -04:00
Lioncache 03410ed4a2 config_generator: Remove unnecessary format() in print_parse_jsonloader_options()
This format call essentially does nothing since it has no specifier.
We can also fix up the indentation of the else case too, for consistency.
2025-09-06 10:51:02 -04:00
Lioncache ac88e7fabf Common: Remove EnvironmentLoader.cpp/.h
This has been unused since the transition to the layered config system.
2025-09-06 10:23:53 -04:00
Tony Wasserka 71d42d7e73 Merge pull request #4840 from lioncash/internal
InternalThreadState: Make bool operator explicit
2025-09-06 09:20:08 +02:00
Lioncache 2613762ac2 InternalThreadState: Make bool operator explicit
We definitely don't want the boolean null test to be able to be implicitly converted
(e.g. to an int or whatever else).

For example a non-explicit bool operator allows for silly things like:

NonMovableUniquePtr<...> ptr;
// ...
auto k = 5 + ptr;

to build without issue, which we should really force the user to be explicit about if it's *really* a desired behavior.
2025-09-05 16:29:19 -04:00
Lioncache 831b21cd39 JitSymbols: Remove unused includes
Reduces header dependencies (and clarifies existing ones).
2025-09-05 14:27:49 -04:00
Lioncache 7869d80073 VolatileMetadata: Add missing include 2025-09-05 14:16:23 -04:00
Lioncache ea39d80e83 StringUtil: Move algorithm header to cpp file
The header doesn't make any direct use of it.
2025-09-05 14:16:23 -04:00
Lioncache 62e481a608 JSONPool: Remove unnecessary bounds check
Given the generic context here, we should just use std::data
2025-09-05 14:16:20 -04:00
Lioncache bac7148d69 JSONPool: Remove unnecessary include
This isn't used in this header.
2025-09-05 13:52:05 -04:00
Lioncache 62f16a3c4b FEXServerClient: Remove unnecessary includes from header
We can reduce includes, such as logging by specifying a concrete size for logging levels,
allowing the enum to be forward declared. We can also move FillHeader into the cpp file,
allowing the syscalls header to be removed.
2025-09-05 13:45:31 -04:00
Lioncache 248ecd54e4 FDUtils: Remove unused headers 2025-09-05 13:30:09 -04:00
Lioncache 1f59accb4e CPUInfo: Remove unused header 2025-09-05 13:28:54 -04:00
Lioncache dd7e4375cf Config: Add missing includes 2025-09-05 13:26:43 -04:00
LC 3df3999138 Merge pull request #4838 from lioncash/dumper
IRDumper: Add missing PrintArg for IndexNamedVectorConstant
2025-09-05 13:06:53 -04:00
Ryan Houdek 1ca6443c9d Merge pull request #4836 from neobrain/refactor_irscript_cleanup
json_ir_generator: Clean up generated assertion statements
2025-09-04 11:21:10 -07:00
Ryan Houdek a804cb5d07 Merge pull request #4833 from lioncash/ptr
CPUBackend: Remove unused variable in CheckCodeBufferUpdate
2025-09-04 11:20:42 -07:00
Ryan Houdek a0883c9615 Merge pull request #4832 from lioncash/threads
Utils/Threads: Minor cleanup
2025-09-04 11:20:24 -07:00
Ryan Houdek 03c968a8fe Merge pull request #4831 from lioncash/header
IntrusiveIRList/RegisterAllocationData: Remove unused headers
2025-09-04 11:20:14 -07:00
Ryan Houdek e34de22cf4 Merge pull request #4829 from lioncash/clz
VixlUtils: Use std::countl_zero where applicable
2025-09-04 11:19:54 -07:00
Ryan Houdek 3db1f1e01f Merge pull request #4826 from Sonicadvance1/classify_features
unittests/ASM: Adds support for classifying more requirements
2025-09-04 11:19:41 -07:00
Ryan Houdek 5e467c737f Merge pull request #4822 from bylaws/ecent
Dispatcher: Move the EC entrypoint check prior to the L1 lookup
2025-09-04 11:19:29 -07:00
Lioncache dfc42fa7c5 IRDumper: Add missing PrintArg for IndexNamedVectorConstant
Results in better named output for these.
2025-09-04 12:04:48 -04:00
LC f1016e82c7 Merge pull request #4834 from neobrain/refactor_async_simplify_ancbuffer
AsyncNet: Simplify declaration of ancillary buffer
2025-09-04 11:28:41 -04:00
Tony Wasserka 5e099a5dbd IR: Fix indentation 2025-09-04 15:39:40 +02:00
Tony Wasserka 8dceec9316 IR: Fix array emptiness check 2025-09-04 15:39:40 +02:00
Tony Wasserka 6f51092998 AsyncNet: Simplify declaration of ancillary buffer 2025-09-04 14:40:39 +02:00
Lioncache 4e83ca67f5 CPUBackend: std::move shared_ptr where applicable
Very minor, just avoids churning reference count increments and decrements.
2025-09-03 21:20:58 -04:00
Lioncache 76e6c6f8ea CPUBackend: Remove unused shared_ptr in CheckCodeBufferUpdate()
Just gets rid of an unused variable.
2025-09-03 21:19:16 -04:00
Ryan Houdek 343a529e63 Merge pull request #4824 from bylaws/tf2
OpDispatcher: Force a recheck of TF after POPF
2025-09-03 15:24:40 -07:00
Ryan Houdek c6ccc68607 Merge pull request #4823 from bylaws/mbooo
Frontend: Always explore fallthrough branches with multiblock
2025-09-03 15:24:29 -07:00
Ryan Houdek 8009f4a6ee Merge pull request #4820 from bylaws/neg
AtomicOps: Reimplement AtomicNeg using 8.1 CAS atomics
2025-09-03 15:23:40 -07:00
Tony Wasserka e126cecb16 Merge pull request #4818 from lioncash/at
General: Replace usages of .at(0) with .data() where applicable
2025-09-03 22:37:50 +02:00
Tony Wasserka 5e19202540 Merge pull request #4816 from lioncash/float
F80Fallbacks: Remove unnecessary memset
2025-09-03 22:36:59 +02:00
Billy Laws 06466bac0e Dispatcher: Drop the L1 lookup
This is almost never hit outside of exceptional cases like signals/SMC,
for which other overheads dominate.
2025-09-03 18:26:47 +01:00
Billy Laws 156edcacb2 Dispatcher: Move the EC entrypoint check prior to the L1 lookup
Profiling shows this is by far the hottest path, as every DX call will
hit this (as arm64ec functions in a vtable). L2 hits are much rarer and
L1 almost never hits since that is looked up inline in most cases.
2025-09-03 18:26:47 +01:00
Lioncache bf4b8d246a Threads: Simplify copying in SetInternalPointers
We can just use assignment here, which lets us also
remove <cstring> since this is the only usage of memcpy.
2025-09-03 11:01:12 -04:00
Lioncache abc37ec35e Threads: Mark default handler functions as static
These aren't used outside of the translation unit.
2025-09-03 11:01:12 -04:00
Lioncache dd735464a2 Threads: Remove unnecessary includes
We have quite a few unnecessary includes here left over from code movement,
so we can clean those out.

Also fix an indirect include in the header.
2025-09-03 11:00:57 -04:00
Lioncache 7041bdb144 IntrusiveIRList: Remove unused headers
Gets rid of unnecessary dependencies and clarifies previously indirect dependencies.
2025-09-03 10:39:45 -04:00
Lioncache 19e8c3fcb6 RegisterAllocationData: Remove unused headers
Gets rid of unnecessary header dependencies.
2025-09-03 10:31:41 -04:00
LC 06c5135d4e Merge pull request #4815 from lioncash/abi
Dispatcher: Reduce header dependencies
2025-09-03 09:45:36 -04:00
LC f66066cbd8 Merge pull request #4819 from lioncash/constprop
ConstProp: Remove file
2025-09-03 09:45:01 -04:00
Lioncache a144622a06 VixlUtils: Remove IsPowerOf2()
std::has_single_bit already does this.
2025-09-03 05:46:29 -04:00
Lioncache 05c814b8c1 VixlUtils: Use std::countl_zero where applicable
Makes for a little less code.
2025-09-03 05:45:02 -04:00
Ryan Houdek e0af51fb06 unittests/ASM: Classify unittests with new flags
Necessary for running on old hardware.
2025-09-02 18:37:59 -07:00
Ryan Houdek 7b749558df HostRunner: Add support for non-fsgsbase hardware 2025-09-02 18:37:59 -07:00
Ryan Houdek 1797c3c617 unittests/ASM: Add support for more feature flags 2025-09-02 18:15:31 -07:00
Billy Laws 28a1c28cb4 Frontend: Track the raw opcode and use for mono tailcall detection
This path was nonfunctional prior to this, as the jump is a group-opcode
which wouldn't be picked up (as Op is rewritten in NormalOp).
2025-09-02 22:42:02 +01:00
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Billy Laws 7c9c97f2b9 Update InstCountCI 2025-09-02 22:23:41 +01:00
Billy Laws d72e121fd4 AtomicOps: Reimplement AtomicNeg using 8.1 CAS atomics 2025-09-02 22:23:14 +01:00
Billy Laws 003fc3dab5 Update InstCountCI 2025-09-02 22:21:56 +01:00
Billy Laws 2c883d7cdc OpDispatcher: Force a recheck of TF after POPF
The dispatcher/block linker will handle this, but if the instruction
following a POPF flag doesn't otherwise trigger one of those the
interrupt would be missed.
2025-09-02 22:21:42 +01:00
Billy Laws fcba49768c Frontend: Force acq/rel semantics for known Unity ringbuffer offsets
Unity games crash with TSO disabled due to the SPSC GfxDevice
ThreadedStreamBuffer read/write pointer updates missing acq/rel
semantics. Rather than attempting to use heuristics to match cases like
this, which could be overzealous and hit more accesses than necessary, just
target the problem directly as this is consist across 32/64 bit Unity versions
for at least the past 10 years. Gate this behind the existing Unity mono
hacks to avoid false-positives in non-Unity games.
2025-09-02 22:05:22 +01:00
Billy Laws 959af9c3af Frontend: Always explore fallthrough branches with multiblock 2025-09-02 22:02:53 +01:00
Lioncache 1edacf05e9 ConstProp: Remove file
This was left in the tree, but the functionality was removed a while ago, so we can delete this.
2025-09-02 14:50:22 -04:00
Lioncache 1f756fb386 General: Replace usages of .at(0) with .data() where applicable
This is generally more straightforward in cases where bounds checking isn't necessary.
2025-09-02 13:03:37 -04:00
Lioncache 5ebdc01783 F80Fallbacks: Remove unnecessary memset
X80SoftFloat instances already initialize zeroed out in the default constructor.
2025-09-02 12:25:03 -04:00
Lioncache aa29201e00 Dispatcher: Reduce header dependencies
Gets rid of some unnecessary dependencies and also resolves some indirect dependencies.
2025-09-02 12:10:05 -04:00
LC b98acdf97a Merge pull request #4814 from neobrain/refactor_code_cache_remove_legacy
CodeCache: Remove legacy interfaces
2025-09-02 07:41:57 -04:00
LC 7aeb4788ed Merge pull request #4813 from Sonicadvance1/movsxd_test
unittests: Adds missing movsxd edge case
2025-09-02 07:40:24 -04:00
Tony Wasserka 7b1db40d4a Config: Remove legacy code caching interfaces 2025-09-02 12:06:57 +02:00
Tony Wasserka dcaa90a855 CodeCache: Remove legacy interfaces 2025-09-02 12:06:52 +02:00
Ryan Houdek 0a74efe3f8 unittests: Adds missing movsxd edge case
nasm doesn't let you encode all of these but we had already supported
them in FEX.

It's a bit of a weird edge case behaviour, since it's supposed to be
sign extended a 32-bit value in to a 64-bit register but with prefixes
(or lack of) it can be a move with zero-extend or a 16-bit insert.

Make sure we're actually testing these edge cases.
2025-09-02 01:16:42 -07:00
LC 5a73134c7c Merge pull request #4812 from Sonicadvance1/disable_gvisor
gvisor_tests: Update expectations for kernel 6.16
2025-09-02 04:15:37 -04:00
Ryan Houdek 43d535cf8d gvisor_tests: Update expectations for kernel 6.16 2025-09-01 16:13:42 -07:00
Ryan Houdek 4a71edf7bc Merge pull request #4811 from pmatos/fix/SPEC_GCC-C-execute-ieee-fp-cmp-8l
Fix IEEE 754 unordered comparison detection in x87 floating-point operations
2025-09-01 14:39:59 -07:00
Ryan Houdek ae90040ed2 Merge pull request #4810 from lioncash/file
FileManagement: Minor interface cleanup
2025-09-01 14:39:48 -07:00
Ryan Houdek fbf982f001 Merge pull request #4809 from lioncash/ir
IR: Remove some unimplemented prototypes
2025-09-01 14:39:39 -07:00
Ryan Houdek 839e3ac775 Merge pull request #4808 from lioncash/sig
Dispatcher: Centralize SignalDelegator config creation
2025-09-01 14:39:30 -07:00
Paulo Matos 65d230f32b instcountci: Fix IEEE 754 unordered comparison detection in x87 floating-point operations 2025-09-01 21:43:23 +02:00
Paulo Matos 0882b0db0c asm_tests: Fix IEEE 754 unordered comparison detection in x87 floating-point operations 2025-09-01 21:42:01 +02:00
Paulo Matos 3d92c1fb32 Fix IEEE 754 unordered comparison detection in x87 floating-point operations
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.

Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
2025-09-01 18:26:48 +02:00
Lioncache d4e679432c FileManagement: Add missing includes
Gets rid of a few indirect dependencies.
2025-08-29 13:42:17 -04:00
Lioncache 54c8cb9fd7 FileManagement: Mark helper functions as const where applicable
These don't modify internal class state, so we can mark them as such.
2025-08-29 13:41:58 -04:00
Lioncache 221ee1a122 IREmitter: Remove unimplemented prototypes
Get/SetPackedRFLAG() functions were moved to the opcode dispatcher, but these were
accidentally left behind.
2025-08-29 12:00:53 -04:00
Lioncache 2b29f90c92 IREmitter: Remove unused <array> header
Just one less include to worry about
2025-08-29 11:57:03 -04:00
Lioncache d7ec0e570f IR: Make IsFragmentExit() internally linked
This isn't used outside of the IR emitter. Plus, IsBlockExit() is already a more general interface to use,
since it handles the fragment exit case as well.
2025-08-29 11:51:49 -04:00
Lioncache 86c2144943 Core: Simplify define in InitCore()
Instead of inverting off a separate define, we can combine both checks together.
2025-08-29 11:30:48 -04:00
Lioncache ee85230db2 Dispatcher: Remove unused IntCallbackReturnAddress member
Turns out that this isn't used anymore. Found due to an unused field warning
when moving all public members into the private section.
2025-08-29 11:19:10 -04:00
Lioncache fde99a8dc1 Dispatcher: Centralize SignalDelegator config creation
Since all of the information comes from the Dispatcher, we can have the dispatcher
provide that information. This way we can also eliminate a bunch of now-redundant
public interface members and simplify the config setup within InitCoreImpl().

Conveniently, this also allows making all members of the dispatcher non-public.
2025-08-29 11:19:07 -04:00
Lioncache 2e692a1146 SignalDelegator: Move struct outside of SignalDelegator
The name itself is already qualified with SignalDelegator, so this can reasonably be outside the class itself.
This also allows for forward declarations of the config struct
2025-08-29 10:23:03 -04:00
Lioncache 73f36fb5b3 SignalDelegator: Move includes into cpp file
Reduces header dependencies
2025-08-29 10:07:56 -04:00
Ryan Houdek f866ad52dc Merge pull request #4807 from Sonicadvance1/remove_visibility
Remove visibility on functions that is no longer necessary
2025-08-28 18:34:37 -07:00
LC 9566910574 Merge pull request #4806 from Sonicadvance1/avx_constexpr
X86Tables: Convert AVX tables to constexpr
2025-08-27 19:18:41 -04:00
Ryan Houdek 4188d4d10b Add hidden visibility to FEXLoader/LinuxEmulation/Common 2025-08-27 16:14:59 -07:00
Ryan Houdek d75152cb74 IR: Remove visibility on functions that is no longer necessary
These used to be required because we leaked IR semantics across the API
boundary. Since we no longer do, we can remove the visibility from
these.
2025-08-27 15:34:57 -07:00
Ryan Houdek add54b8089 X86Tables: Convert AVX tables to constexpr
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
2025-08-27 14:55:10 -07:00
LC 7431495f13 Merge pull request #4797 from Sonicadvance1/more_brk_fixes
More BRK fixes
2025-08-27 16:49:55 -04:00
Ryan Houdek 5ed5566106 X86Tables/AVX128: Stop installing state functions a second time 2025-08-27 13:09:53 -07:00
Ryan Houdek 2d9a4412a4 X86Tables: Have AVX128 table always have PCLMUL 2025-08-27 13:08:27 -07:00
Ryan Houdek 982abd5719 X86Tables: Have AVX256 table always have PCLMUL 2025-08-27 13:06:43 -07:00
Ryan Houdek f09db0cbe5 Merge pull request #4805 from Sonicadvance1/hammer_hammer
FEXCore: Move most instruction tables to be constexpr
2025-08-27 12:57:21 -07:00
Ryan Houdek 80cacc462e X86Tables: Move x87 tables to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 1a2b2f8870 X86Tables: Move Second ModRM table to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek e8c576047f X86Tables: Moves Base ops to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 995d3152eb X86Tables: Convert Secondary tables to constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek ddc40cc273 X86Tables: Converts PrimaryGroup table to constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 5ee311b831 X86Tables: Converts H0F3A table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek ea90296e22 X86Tables: Converts H0F38 table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek 58c1a8d148 X86Tables: Convert DDDNow table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek 87a19c7938 FEXCore: Move SecondaryGroupTables Arch specific ops to a table
No runtime-installation necessary.
2025-08-27 12:24:24 -07:00
Ryan Houdek 5b0709de2f Syscalls: Removes an old workaround for MAP_FIXED_NOREPLACE
We don't support these old kernels, so remove this workaround.
2025-08-27 11:52:10 -07:00
Ryan Houdek 87db37c65e Syscalls: Allocate brk granules at page size
This matches kernel behaviour for how brk works. I initially implemented
this as a way to speed up some early applications that relied heavily on
brk. We're long past that time and we should match Linux behaviour
instead.

This fixes issues where applications (FEXLoader really) can see we've
allocated a larger granule and breaks self-hosting. There's no reason to
be smarter around this, if applications want more perf then they'll
switch to a better allocator than brk.

Also need to make sure that after the brk region has been reserved, to
unmap its initial mapping to allow future mmap syscalls to overwrite it.
We need to reserve it early to ensure the region initially exists and we
don't accidentally map other things in that space. Then it is up to the
guest if they don't want to overwrite it.
2025-08-27 11:51:52 -07:00
Ryan Houdek 100db13f99 Syscalls: Make sure to munmap on full brk zero
Failed to munmap in this case.
2025-08-27 11:51:52 -07:00
Ryan Houdek d2061b27a2 Syscalls: Rename DataSpaceMaxSize
Implied the maximum size that brk could grow to. Which isn't correct, it
is the current maximum size mapped which is page size, versus the
current brk offset which is byte ranged.
2025-08-27 11:51:52 -07:00
Ryan Houdek 002dfb9b84 Syscalls: Remove unused DataSpaceStartingSize 2025-08-27 11:51:52 -07:00
Ryan Houdek 0ab1241f09 X86Tables: Switch OpDispatch pointer over to a union
A wild use case of union over a variant because we don't want to
increase the encoding size from 128-bit to 256-bit (because of padding).
The type of operation is encoded with the table operation type, so a
variant is unnecessary and we get to keep the 128-bit encoding.

This allows us to have "recursive" x86 table descriptions. But in
reality this is going to only be one layer deep. As this will allow the
Frontend decoder to select instruction encodings based on arch bitness
once the tables are generated correctly.
2025-08-26 15:36:37 -07:00
Ryan Houdek aa63656ff1 X86Tables: Removes nullptr on DispatchPtr
Let it value initialize to zero which will be necessary when it switches
to a union.
2025-08-26 15:36:33 -07:00
LC 6ba32da73a Merge pull request #4804 from Sonicadvance1/portable_tests
Github: Force FEX portable config
2025-08-26 12:52:22 -04:00
Ryan Houdek 1129e71069 Github: Force FEX portable config
Noticed some runners had generated a .fex-emu/Config.json, run the tests
in portable mode to ensure they don't pick up a spurious config option.

FEXLinuxTests still ran non-portable because it's required for the thunk
tests.
2025-08-25 13:32:51 -07:00
Ryan Houdek 3b32fd5c49 Merge pull request #4794 from neobrain/fix_reformat_script
Scripts: Fix use of removed helper script
2025-08-25 12:20:33 -07:00
Ryan Houdek 7df1ed271f Merge pull request #4799 from neobrain/refactor_ranges
VolatileMetadata: Clean up implementation using ranges::views::split
2025-08-25 11:45:44 -07:00
Ryan Houdek 67e9b40bab Merge pull request #4803 from neobrain/refactor_irdumper_cleanup
IRDumper: Clean up formatting using fmt
2025-08-25 11:43:36 -07:00
Ryan Houdek 62b80903f0 Merge pull request #4801 from neobrain/refactor_maybe_unused
Reduce use of maybe_unused attributes
2025-08-25 11:43:09 -07:00
Ryan Houdek 993d917478 Merge pull request #4800 from neobrain/fix_assert_logargs
LogManager: Evaluate non-predicate arguments for assertions lazily
2025-08-25 11:41:05 -07:00
Ryan Houdek ac20880a14 Merge pull request #4802 from neobrain/fix_fedora_rootfs
FEXConfig: Fix RootFS override in muvm-based setups
2025-08-25 11:40:26 -07:00
Tony Wasserka 99cfe05ee5 IRDumper: Clean up formatting using fmt 2025-08-25 17:05:10 +02:00
Tony Wasserka 7c187e82b7 FEXConfig: Don't set an empty RootFS if none is selected 2025-08-25 16:30:58 +02:00
Tony Wasserka 2a760660a9 Merge pull request #4795 from pmatos/formatCI
Do not use diff_from_common_commit
2025-08-25 14:03:24 +02:00
Tony Wasserka 4b11826a7d LogManager: Evaluate non-predicate arguments for assertions lazily
This allows assertions like "some_iterator == end()" to dereference the
unexpected iterator value when formatting the log message on failure.
2025-08-25 10:54:39 +02:00
Tony Wasserka 887f586874 IRDumper: Remove unneeded maybe_unused attribute 2025-08-25 10:36:01 +02:00
Tony Wasserka db94c04179 FEXCore: Use inline functions instead of maybe_unused static ones 2025-08-25 10:36:01 +02:00
Tony Wasserka 51e64c69f3 LogManager: Unconditionally evaluate assertion conditions
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
2025-08-25 10:36:01 +02:00
Tony Wasserka 97c46843d2 LogManager: Evaluate non-predicate arguments for assertions lazily
This allows assertions like "some_iterator == end()" to dereference the
unexpected iterator value when formatting the log message on failure.
2025-08-25 10:21:14 +02:00
Tony Wasserka 16310c91c9 VolatileMetadata: Clean up implementation using ranges::views::split 2025-08-25 09:29:27 +02:00
Tony Wasserka 220e421c22 External: Add range-v3 submodule 2025-08-25 08:47:56 +02:00
LC a156752fae Merge pull request #4796 from Sonicadvance1/brk_fixes
ELFCodeLoader: Fixes BRK allocation with non-PT_DYN ELF files
2025-08-23 20:03:54 -04:00
Ryan Houdek 5d1c574dc5 ELFCodeLoader: Fixes BRK allocation non PT_DYN ELF files
When loading an ELF file (usually PT_EXEC), we would have BSS regions
that ended up in the high pages of the address space. We would then
attempt mapping BRK after whatever the highest address ending up being.

This was problematic because we used `MAP_FIXED` which means that the
BRK region could cross in to 64-bit address space, overwrite VDSO,
overwrite the stack, vsyscall, maybe a couple of other things.

Instead of letting that happen, reorder some of the logic so that in
`LoadElfFile` will ensure the full `BRK_SIZE` is allocated (or return
zero) and then we can ensure we're never overwriting other various
memory regions.

This doesn't fix two fundamental issues that FEX has with brk:
* If BRK gets mapped below the ELF, that should mean it isn't mapped at all
  * Linux-isms, doubt this breaks new software
* We still map the full 8MB, which isn't correct
  * I'm working on fixing this issue, which is why I encountered this.
2025-08-22 20:30:04 -07:00
Ryan Houdek 6310c20217 Merge pull request #4773 from Sonicadvance1/extended_volatile_metadata
Implement support for additional provided volatile metadata
2025-08-22 15:15:50 -07:00
Ryan Houdek 0dfc30fdd8 APITests: Adds extendedvolatile 2025-08-22 15:03:56 -07:00
Ryan Houdek f050571696 Implement support for additional provided volatile metadata
Only wired up for wow64 and arm64ec. Gives more granular control over
TSO enabling and disabling. Matches arm64ec volatile metadata except
with one more additional feature that whole modules can be disabled at a
time.

Once we know the mapped size of files in Linux then we'll be able to do
the same thing there, but there's not a full mechanism wired up for that
yet.
2025-08-22 15:03:56 -07:00
Ryan Houdek 6db5bc3721 FEXCore/Utils/IntervalList: Adds a couple functions 2025-08-22 15:03:56 -07:00
LC 4f84c281ef Merge pull request #4791 from bylaws/norace
ARM64EC: Cooperatively suspend if needed when leaving the JIT
2025-08-22 12:21:09 -04:00
LC 0ca438e081 Merge pull request #4790 from bylaws/nomorewine
WinAPI: Force WOW64 allocations to start >4GB
2025-08-22 12:17:28 -04:00
LC 2472aee473 Merge pull request #4786 from Sonicadvance1/noise_fcw
Dispatcher: Most minor of optimizations for f64 x87
2025-08-22 12:16:43 -04:00
LC 714856f74c Merge pull request #4788 from Sonicadvance1/modify_ldt_x64
LinuxSyscalls: Support modify_ldt on x64
2025-08-22 12:15:17 -04:00
Paulo Matos e98932df92 Do not use diff_from_common_commit 2025-08-22 15:10:17 +02:00
Tony Wasserka 95623dac63 Scripts: Fix use of removed helper script
This was accidentally changed back to clang-format.py in
3b322d8f37.
2025-08-22 11:05:01 +02:00
Billy Laws 827015a8e3 ARM64EC: Cooperatively suspend if needed when leaving the JIT 2025-08-21 17:24:56 +01:00
Billy Laws a379818386 WinAPI: Force WOW64 allocations to start >4GB
Avoids stealing address space from the guest.
2025-08-21 17:19:46 +01:00
Ryan Houdek 4516e4e9f6 LinuxSyscalls: Support modify_ldt on x64
This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.

In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.

This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.

With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
2025-08-20 13:22:56 -07:00
Ryan Houdek 83e44961fc Merge pull request #4784 from Sonicadvance1/more_ldt_gdt
FEXCore: Move LDT/GDT management to the frontend
2025-08-20 13:22:08 -07:00
Ryan Houdek bc2e959897 Merge pull request #4742 from pmatos/fninit_IE
Clear IE flag on fninit
2025-08-19 12:57:25 -07:00
Ryan Houdek 1b0b4ba416 InstCountCI: Add frontend handling of GDT 2025-08-18 14:03:51 -07:00
Ryan Houdek 25a0752e58 FEXCore: Move LDT/GDT management to the frontend
This will allow the frontend to manage LDT and GDT so that Linux will be
able to install 32-bit code segments.
2025-08-18 14:03:51 -07:00
Paulo Matos b0a7e0bbcc instcountci: Clear IE flag on fninit 2025-08-18 15:49:10 +02:00
Paulo Matos af561abc71 asm_tests: Clear IE flag on fninit 2025-08-18 15:40:46 +02:00
Paulo Matos 3bf323aa2b Clear IE flag on fninit 2025-08-18 15:40:46 +02:00
Ryan Houdek c9a33c638a Dispatcher: Most minor of optimizations for f64 x87
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
2025-08-17 19:50:21 -07:00
Ryan Houdek 8d1042859d CoreState: Use a named constant for LDT/GDT
Just to make this easier, NFC.
2025-08-15 19:09:33 -07:00
LC 9072571fc1 Merge pull request #4783 from Sonicadvance1/save_restore_segments
LinuxSyscalls: Save and restore segment registers on signal
2025-08-14 15:34:10 -04:00
Ryan Houdek d8748b8216 LinuxSyscalls: Save and restore segment registers on signal
Effectively NFC today while we don't allow installing LDT entries, but
is required to have signal handlers execute in the correct bitness if
32-bit code segments are enabled.

Nothing too crazy, just that four segments are stuffed in to the
`REG_CSGSFS` data value, and then CS/SS gets restored on restore.
FS and GS are ignored in this instance because edge case behaviour where
those don't get reset on signal entry, so if they get changed then it is
on the userspace application to fix it.

Gets another change out of my stashes.
2025-08-14 12:22:28 -07:00
Ryan Houdek 2f930b201a Merge pull request #4782 from bylaws/oo0
OpcodeDispatcher: Simplify CALLOp return addr calc
2025-08-14 11:23:21 -07:00
Ryan Houdek 26b4195f1f Merge pull request #4779 from asLody/environ
Windows: Support EnvLoader via _environ
2025-08-14 11:22:23 -07:00
Billy Laws 5c078d13c0 Update InstCountCI 2025-08-14 15:35:23 +01:00
LC 2825ac282a Merge pull request #4781 from Sonicadvance1/implement_call_ret_far
OpcodeDispatcher: Implement support for CALLF/RETF
2025-08-14 00:15:31 -04:00
Lody cd31a79f41 Windows: Support EnvLoader via _environ 2025-08-14 10:41:11 +08:00
Ryan Houdek 52eb9f46d2 InstcountCI: Add callf/retf 2025-08-13 16:33:23 -07:00
Billy Laws 08c430dcfb OpcodeDispatcher: Simplify CALLOp return addr calc 2025-08-14 00:25:23 +01:00
Ryan Houdek b486aeeac5 unittests/ASM: Unittests for callf/retf
Not too crazy, just make sure to check RSP locations on both sides.
2025-08-13 16:20:18 -07:00
Ryan Houdek 13da6102b3 OpcodeDispatcher: Implement support for CALLF/RETF
Similar to the previous far jmp, if the CS changes operating mode then
things will still explode with other FEX asserts. But this gets another
change out of my stashes.

These instructions go hand-in-hand obviously so they get implemented as
a pair.
2025-08-13 16:18:12 -07:00
LC 1e1c4e017e Merge pull request #4780 from Sonicadvance1/implement_jump_far
OpcodeDispatcher: Implement support for far jump
2025-08-13 15:58:57 -04:00
Ryan Houdek ceaaab5a86 unittests/ASM: Test jmpf without cs change 2025-08-13 12:26:34 -07:00
Ryan Houdek 2b7d03d4f9 OpcodeDispatcher: Implement support for far jump
If anything actually attempts to use this to change CS then it'll very
quickly hit some other asserts, but it gets one more change off my list.
2025-08-13 12:19:16 -07:00
LC c8810557c1 Merge pull request #4777 from Sonicadvance1/move_gdt_up
FEXCore: Split GDT and LDT prep work
2025-08-13 12:06:41 -04:00
Ryan Houdek 2a4bfe49f5 Merge pull request #4743 from pmatos/fclex
Implementation of FCNLEX to clear IE flag
2025-08-12 12:35:15 -07:00
Ryan Houdek b71879e90e Merge pull request #4776 from bylaws/correct
Windows: Misc race and correctness fixes
2025-08-12 12:07:53 -07:00
Billy Laws 63ae262752 Merge pull request #4778 from FEX-Emu/fix_3dnow_doc
HostFeatures: Fix insufficient documentation to reproduce 3DNow! bug
2025-08-12 15:06:36 +01:00
Tony Wasserka 2c3a4fef26 HostFeatures: Fix insufficient documentation to reproduce 3DNow! bug
This ensures we don't forget about it if bug diagnosis leads nowhere.
2025-08-12 15:55:45 +02:00
Ryan Houdek e696e32b95 InstcountCI: Update 2025-08-11 20:48:47 -07:00
Ryan Houdek 3f9df49e3e FEXCore: Split GDT and LDT prep work
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.

Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
2025-08-11 20:45:51 -07:00
Billy Laws 1ac3d3c5a6 Windows: Extend user handle access masks in callbacks where necessary
While most applications will just create ALL_ACCESS handles which
support the additional operations FEX needs over the regular windows syscall
(usually just QUERY_INFORMATION to lookup the thread object), it is
valid for them to pass in the minimal set required for each operation
and windows/wine will reject any other operations (e.g. querying the TEB
base on a handle with only the SUSPEND right). Windows permits
duplicating handles to extend their permissions given the process itself
has such permissions, so do that as necessary for the operations FEX
needs.
2025-08-12 03:19:20 +01:00
LC b3d1aade71 Merge pull request #4775 from Sonicadvance1/fix_32bit_cmpxchg
FEXCore: Fixes 32-bit zero-extend semantics on cmpxchg
2025-08-11 21:53:14 -04:00
Ryan Houdek c300c723f7 InstcountCI: Update 2025-08-11 17:50:18 -07:00
Ryan Houdek d0ec069dcb FEXCore: Fixes 32-bit zero-extend semantics on cmpxchg
We were missing the semantics when the result was a success.
2025-08-11 17:49:17 -07:00
LC 6811406b24 Merge pull request #4774 from Sonicadvance1/fix_o16_enterleave
FEXCore: Fixes 16-bit operand size Enter and Leave instructions
2025-08-11 20:15:44 -04:00
Ryan Houdek 0879994aff FEXCore: Fixes 16-bit operand size Leave
Weren't testing this edge case just like enter
2025-08-11 17:01:49 -07:00
Billy Laws b2b5ccf69c Windows: Validate any user handle permissions for syscall callbacks
In the terminate case, if we don't return early then we will delete the
FEX thread object but the actual terminate call that follows will
fail, leaving the thread in a bad state.
2025-08-12 00:51:00 +01:00
Billy Laws 6f5d4fbf30 Windows: Avoid potential deadlocks is a thread is terminated while in the JIT
Try to suspend the thread before termination to ensure no important JIT locks
are held that could be left locked if the thread is compiling code etc.
2025-08-12 00:51:00 +01:00
Billy Laws 8ad1ea499d Windows: Avoid a double-free if a thread is terminated twice
Unlikely to show up in any real code, but if an attempt is made to terminate
a terminated thread we should handle this gracefully.
2025-08-12 00:51:00 +01:00
Billy Laws 4df17abea7 Windows: Avoid potential race between thread init and termination
With the prior locking, we could try to terminate a thread while it
was initializing and perhaps holding global locks leading to deadlock.
2025-08-12 00:51:00 +01:00
Billy Laws 972c8a97dc ARM64EC: Disable syscall callbacks when handling OvercommitTracker AVs
We don't need to be informed of our own mappings, avoids a deadlock that
occurs if the ThreadCreationMutex lifetime is extended when creating a
thread.
2025-08-12 00:50:51 +01:00
Ryan Houdek 135d17fbd3 FEXCore: Fixes 16-bit operand size Enter
We weren't handling this edge case. Noticed it while looking at some
other test failures.
2025-08-11 16:50:30 -07:00
LC 095b2a8926 Merge pull request #4772 from Sonicadvance1/fix_3dnow_test
unittests/3DNow: Fix femms test
2025-08-09 17:19:17 -04:00
Alyssa Anne Rosenzweig 297677c2c6 Merge pull request #4771 from alyssarosenzweig/opts
Local JIT translation time wins
2025-08-09 09:44:23 -04:00
Alyssa Rosenzweig 2a00f23459 JIT: replace JumpTargets map with vector
Hashmaps are super expensive and there's no reason not to use a vector - we
already have compact block IDs so we don't benefit from the sparseness. Huge win
for very little effort.

Spotted when profiling FEX. CondJump() in the JIT was almost 4% of our time (?!)
and all because of map slowness. Easy fix.

Difference at 95.0% confidence
	-0.0196494 +/- 0.00194956
	-3.92827% +/- 0.389753%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Alyssa Rosenzweig 68ad916a91 Addressing: remove closures from SelectAddressMode
this is hot and closures make the assembly harder to reason about.

No difference at 95%.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Alyssa Rosenzweig 0f3883152b RegisterAllocationPass: call GetRAArgs less
..and be consistent about signedness.

clang doesn't seem able to do this itself.

Difference at 95.0% confidence, n=100
	-0.272326% +/- 0.242353%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Ryan Houdek 60baee4020 unittests/3DNow: Fix femms test
We are onyl wanting to test the FTW, which is 16-bit but we were reading
a whole 32-bits. The upper 16-bits was ending up in the reserved
section, which is undefined as to what it contains. In FEX it'll contain
0 but on a Zen 4 CPU it contains 0xFFFF.

Only read the bits we care about.
2025-08-08 12:08:52 -07:00
Ryan Houdek 5a7ff4780f Merge pull request #4770 from bylaws/arm64ec
InvalidationTracker: Disable mono hacks when multiblock is disabled
2025-08-07 13:59:33 -07:00
Billy Laws dd2845401c InvalidationTracker: Disable mono hacks when multiblock is disabled 2025-08-07 21:27:04 +01:00
Ryan Houdek a47c947065 Merge pull request #4768 from bylaws/memover
Avoid OOB reads for mem sources in some float conversion ops
2025-08-07 09:55:20 -07:00
Billy Laws 203ac51681 Update InstCountCI 2025-08-07 14:36:12 +01:00
Billy Laws a7986993bf InstCountCI: Add tests for AVX float conversion ops with mem src 2025-08-07 14:36:12 +01:00
Billy Laws ce59d40570 Vector: Avoid OOB reads in (v)cvt(t)s{s,d}2s{s,d} with mem src 2025-08-07 14:36:12 +01:00
Billy Laws 19f33e22d7 AVX128: Avoid OOB reads in vcvtsi2s{s,d} with mem src 2025-08-07 14:36:12 +01:00
Ryan Houdek 7d818e18be Merge pull request #4767 from bylaws/steamy
Windows: Ignore MEM_RESET(_UNDO) allocation protection notifications
2025-08-06 18:33:35 -07:00
Billy Laws 52af0aceea Windows: Ignore MEM_RESET(_UNDO) allocation protection notifications
When these are set the allocation permissions aren't actually updated.
2025-08-07 02:05:35 +01:00
Billy Laws cd1cf6f615 ARM64EC: Drop redundant ThreadCreationMutex locks 2025-08-07 02:04:48 +01:00
Ryan Houdek 2f925aeb22 Merge pull request #4764 from alyssarosenzweig/opt/drop-constprop-2
Drop ConstProp
2025-08-06 18:03:06 -07:00
Alyssa Rosenzweig f99b42da5e InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:42:36 -04:00
Alyssa Rosenzweig 40beef061f ConstProp: drop pass
now obsolete!

Results for the whole series are excellent:

Difference at 95.0% confidence
	-3.97603% +/- 0.254656%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f0a434c167 IR: inline as we go
Augment IREmitter to inline constants as we generate code, rather than
needing a later clean up pass. This replaces the last function of ConstProp.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 92b66dbf17 IR: describe immediate inlining in json
This drops some register class validation since InlineConstants don't have a
register class.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f336261d6b IR: introduce & use more add helpers
so we can stash all the opts.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig c6324b83ea IR: introduce & use constant add helpers
more ergonomic and gives us a place to stash more logic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 54873992ab OpcodeDispatcher: optimize storecontext(0) in dispatcher
It's actually slightly less code to do it here than ConstProp, lol.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 38d568a807 OpcodeDispatcher: rework SelectCC
inline as we go, while trimming the internal interfaces.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig dcef1212c6 OpcodeDispatcher: add and use NZCVSelect01 helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 2030b70ff8 OpcodeDispatcher: do Select 0/1 inline in dispatcher
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 41a188eaec OpcodeDispatcher: inline entrypoint offsets manually
it's literally less code with some helpers.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f3ddaf455c OpcodeDispatcher: drop dead constructor
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Ryan Houdek aa979e4bc7 Merge pull request #4728 from bylaws/monohacks
Implement mono specific hacks
2025-08-06 15:51:35 -07:00
Ryan Houdek baddc56f01 Merge pull request #4766 from bylaws/3dnow
Disable 3DNow by default on WOW64 FEX
2025-08-06 15:30:27 -07:00
Billy Laws 9836e43d10 Update InstCountCI 2025-08-06 22:39:44 +01:00
Billy Laws c8cecd9f46 Frontend: Adjust forward branch distance limit
Now executable page tracking is implemented, this limit is technically
unnecessary and effectively never hit in normal code. Keep a reasonable
limit however to avoid accidentally inlining tail calls and exploring
dead branches in obfuscated code.
2025-08-06 22:39:17 +01:00
Billy Laws 38c59a44de Frontend: Disable multiblock across calls in mono JIT code 2025-08-06 22:39:17 +01:00
Billy Laws fddc86c2b1 Windows: Opportunistically detect and replace the mono backpatcher 2025-08-06 22:39:17 +01:00
Billy Laws 0dddc9b5cb InvalidationTrack: Support disabling SMC detection 2025-08-06 22:39:17 +01:00
Billy Laws 8b14bd4e87 FEXCore: Fix broken RemoveCustomIREntrypoint
Unused, but would have crashed prior due to providing a nullptr thread.
2025-08-06 22:39:17 +01:00
Billy Laws 5c77969e83 Frontend: Force full SMC detection for mono jump thunk callsites
See IsBranchMonoTailcall
2025-08-06 22:39:17 +01:00
Billy Laws e049596252 FEXCore: Use MonoBackpatcherWrite for XCHG ops in the mono backpatcher block 2025-08-06 22:39:17 +01:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Billy Laws da3b7f5a41 FEXCore: Add an option to enable future mono-specific hacks 2025-08-06 22:39:17 +01:00
Billy Laws 100f61ee24 FEXCore: Allow the frontend to force full SMC detection for a block 2025-08-06 22:39:17 +01:00
Billy Laws dcd2794ff5 FEXCore: Make ThreadRemoveCodeEntryFromJit invalidate over all threads
Avoids redundant CompileBlock hits and generally easier to reason about.
2025-08-06 22:39:17 +01:00
Billy Laws 8d1d3fe12b Windows: Implement SyscallHandler code range invalidation 2025-08-06 22:39:17 +01:00
Billy Laws 87ae0f058a LinuxEmulation: Implement SyscallHandler code range invalidation 2025-08-06 22:39:17 +01:00
Billy Laws 7e1be6625f SyscallHandler: Add a method for invalidating a guest code range 2025-08-06 22:39:17 +01:00
Billy Laws 0c6fcd1678 FEXCore: Add function to recover the current block entrypoint 2025-08-06 22:39:17 +01:00
Billy Laws 37ac46273d Windows: Print extra debug info on image map 2025-08-06 22:39:17 +01:00
Billy Laws 1e21416ccb Disable 3DNow by default on WOW64 FEX 2025-08-06 22:30:38 +01:00
Ryan Houdek e2353959c1 Merge pull request #4755 from ChanthMiao/fix/wine32-v4l2
LinuxEmulation: add basic supports for v4l2 driver of wine32.
2025-08-05 21:36:13 -07:00
LC 61d3d5b068 Merge pull request #4763 from Sonicadvance1/fix_hostrunner_bug
HostRunner: Fixes bug with saving CSFSGS
2025-08-05 23:23:53 -04:00
Ryan Houdek 572e5ed9c6 HostRunner: Fixes bug with saving CSFSGS
We were failing to save the entire 64-bit register, cutting off the top
32-bits which contain SS and CS. This ...happened to not cause issues
but is a bug that I encountered with far jump handling.
2025-08-05 19:45:53 -07:00
Ryan Houdek 4dbfa2febb Merge pull request #4741 from pmatos/fninit_tests
asm_tests: Tests for fninit
2025-08-05 19:29:58 -07:00
Changwei Miao 0293d2027d LinuxEmulation: add basic supports for v4l2 driver of wine32.
This patch does cover up all v4l2 syscalls in i386. It just works
well with the poor v4l2 driver of wine32.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-08-06 04:49:49 +08:00
LC c28b94890c Merge pull request #4761 from neobrain/fix_debug_strace_build
Syscalls: Fix DEBUG_STRACE build
2025-08-05 11:17:19 -04:00
LC 63a910d2fc Merge pull request #4760 from Sonicadvance1/vex_log
Frontend: Remove log about VEX map_select
2025-08-05 11:15:53 -04:00
Tony Wasserka fbef5c6a6a Syscalls: Fix DEBUG_STRACE build 2025-08-05 15:51:34 +02:00
Ryan Houdek 674b6e9f43 Frontend: Remove log about VEX map_select
During multiblock code discovery this fires a lot and it isn't
interesting. Just remove the log, it'll SIGILL correctly if it actually
hits.
2025-08-04 15:46:36 -07:00
Ryan Houdek 7897b6ad55 Merge pull request #4759 from alyssarosenzweig/ra/simplify
RegisterAllocationPass: simplify next-use logic
2025-08-04 14:59:13 -07:00
Alyssa Rosenzweig 6769b54e1c unittests: add blake3 test
this provokes RA spilling and hit an assertion fail on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 17:06:15 -04:00
Alyssa Rosenzweig 734ab4429f RegisterAllocationPass: fix SRA spilling corner
I hate this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 17:06:15 -04:00
Ryan Houdek 5db6a64f3f Merge pull request #4758 from alyssarosenzweig/opt/x87ftw-2
OpcodeDispatcher: optimize SetX87FTW
2025-08-04 13:14:57 -07:00
Alyssa Rosenzweig dcc458ea94 RegisterAllocationPass: simplify next-use logic
I doubt this will fix the regression but it might make it easier to identify.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:31:08 -04:00
Alyssa Rosenzweig 7354e7e3c9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:13:10 -04:00
Alyssa Rosenzweig 07ef765f7e OpcodeDispatcher: optimize SetX87FTW
In b8dd5d95b ("OpcodeDispatcher: optimize X87FTWTag"), we optimized
X87FTWTag using an efficient Morton interleave operation. Here, we do the
inverse, optimizing SetX87FTW using an efficient Morton deinterleave
operation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:13:10 -04:00
Alyssa Rosenzweig c58ad8b593 IR: add AndShift op
will use it for next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:09:20 -04:00
Alyssa Rosenzweig 31091c1053 RegisterAllocationPass: remove some indentation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 08:16:28 -04:00
Ryan Houdek 59659a7184 Docs: Update for release FEX-2508 2025-08-01 18:24:33 -07:00
Ryan Houdek e6de17e72e Merge pull request #4752 from alyssarosenzweig/opt/defer-next-use-analysis-2
RegisterAllocationPass: defer next-use analysis
2025-08-01 16:38:31 -07:00
Ryan Houdek 0457bdc7ce Merge pull request #4751 from lioncash/perm
ASIMDOps: Remove unused permute overloads
2025-08-01 12:45:50 -07:00
Ryan Houdek 7e54c2735c Merge pull request #4750 from lioncash/op6
ASIMDOps: Move remaining base opcodes into implementing function
2025-08-01 12:45:24 -07:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Alyssa Rosenzweig 6a03df7b8d RedundantFlagCalculationElimination: drop dead syscall/atomic opts
I don't think these are worth it, and also currently they don't trigger ever.

n=100:
Difference at 95.0% confidence
	-0.00245961 +/- 0.00139573
	-0.524468% +/- 0.297615%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:37 -04:00
Alyssa Rosenzweig 7ee5a8065c RegisterAllocationPass: defer next-use analysis
This is expensive and only needed for spilling, so only do it for spilling. This
complicates the RA a bit but speeds us up on average since most blocks
don't spill. Total results of this change (including the prep commits that
slowed things down temporarily):

Difference at 95.0% confidence
	-1.71952% +/- 0.455996%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 2f2353765f RegisterAllocationPass: use kill bits
this is a lot lighter weight than next uses for the same purpose.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 8b4c2f5093 RegisterAllocationPass: consider AnySpilled at start of iteration
if we spill for SRA, we don't need/want to execute this code path. this will be
load bearing by the end of this series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig d80662a9ba RegisterAllocationPass: ignore kill bit in SRA
needed for the backwards pass internally due to ordering. a little awkward but
shrug.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:11:53 -04:00
Alyssa Rosenzweig 68c5c72dbc IR: model kill bits
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:04:50 -04:00
Lioncache b460cbc91a ASIMDOps: Remove unused permute overloads
These aren't used at all and don't really provide anything that the existing
non-templated overloads can't.
2025-08-01 11:51:55 -04:00
Lioncache f7939078c2 ASIMDOps: Remove unnecessary ins() overload
There's no difference between this and the non-templated version, in fact,
this variant wasn't even used at all, so we can just remove it.
2025-08-01 11:24:44 -04:00
Lioncache e103af3b93 ASIMDOps: Move base opcode into ASIMDScalarCopy()
Now we have no more duplicated opcodes in ASIMDOps.
2025-08-01 11:13:58 -04:00
Lioncache 45bec06d5b ASIMDOps: Move base opcode into ASIMDFloatConvBetweenInt()
Deduplicates the second last remaining instruction category in ASIMDOps
2025-08-01 10:49:52 -04:00
Alyssa Rosenzweig c82efe7621 RegisterAllocationPass: set AnySpilled less
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 10:39:11 -04:00
Ryan Houdek ba85dbf526 Merge pull request #4749 from lioncash/op5
ASIMDOps: Move most remaining base opcodes into their implementing functions
2025-07-31 12:21:31 -07:00
Lioncache e624e87b29 ASIMDOps: Move base opcode into ASIMDExtract()
Minor deduplication of the base opcode.
2025-07-31 10:44:06 -04:00
Lioncache 6ff8d0953f ASIMDOps: Move base opcode into ASIMDPermute()
Gets rid of the need to respecify the base opcode in every instruction implementation
2025-07-31 10:43:17 -04:00
Lioncache 9f2fcdf0b3 ASIMDOps: Move base opcode into Crypto2RegSHA()/Crypto3RegSHA()
Removes the need to specify the base opcode multiple times.
2025-07-31 10:23:38 -04:00
Lioncache cecb9bfe32 ASIMDOps: Move base opcode into CryptoAES()
Minor deduplication.
2025-07-31 10:15:15 -04:00
Lioncache 6a989d844a ASIMDOps: Move base opcode into ASIMDTable()
Gets rid of duplication of the opcode in all instruction implementations.
2025-07-31 10:10:35 -04:00
Ryan Houdek 318b3115a8 Merge pull request #4748 from lioncash/op4
ASIMDOps: Constrain more instructions with IsQOrDRegister
2025-07-30 16:22:14 -07:00
Lioncache 017952e89e ASIMDOps: Constrain more instructions with IsQOrDRegister
We had quite a few instructions that we weren't constraining with this,
now the instruction itself will show up in the failure output more immediately
instead of failing on the internal implementation function if anything is
incorrectly passed through.
2025-07-30 19:03:33 -04:00
LC b0c74e1b66 Merge pull request #4747 from Sonicadvance1/cpuid
CPUID: Update documentation comments
2025-07-30 18:31:54 -04:00
Ryan Houdek a190086504 Merge pull request #4746 from lioncash/op3
ASIMDOps: Move base opcode into ASIMDModifiedImm()/ASIMDShiftByImm()
2025-07-30 15:30:37 -07:00
Ryan Houdek 84704d1cc2 CPUID: Update documentation comments
Additional reserved bits have set uses now.
Additionally set the cpuid bit for bus-lock-detect, because FEX
definitely detects bus-locks.
2025-07-30 15:13:08 -07:00
Lioncache b1ef8bbc5a ASIMDOps: Move base opcode into ASIMDModifiedImm() 2025-07-30 18:03:43 -04:00
Lioncache f19a6343dd ASIMDOps: Move base opcode into ASIMDShiftByImm()
Avoids needing to duplicate the opcode all over the place
2025-07-30 18:03:36 -04:00
Ryan Houdek 23d0d7d4a4 Merge pull request #4745 from lioncash/op2
ASIMDOps: Move base opcode into ASIMD2RegMisc()
2025-07-30 13:50:04 -07:00
Lioncache ed7bc28c21 ASIMDOps: Simplify size conditionals in several 2-reg misc instructions
In many of these, the long ternaries are equivalent to a subtraction by 1,
which is much more efficient
2025-07-30 16:31:33 -04:00
Lioncache 05dcb160bd ASIMDOps: Move base opcode into ASIMD2RegMisc()
Avoids the need to specify the opcode repeatedly, making implementations less noisy.
2025-07-30 16:25:24 -04:00
Ryan Houdek e49155ff90 Merge pull request #4744 from lioncash/op
ASIMDOps: Move base opcode into implementation function for some categories
2025-07-30 12:46:07 -07:00
Ryan Houdek 265525e472 Merge pull request #4740 from Sonicadvance1/multiple_segments
FEXCore/Frontend: Ensure multiple prefix bytes work
2025-07-30 12:43:51 -07:00
Ryan Houdek d4dcbfa90b FEXCore/Frontend: Ensure multiple prefix bytes work
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.

Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
2025-07-30 12:30:13 -07:00
Ryan Houdek 6724673482 Merge pull request #4739 from Sonicadvance1/remove_check
Frontend: Remove arbitrary check
2025-07-30 12:26:13 -07:00
Ryan Houdek d04f75df29 Frontend: Remove arbitrary check
REX prefix isn't even encoded in to the instruction tables if a 32-bit
process is running. Just remove this.
2025-07-30 11:53:15 -07:00
Ryan Houdek a5260f4233 Merge pull request #4738 from bylaws/peggle
Implement inline SMC handling for linux FEX
2025-07-30 11:48:28 -07:00
Ryan Houdek 369ca5cb72 Merge pull request #4719 from Sonicadvance1/runtime_mode_switch_take2
Runtime mode switch take 2
2025-07-30 11:47:54 -07:00
Lioncache a3e02cbafa ASIMDOps: Simplify conditionals in saddlv/uaddlv
Really all these size conversions are emulating is a subtraction by 1.
Also we can drop in an assert that was missed in uaddlv
2025-07-30 11:21:06 -04:00
Lioncache 4cca2f6915 ASIMDOps: Move base opcode into ASIMDAcrossLanes() 2025-07-30 11:07:08 -04:00
Lioncache 64ca48c4b1 ASIMDOps: Move base opcode into ASIMD3Different()
Deduplicates the open-coded base opcode in the instruction implementations.
2025-07-30 10:59:11 -04:00
Lioncache 74a897271f ASIMDOps: Move base opcode into ASIMD3Same()
Moves the base opcode into the actual implementation, so that we
aren't open-coding it into every relevant instruction function.
2025-07-30 10:40:17 -04:00
Paulo Matos c95f9f5e6a instcountci: Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:32 +02:00
Paulo Matos 93c3eb064d asm_tests: Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:21 +02:00
Paulo Matos b13162c5bf asm_tests: Division by zero does not set IE flag 2025-07-30 10:50:20 +02:00
Paulo Matos ee1725684a Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:16 +02:00
Tony Wasserka c6733a6eec Merge pull request #4731 from Sonicadvance1/armtifacts
github: Upload armtifacts
2025-07-30 10:22:49 +02:00
Tony Wasserka f7e99678b4 Merge pull request #4730 from Sonicadvance1/fix_thunk_functional_path
unittests: Fixes thunk unittest path
2025-07-30 10:20:46 +02:00
Tony Wasserka 976b68ffec Merge pull request #4713 from Sonicadvance1/free_the_stats
Profiler: Decouple profile stats from the profiler option
2025-07-30 10:19:07 +02:00
Paulo Matos a07b234659 asm_tests: Tests for fninit
NFC.
These tests existed for 32bits but were missing for 64bits and reduced precision.
Adjust 32bit test to match new tests.
2025-07-30 09:54:08 +02:00
Billy Laws 7b656fb009 SyscallsSMCTracking: Support inline SMC 2025-07-30 00:13:27 +01:00
Billy Laws 20331d52c3 LinuxSyscalls: Always reconstruct RIP and EFLAGS when spilling from the JIT 2025-07-30 00:13:27 +01:00
Billy Laws 2556acb82d TestHarnessRunner: Avoid frontend SMC handling 2025-07-30 00:13:27 +01:00
Ryan Houdek 248f0948b8 github: Upload armtifacts 2025-07-29 12:02:57 -07:00
Ryan Houdek 6c12db500e InstCountCI: Update for segment changes 2025-07-29 12:02:38 -07:00
Ryan Houdek 2a8f4dbbb6 Arm64EC: Update for GDT 2025-07-29 12:02:38 -07:00
Ryan Houdek 91598d5178 WOW64: Update for GDT 2025-07-29 12:02:37 -07:00
Ryan Houdek a6bb9739d4 OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2025-07-29 12:02:37 -07:00
Ryan Houdek aa871c797b FEXCore: Accurately store segment descriptors
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.

In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.

Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
2025-07-29 12:02:37 -07:00
Ryan Houdek 153d20ca59 Rename SHMStats 2025-07-29 12:02:19 -07:00
Ryan Houdek 03b04771be TestHarnessRunner: Become a real thread
Stop being so special as a host runner.
2025-07-29 12:02:19 -07:00
Ryan Houdek 392fa62dae Profiler: Decouple profile stats from the profiler option
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
2025-07-29 12:02:19 -07:00
Ryan Houdek af3918491c unittests: Fixes thunk unittest path
This hasn't been run in CI for awhile so this was missed when I was
testing it. I need to see what I can do to get this going in CI again.
2025-07-29 11:46:30 -07:00
LC c0762bbd82 Merge pull request #4737 from Sonicadvance1/i_dislike_tuple_13
FileManagement: Remove pair usage from GetEmulatedFDPath
2025-07-29 10:48:38 -04:00
LC b68272b413 Merge pull request #4733 from Sonicadvance1/i_dislike_tuple_9
OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
2025-07-29 10:46:08 -04:00
LC 90b9414b5f Merge pull request #4732 from Sonicadvance1/i_dislike_tuple_8
FEXCore: Remove unused refcount_shared_mutex
2025-07-29 10:43:25 -04:00
LC 5db494d4f7 Merge pull request #4736 from Sonicadvance1/i_dislike_tuple_12
x87StackOptimizationPass: Removes pair usage
2025-07-29 10:42:55 -04:00
Tony Wasserka d073066293 Merge pull request #4722 from Sonicadvance1/fix_4124
CMake: Work around QCom disabling SVE in their chips
2025-07-29 10:54:18 +02:00
LC e80270ae43 Merge pull request #4735 from Sonicadvance1/i_dislike_tuple_11
64BitAllocator: Removes pair usage in allocator
2025-07-28 21:41:36 -04:00
LC cd56b85e88 Merge pull request #4734 from Sonicadvance1/i_dislike_tuple_10
Arm64: Remove pair usage in 128-bit loader
2025-07-28 21:40:57 -04:00
Ryan Houdek 64fbf55bdd FileManagement: Remove pair usage from GetEmulatedFDPath
NFC
2025-07-28 16:21:49 -07:00
Ryan Houdek 14c1ee10b6 x87StackOptimizationPass: Removes pair usage
NFC
2025-07-28 16:14:49 -07:00
Ryan Houdek 96ae671738 64BitAllocator: Removes pair usage in allocator
NFC
2025-07-28 16:07:07 -07:00
Ryan Houdek 16552d1194 Arm64: Remove pair usage in 128-bit loader
NFC
2025-07-28 15:58:11 -07:00
Ryan Houdek 6470c98ee2 OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
NFC
2025-07-28 15:50:31 -07:00
Ryan Houdek 126c4efa0e FEXCore: Remove unused refcount_shared_mutex 2025-07-28 15:43:44 -07:00
Ryan Houdek f6fd9e18d9 Merge pull request #4727 from bylaws/meopd
Windows: Lock invalidation tracking for the entire duration of memory ops
2025-07-28 11:59:23 -07:00
Ryan Houdek 1f554867e0 CMake: Work around QCom disabling SVE in their chips
Fixes #4124

clang feature checking can't check beyond MIDR, so compiling for a
specific cortex version means compiling SVE on these CPUs that disabled
them.

Just detect the particular situation in-which SVE isn't inside
proc/cpuinfo and is one of the snapdragon cores that are supposed to
support SVE. Then compile for Cortex-a78 instead.
2025-07-28 11:57:02 -07:00
Ryan Houdek d226331ccc Merge pull request #4726 from bylaws/ijwidjn
WOW64: Wrap BTCpuSimulate to ensure correct unwinding
2025-07-28 11:50:34 -07:00
Ryan Houdek 123f8c93b5 Merge pull request #4721 from alyssarosenzweig/ir/pool-as-you-go
Pool constants as we go
2025-07-28 11:50:08 -07:00
Ryan Houdek 6e93b0a9df Merge pull request #4725 from bylaws/sver
Windows: Force-disable SVE usage for now
2025-07-28 11:16:20 -07:00
Billy Laws a5d6d8c928 Windows: Lock invalidation tracking for the entire duration of memory ops 2025-07-28 17:46:14 +01:00
Billy Laws 561ee6601b WOW64: Wrap BTCpuSimulate to ensure correct unwinding
Some compiler versions generated FP-relative operations before loading
it from the stack, which would crash wine when APCs were used.
2025-07-28 17:24:49 +01:00
Billy Laws f44a7c95f6 Windows: Force-disable SVE usage for now 2025-07-28 17:21:57 +01:00
Tony Wasserka ac420347d1 Merge pull request #4712 from Sonicadvance1/move_legacy_binfmt
cmake: Move legacy binfmt arch-specific targets to combined
2025-07-28 09:46:10 +02:00
Tony Wasserka 493a7ccdb6 Merge pull request #4717 from Sonicadvance1/i_dislike_tuple_6
ArchHelpers: Remove pair usage in unaligned handler
2025-07-28 09:40:09 +02:00
Alyssa Rosenzweig 387db78120 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig 1338a99add IR: cap constant pool
If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.

Difference at 95.0% confidence
        -0.00138911 +/- 0.00104724
        -0.418608% +/- 0.315587%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig f9edd20bf6 IR: pool constants in the emitter
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.

Difference at 95.0% confidence
	-0.00474273 +/- 0.00119189
	-1.40908% +/- 0.354114%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 10:50:17 -04:00
Alyssa Rosenzweig 6b3c7319c4 IR: wrap _Constant as Constant
flag day rename/wrapping. no functional change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:33:15 -04:00
Alyssa Rosenzweig bf51fc7c36 OpcodeDispatcher: do not use Constant as an identifier
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:32:42 -04:00
Alyssa Rosenzweig e910a81c12 OpcodeDispatcher: remove sized constant use
instcountci squashed for visibility.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:26:43 -04:00
Alyssa Rosenzweig 027bd93df9 IREmitter: remove unused
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:03:00 -04:00
LC 57c0d0f25e Merge pull request #4718 from Sonicadvance1/i_dislike_tuple_7
ELFContainer: Remove tuple usage
2025-07-25 23:04:52 -04:00
Ryan Houdek 24ff6bfde5 ELFContainer: Remove tuple usage 2025-07-25 19:33:32 -07:00
LC 78770683fc Merge pull request #4716 from Sonicadvance1/i_dislike_tuple_5
FEXCore: Replace CustomIREntry tuple with struct
2025-07-25 21:38:55 -04:00
LC c0008af877 Merge pull request #4714 from Sonicadvance1/i_dislike_tuple_3
IR: Remove tuple usage from NodeIterator
2025-07-25 21:37:56 -04:00
LC bffb81241a Merge pull request #4715 from Sonicadvance1/i_dislike_tuple_4
FEXCore: Remove reference SHA implementation
2025-07-25 21:34:29 -04:00
Ryan Houdek 53ff6d54c3 Merge pull request #4720 from tobhe/readme
Readme.md: Mention Ubuntu 25.04 as supported
2025-07-25 16:40:48 -07:00
Tobias Heider 2278e3334e Readme.md: Mention Ubuntu 25.04 as supported
Support was added in 1786c2f157
2025-07-26 00:38:26 +02:00
Ryan Houdek b0c61e2b69 ArchHelpers: Remove pair usage in unaligned handler
It's being treated like optional, where a value means it has been
handled, and no value means it hasn't been handled. Stop using pair in
this case.
2025-07-25 13:01:18 -07:00
Ryan Houdek a402c308ad FEXCore: Replace CustomIREntry tuple with struct 2025-07-25 12:33:25 -07:00
Ryan Houdek 6d83a195c6 InstcountCI: Remove sha 2025-07-25 12:22:13 -07:00
Ryan Houdek 0d69e88d53 FEXCore: Remove reference SHA implementation
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.

This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!

But now as we are no longer utilizing it, it is time to remove it.
2025-07-25 12:18:13 -07:00
Ryan Houdek 7309a5c3e1 HostFeatures: Fixes SHA check for vixl sim
Oops, was accidentally checking for rand.
2025-07-25 12:17:45 -07:00
Ryan Houdek 9e3c7aa9b3 IR: Remove tuple usage from NodeIterator 2025-07-25 12:10:39 -07:00
LC 30fb992cf7 Merge pull request #4708 from Sonicadvance1/geekbench_microtest
unittests/instcountci: Adds a long-lived ymm_high test
2025-07-24 21:31:04 -04:00
LC 8e5db99ba1 Merge pull request #4709 from Sonicadvance1/lock_xadd
unittests/instcountci: Adds missing lock xadd tests
2025-07-24 21:30:22 -04:00
Ryan Houdek 8103cdbec4 Merge pull request #4711 from bylaws/rwxinf
InvalidationTracker: Fix queries in the non-intersecting RWX case
2025-07-24 16:23:11 -07:00
Ryan Houdek b11835aab5 cmake: Move legacy binfmt arch-specific targets to combined
This matches the systemd path, no more _32 and _64 versions, just
binfmt_misc.

Having a mixed install where one program does one architecture and
another is weird and unsupported anyway.
2025-07-24 16:18:35 -07:00
Ryan Houdek 5b855d9baf unittests/instcountci: Adds missing lock xadd tests
Realized we were missing these, Looks like some moves could be
eliminated.
2025-07-24 15:20:02 -07:00
Billy Laws 56b6b96def InvalidationTracker: Fix queries in the non-intersecting RWX case
This needs to return the base of the non-RWX interval, with a size
that when added to the base is the start of the RWX interval.
2025-07-24 22:58:02 +01:00
Ryan Houdek 3744cad3d8 unittests/instcountci: Adds a long-lived ymm_high test
Spills in to spill-slots when it should instead spill in to the context.
2025-07-24 10:24:53 -07:00
Ryan Houdek bf8b5ed9ad Merge pull request #4670 from bylaws/callret
Implement call-ret stack optimisations
2025-07-24 10:24:30 -07:00
Billy Laws 49482fe963 InstCountCI: Update 2025-07-24 14:53:09 +01:00
Billy Laws 3497870a45 JIT: Guard lookupcache locks with the code invalidation mutex
Avoids issues with forking, as the code invalidation mutex is fork-safe.
2025-07-24 14:53:09 +01:00
Billy Laws 9a1efacefe OpcodeDispatcher: Allow for direct linking of non-multiblock direct jumps
Using an add here prevents ExitFunction from taking the direct path.

Reported by chengmingtang on Discord.
2025-07-24 14:53:09 +01:00
Billy Laws 3efb2379ee Dispatcher: Keep the call-ret stack balanced for thunk callbacks 2025-07-24 14:53:09 +01:00
Billy Laws 20efadbe66 OpcodeDispatcher: Treat ThunkOp ExitFunction as a return
ThunkOp acts as an implicit return, mark it as such so the call-ret
stack entry from the caller is popped
2025-07-24 14:53:09 +01:00
Billy Laws 31f8abfa61 Linux: Manage the call-ret stack 2025-07-24 14:53:09 +01:00
Billy Laws cf4cc71010 Dispatcher: Opportunistically perform a call-ret stack return on EC entry 2025-07-24 14:53:09 +01:00
Billy Laws 44107757a3 JIT: Rewrite block linking to support direct ExitFunction calls
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.

To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset

If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.

Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.

(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.

(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).

(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).

(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).

(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).

(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
2025-07-24 14:53:09 +01:00
Billy Laws 45ba1af388 BranchOps: Use the call-ret stack to optimise indirect ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws ba9884a26a JIT: Emit entrypoint code for call return target blocks
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
2025-07-24 14:53:09 +01:00
Billy Laws a40d53497b OpcodeDispatcher: Emit hints for call/ret instructions 2025-07-24 14:53:09 +01:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws c604893944 WOW64: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 2f16d25ab8 ARM64EC: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 132de3160a Windows: Implement common call-ret stack helpers
Allocate the stack with uncommited guard pages on either side, if
any faulting accesses to these occur then the call-ret SP value is
reset to the default and execution resumed.
2025-07-24 14:53:09 +01:00
Billy Laws d67b1645c9 FEXCore: Hold a frontend allocation for the call-ret stack
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
2025-07-24 14:53:09 +01:00
Billy Laws 68270ad425 LookupCache: Return whether Erase removed any cache entries 2025-07-24 14:53:09 +01:00
Billy Laws 963a8c2f08 FEXCore: Save and restore the call/ret SP from CPUState
For simplicity in cases like signal handling, always load it in
fill and store in spill, even the SP is stored in a callee save
register.
2025-07-24 14:53:09 +01:00
Billy Laws a80581e52b Arm64Emitter: Allocate a register for the call-ret SP
The host stack pointer can't be reused since explicit bounds checks
would be far too expensive, and on-stack signals prevent implicit ones
using guard pages from working.
2025-07-24 14:53:09 +01:00
Billy Laws 882cdaaeda Frontend: Move to initializing persistent sets directly 2025-07-24 14:52:54 +01:00
Billy Laws fd58f17dbe FEXCore: Track block executable ranges prior to adding cache entries
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.

This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
2025-07-24 14:52:54 +01:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Ryan Houdek 525462e982 Merge pull request #4642 from pmatos/feature/x87-invalid-operation-bit
Implement x87 invalid operation bit on F80 mode
2025-07-23 11:05:28 -07:00
Ryan Houdek 507b3a69f0 Merge pull request #4699 from bylaws/srawow64
Inline SMC fixes and support for WOW64
2025-07-23 11:03:18 -07:00
Billy Laws b15e49910a WOW64: Support inline SMC
Matches the ARM64EC handling
2025-07-23 18:46:28 +01:00
Billy Laws 573d0858ac ARM64EC: Fix inline SMC handling for writes in the same page as the current block
Consider a page with two blocks in it, A and B. Block A performs SMC on B then A.
With the previous logic, the SMC write of B would unprotect the page and then
the inline SMC touching A would not be detected and a single-step would not be forced.
2025-07-23 18:46:28 +01:00
Billy Laws c3c2789cdf Windows: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:28 +01:00
Billy Laws e9f14e32dd Linux: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:26 +01:00
Billy Laws 9b65d819f4 Dispatcher: Add an argument to request a single-step on SRA fill
This already exists for ARM64EC, but this path is required for linux
and WOW64
2025-07-23 18:29:12 +01:00
Billy Laws 287344986c FEXCore: Switch ENTRY_FILL_SRA_SINGLE_INST_REG to TMP2
Will allow this to be taken as the second dispatcher argument
2025-07-23 18:29:12 +01:00
Paulo Matos 7386b4b037 instcountci: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 9a2f62c479 asm_tests: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 726656d0bf Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:10 +02:00
Ryan Houdek fe61aabc52 Merge pull request #4701 from bylaws/noexecwin
Windows: Correctly report noexec faults
2025-07-22 12:54:28 -07:00
Ryan Houdek 564e966730 Merge pull request #4698 from lioncash/float2
ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
2025-07-22 12:36:16 -07:00
LC 98abfa4395 Merge pull request #4700 from bylaws/geiex
WinAPI: Avoid sign-extension of processor count in GetSystemInfo
2025-07-22 15:02:17 -04:00
Billy Laws 8e6fbe7615 Windows: Correctly report noexec faults 2025-07-22 19:30:31 +01:00
Billy Laws 27fd5505be OpcodeDispatcher: Set correct trap number for noexec faults 2025-07-22 19:30:23 +01:00
Billy Laws e383cc89eb unittests: Test for the noexec pagefault trap number 2025-07-22 19:29:17 +01:00
Alexandre Julliard a71a0ed137 WinAPI: Avoid sign-extension of processor count in GetSystemInfo 2025-07-22 19:28:36 +01:00
Lioncache ab6fc279ce ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
Forgot in #4697 to take into account that there are still parts of the
emitter that qualify with its own namespace.

This also removes those where applicable to be more in line with the rest
of the cases.
2025-07-22 08:25:35 -04:00
Ryan Houdek dd5c17291a Merge pull request #4697 from lioncash/float
SVEOps: Make use of IsStandardFloatSize() more
2025-07-21 12:37:03 -07:00
Ryan Houdek 5fe06b7413 Merge pull request #4696 from lioncash/opcode
OpcodeDispatcher: Mark functions as const/static where applicable
2025-07-21 12:36:22 -07:00
Lioncache 2a8efa9891 SVEOps: Make use of IsStandardFloatSize() more
Initially introduced in #4687 as a utility for ASIMD ops, this can be used
elsewhere as well to reduce the verbosity of some other assertions.
2025-07-21 13:51:37 -04:00
LC cf43e5eaaf Merge pull request #4694 from Sonicadvance1/fix_moffset_instcountci
InstcountCI: Fix bad encoded moffset instructions
2025-07-21 10:53:26 -04:00
Lioncache cafc906b9e OpcodeDispatcher: Mark functions as const/static where applicable
Not a functional change, but makes it more obvious how these rely on object state.
2025-07-21 10:48:55 -04:00
Tony Wasserka 3387f51e24 Merge pull request #4663 from Sonicadvance1/update_format_requires
External/code-format-helper: Update requirements
2025-07-21 10:17:52 +02:00
Ryan Houdek ccf1bb26af InstcountCI: Fix bad encoded moffset instructions
These were actually encoded incorrectly, and I noticed they changed in
PR #4670 for some reason. So I dove in to the nasm source to figure out
what it takes to encode these sanely rather than with raw bytes, turns
out it's fairly trivial, just annoying to track in their parser through
iwdq->mem_offset->MEM_OFFS->qword.

Should remove the change in #4670 once it rebases.
2025-07-18 12:32:56 -07:00
Ryan Houdek fc1ca01a0b Merge pull request #4693 from neobrain/fix_libfwd_wl
LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options
2025-07-18 10:57:50 -07:00
Tony Wasserka e508f0900c LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options 2025-07-18 11:38:46 +02:00
LC 5123ca52ba Merge pull request #4692 from Sonicadvance1/reintroduce_cssc
FEXCore: Reintroduce support for CSSC
2025-07-17 21:21:46 -04:00
Ryan Houdek 7e1ee5bb07 FEXCore: Reintroduce support for CSSC
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.

Been a while since I last looked at this, added a new instcountci file
to show the improvement.
2025-07-17 15:09:27 -07:00
LC b7ac641aa0 Merge pull request #4691 from Sonicadvance1/ignore_jemalloc_config_sources
External: Update jemallocs
2025-07-17 11:29:20 -04:00
Ryan Houdek 251ac145d9 Merge pull request #4650 from pmatos/feature/ReformatChanged
Add --changed flag to reformat.sh script
2025-07-16 23:33:29 -07:00
Ryan Houdek e17c2c8c73 Merge pull request #4661 from pmatos/feature/reformat-clang-format-19
Whole-tree reformat with clang-format-19
2025-07-16 23:33:09 -07:00
Paulo Matos 4473054d33 Add reformat sha to ignore revs for git-blame 2025-07-17 08:11:08 +02:00
Paulo Matos 5267cde60e Whole-tree reformat with clang-format-19 2025-07-17 08:10:00 +02:00
Paulo Matos 3c7ece2ef3 Update .clang-format 2025-07-17 08:09:25 +02:00
Paulo Matos 3b322d8f37 Add --changed flag to reformat.sh script 2025-07-17 08:04:41 +02:00
LC 1d4b6c6a74 Merge pull request #4689 from Sonicadvance1/disable_gcs_protection
Disable GCS in simulator and userspace
2025-07-16 20:25:16 -04:00
Billy Laws f5efe1d251 Merge pull request #4690 from Sonicadvance1/moar_padding
JIT: Add more padding
2025-07-17 00:58:37 +01:00
Ryan Houdek 09db3aba0a External: Update jemallocs
Ignore any additional jemalloc configuration options, can come from
`/etc/malloc.conf` or even the `MALLOC_CONF` environment variable.

Ensure nothing can override our options.
2025-07-16 16:35:48 -07:00
Ryan Houdek aec7deca9d Merge pull request #4688 from lioncash/hfloat2
ASIMDOps: Merge half-float 3-reg same with single/double variants
2025-07-16 15:42:08 -07:00
Ryan Houdek 8f1d4bc710 FEXLoader: Check for GCS being enabled
There is a ELF note for this but currently clang doesn't support a
`no-gcs` flag. The best we can do is check if the kernel has GCS enabled
for the current process and early exit.

Then continue to use the kernel's locking functionality to disable it if
the guest happens to try, ensuring safety.
2025-07-16 15:41:23 -07:00
Ryan Houdek 7edc8417b6 JIT: Add more padding
Go to a whole page of additional padding, #4670 adds some more size to a
block and overran the padding causing a unittest to fail.
2025-07-16 15:28:05 -07:00
Ryan Houdek 5b01682641 Arm64Emitter: Disable GCS in the simulator
FEX isn't going to be compatible with this.
PR #4670 requires this
2025-07-16 15:23:43 -07:00
Lioncache a615f8970a ASIMDOps: Constrain 3-reg same float arguments with IsQOrDRegister
These are only intended to be used with QRegisters or DRegisters, so we can constrain these
so that the proper types are enforced at compile-time.
2025-07-16 17:46:28 -04:00
Lioncache 708ea9d50a ASIMDOps: Merge half-float 3-reg same with single/double variants 2025-07-16 17:36:24 -04:00
Ryan Houdek e36a64ecc0 Merge pull request #4681 from Sonicadvance1/i_dislike_tuple_2
FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
2025-07-16 14:23:18 -07:00
Ryan Houdek 82be09b6a3 FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
NFC
2025-07-16 13:57:16 -07:00
Ryan Houdek eb7655d93c External/code-format-helper: Update requirements
Latest of everything, let's see what happens.
To get rid of dependabot alerts.
2025-07-16 13:54:43 -07:00
Ryan Houdek c1fe841bf3 Merge pull request #4687 from lioncash/hfloat
ASIMDOps: Merge half-float 2-reg misc with single/double variants
2025-07-16 13:50:53 -07:00
Lioncache 82ffb379c4 ASIMDOps: Merge half-float 2-reg misc with single/double variants
Unifies the interface, so that there's no need for a stark difference.

Previously, calling the single/double variant didn't require explicit template
arguments, but the half-float version did, which is inconsistent.

Technically it also made the interface more cumbersome to use in the event the
element size isn't able to be determined as a constant ahead of time.
2025-07-16 10:16:49 -04:00
Tony Wasserka befae52993 Merge pull request #4685 from pmatos/fix/clang-format-19-ignore
Ensure .clang-format-ignore is compatible with clang-format-19
2025-07-16 15:32:09 +02:00
Paulo Matos 48597d7682 Ensure .clang-format-ignore is compatible with clang-format-19
Since clang-format-19 doesn't support globstar yet, add .clang-format
to disable formatting inside External/.
2025-07-16 15:20:03 +02:00
Lioncache 05d45ccb7f Emitter: Add helper for sanitizing floating point element sizes
Will be used in a following change to reduce the amount of duplication
made in some floating point instructions.
2025-07-16 09:19:45 -04:00
LC e0ca04bfec Merge pull request #4684 from Sonicadvance1/spurious_syscall_crash
FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
2025-07-15 21:44:51 -04:00
LC 34566861ba Merge pull request #4683 from Sonicadvance1/waitpkg_mostly_nop
FEXCore: Implement a mostly NOP implementation of waitpkg
2025-07-15 21:43:11 -04:00
Ryan Houdek dfe3b4502a FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
If multiblock discovers a codepath with an `int 0x80` then it would
crash the emulator even if it never gets executed. Ensure that this
ERROR_AND_DIE_FMT instead just gets handled as an UnhandledOp to ensure
the Core early terminates the block.

Found by having steamwebhelper spuriously crash when it hit this.
Adds a simple unittest to ensure discovery doesn't break again.
2025-07-15 17:43:16 -07:00
Ryan Houdek 4884aeef20 Config: Adds option to disable wfxt 2025-07-15 16:38:15 -07:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Ryan Houdek 29fccd0e1b Merge pull request #4682 from bylaws/asahi
Windows: Support enabling hardware TSO on Asahi Linux
2025-07-15 14:25:03 -07:00
Ryan Houdek 35e4ac5a64 Merge pull request #4668 from Sonicadvance1/implement_nx
FEXCore: Implement support for NX bit.
2025-07-15 12:40:48 -07:00
Ryan Houdek 43bba77840 FEXCore: Implement support for NX bit.
Long time coming but thanks to bylaw's changes in #4474, this is now
trivial to implement.

Fixes #2175
2025-07-15 10:41:24 -07:00
Billy Laws 0361274366 Windows: Support enabling hardware TSO on Asahi
Relies on a wine-side patch to expose the prctl to the PE-side
2025-07-15 17:45:21 +01:00
Ryan Houdek 5b0e703811 Merge pull request #4678 from neobrain/refactor_unuse_unused
Drop unnecessary uses of maybe_unused
2025-07-15 09:03:55 -07:00
Ryan Houdek 05742a8e28 Merge pull request #4677 from neobrain/refactor_misc_logs
Miscellaneous log changes
2025-07-15 09:03:12 -07:00
Ryan Houdek 170fbb596c Merge pull request #4679 from neobrain/fix_signed_overflow
Arm64Emitter: Fix signed overflow
2025-07-15 09:02:29 -07:00
Ryan Houdek b72bdf9948 Merge pull request #4680 from lioncash/vfaddv
VectorOps: Correct benign VFAddV op cast
2025-07-15 09:00:41 -07:00
Lioncache 678e03b470 VectorOps: Correct benign VFAddV op cast
This was using the non-float VAddV variant, but had the same behavior,
since the fields were named the same. So this is just a correctness fix.
2025-07-15 11:41:28 -04:00
Tony Wasserka 0f45f3a24d Arm64Emitter: Fix signed integer overflows 2025-07-15 17:17:48 +02:00
Tony Wasserka 9b503f3702 IRDumper: Remove unnecessary use of maybe_unused 2025-07-15 17:10:37 +02:00
Tony Wasserka 7f3f619518 GDBJIT: Remove unnecessary use of maybe_unused 2025-07-15 17:10:27 +02:00
Tony Wasserka f7d59880cc ELFCodeLoader: Drop unused maybe_unused attribute 2025-07-15 17:07:36 +02:00
Tony Wasserka ec9e266c3c FDUtils: Mark get_fdpath as nodiscard 2025-07-15 17:07:36 +02:00
Tony Wasserka ed62c02494 AllocatorOverride: Make error message more prominent 2025-07-15 16:55:18 +02:00
Tony Wasserka d5b166bdd9 OpcodeDispatcher: Strengthen error message to fatal 2025-07-15 16:55:18 +02:00
Tony Wasserka 648c726bb1 LogManager: Use assert log level for ERROR_AND_DIE_FMT 2025-07-15 16:55:18 +02:00
Tony Wasserka 88f30afb26 Merge pull request #4672 from neobrain/refactor_format_formatter
code-format-helper: Fix formatting for the format helper
2025-07-15 14:27:22 +02:00
LC e6f319d6b3 Merge pull request #4674 from neobrain/refactor_todrop
Drop TODO defines
2025-07-15 07:12:35 -04:00
LC a80d8b15c2 Merge pull request #4673 from neobrain/refactor_value_nor
XXFileHash: Drop unnecessary use of value_or
2025-07-15 07:09:47 -04:00
LC 2a372895e4 Merge pull request #4671 from Sonicadvance1/xsaveopt
FEXCore: Implement xsaveopt
2025-07-15 07:08:15 -04:00
Tony Wasserka 3b67e573b5 Revert "FexHeaderUtils: Add TodoDefines"
This reverts commit ad1fd7f54b.
2025-07-15 10:04:01 +02:00
Tony Wasserka a3a55d19b8 Revert "FEX_TODO: Convert some XXX to FEX_TODO"
This reverts commit 256df76674.
2025-07-15 10:03:56 +02:00
Tony Wasserka 4e25bce616 FEXServer: Drop use of FEX_TODO macro 2025-07-15 10:03:49 +02:00
Tony Wasserka 135477e539 XXFileHash: Drop unnecessary use of value_or 2025-07-15 09:47:52 +02:00
Tony Wasserka 147b1f2293 code-format-helper: Replace spurious tab with spaces 2025-07-15 09:43:55 +02:00
Ryan Houdek 401dc69586 InstcountCI: Add xsaveopt
Just in-case it changes.
2025-07-14 17:32:45 -07:00
Ryan Houdek 0dff992fab FEXCore: Implement xsaveopt
It's the same as xsave because we don't track a hidden "xinuse" hardware
mask. So this is a trivial implementation.
2025-07-14 17:31:47 -07:00
LC 4f605739f1 Merge pull request #4667 from Sonicadvance1/i_dislike_tuple
XXFileHash: Remove a tuple usage
2025-07-14 18:21:54 -04:00
LC e71e10adcb Merge pull request #4669 from Sonicadvance1/handful_cpuinfo_missing
EmulatedFiles/cpuinfo: Add a few missing flags
2025-07-14 18:21:03 -04:00
Ryan Houdek 9f1bc05644 EmulatedFiles/cpuinfo: Add a few missing flags
bus_lock_detect might be interesting to expose if applications ever
change behaviour depending on the underlying fault behaviour.
2025-07-14 14:56:24 -07:00
Ryan Houdek 603878b3e9 XXFileHash: Remove a tuple usage 2025-07-14 13:09:27 -07:00
Ryan Houdek 70ce323537 Merge pull request #4666 from pmatos/fix/git-clang-format-19
Set clang-format-19 as the version git-clang-format should run in wor…
2025-07-14 10:38:14 -07:00
Paulo Matos b5f8dd93f7 Set clang-format-19 as the version git-clang-format
Unfortunately git-clang-format-19 will not call clang-format-19 but the system clang-format
so we need to hardcode it here.
2025-07-14 18:37:47 +02:00
Ryan Houdek 2e7b86e452 Merge pull request #4474 from bylaws/mentry
Support multiple entrypoints into a multiblock and executable permission tracking
2025-07-11 19:05:42 -07:00
Ryan Houdek b2eeaf75a9 Merge pull request #4659 from bylaws/sysfix
ARM64EC: Rely on syscall export sorting for deriving their IDs
2025-07-11 18:09:07 -07:00
Ryan Houdek 5d4bddcfb9 Merge pull request #4662 from pmatos/patch-2 2025-07-11 08:49:45 -07:00
Paulo Matos 13ef676115 Do not reformat files in External/ 2025-07-11 15:54:02 +02:00
Ryan Houdek 67bfd3877c Merge pull request #4651 from pmatos/feature/ClangFormat19Upgrade
Upgrade to clang-format-19
2025-07-11 01:03:49 -07:00
Ryan Houdek 2be619b011 Merge pull request #4658 from bylaws/fmt
External: Update libfmt to master to support clang 21
2025-07-10 11:08:33 -07:00
Ryan Houdek ebd0a3ceaf Merge pull request #4660 from lioncash/null
VixlUtils: Fix null pointer dereference vector in IsImmLogical()
2025-07-10 11:07:01 -07:00
Lioncache d8da14d550 VixlUtils: Fix null pointer dereference vector in IsImmLogical()
Technically we can end up doing null pointer dereferencing here since checks were being
chained with || instead of &&. So we can just separate the checks out.
2025-07-10 12:09:20 -04:00
Billy Laws 076c156cbb Frontend: Warn on invalid instructions in entry blocks 2025-07-10 16:51:34 +01:00
Billy Laws 6eaaab8dc2 IntervalList: Also return the full matching interval on query 2025-07-10 16:51:34 +01:00
Billy Laws 46bc8e499a Frontend: Always treat FEXCore X86 callbacks as executable 2025-07-10 16:51:34 +01:00
Billy Laws 21865d09a2 SyscallsSMCTracking: Handle READ_IMPLIES_EXEC for executable mapping queries 2025-07-10 16:51:34 +01:00
Billy Laws 31903d0c0b Frontend: Treat instructions in non-executable memory as invalid 2025-07-10 16:51:33 +01:00
Paulo Matos c7b7cbcac1 Upgrade to clang-format-19
Fixes #4577
2025-07-10 17:11:44 +02:00
Billy Laws d0af858a9c LinuxEmulation: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 2e5c6283f5 CodeSizeValidation: Add dummy QueryGuestExecutableRange impl 2025-07-10 16:00:24 +01:00
Billy Laws 261df26cff DummyHandlers: Stub SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 96bd6ce3e5 Windows: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 14eed89bb6 SyscallHandler: Add method to query executable memory ranges 2025-07-10 16:00:24 +01:00
Billy Laws 3feb354186 Frontend: Keep the associated thread object as a member
Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.
2025-07-10 16:00:24 +01:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Billy Laws cdef1ed0c5 Frontend: Explore after call instructions with multiblock
Each instruction after a call instruction can be treated as an
additional entrypoint to the multiblock
2025-07-10 16:00:24 +01:00
Billy Laws 4f79acc32a X86Tables: Mark call instructions with a flag 2025-07-10 16:00:24 +01:00
Billy Laws 9328b099c3 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 16:00:24 +01:00
Billy Laws 378d0351cf External: Update libfmt to master to support clang 21 2025-07-10 16:00:24 +01:00
Billy Laws 793e3adbb4 ARM64EC: Rely on syscall export sorting for deriving their IDs
Rather than relying on wine-specific alphabetical behaviour, that
has since been changed. Rely on their addresses being sorted which
is more stable behaviour also present in Windows.
2025-07-10 16:00:07 +01:00
Billy Laws 2f4f4253e7 External: Update libfmt to master to support clang 21 2025-07-10 15:52:17 +01:00
Ryan Houdek 10c69d561e Merge pull request #4656 from bylaws/mingw-new
CMake: Set CMAKE_AR for MinGW toolchains
2025-07-09 19:26:30 -07:00
Billy Laws bbf6e8e0c5 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 00:33:37 +01:00
Ryan Houdek d8a4d03501 Merge pull request #4655 from alyssarosenzweig/bug/divisor-mask
Fix divisor masking
2025-07-09 12:15:16 -07:00
Alyssa Rosenzweig 8ecc8ef75d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig cfd4318080 unittests: add 32-bit masking divide unit test
fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig d23bb01e96 JIT: fix divisor masking
oversight. should fix Steam.

Fixes: de4becc26 ("OpcodeDispatcher: mask certain divisors")
Closes: #4652
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Ryan Houdek 60f63ca467 Merge pull request #4649 from bylaws/rexvex
Frontend: Raise #UD on invalid VEX/REX encodings
2025-07-09 10:37:50 -07:00
LC 40378cd4d8 Merge pull request #4654 from Sonicadvance1/vcvtsd2si_test
unittests: Update vcvtsd2si test
2025-07-09 12:16:02 -04:00
Ryan Houdek f44af48751 Merge pull request #4648 from bylaws/sfmodreg
SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops
2025-07-09 09:02:12 -07:00
Ryan Houdek d01a48b773 unittests: Update vcvtsd2si test
This would have failed prior to #4647 getting merged.
2025-07-09 08:50:19 -07:00
Ryan Houdek bd136eae71 Merge pull request #4647 from bylaws/sizrfix
OpcodeDispatcher: Fix several cases where incorrectly sized loads/stores could be used
2025-07-09 08:49:38 -07:00
LC e669370629 Merge pull request #4646 from bylaws/uaffix
PoolBufferWithTimedRetirement: Unclaim in dtor
2025-07-09 10:14:13 -04:00
Billy Laws 26215d5425 VEXTables: Complete decode flag information 2025-07-09 00:53:23 +01:00
Billy Laws baa5c6f79b Frontend: Require VEX.vvvv is 0 when unused 2025-07-09 00:53:23 +01:00
Billy Laws c0fcb27aa5 Frontend: Add instruction flags to specify valid REX.W encodings 2025-07-09 00:53:23 +01:00
Billy Laws 02b7938b93 Frontend: Add instruction flags to specify valid VEX.L encodings 2025-07-09 00:53:23 +01:00
Billy Laws 00fec3d51f Update InstCountCI 2025-07-09 00:52:59 +01:00
Billy Laws 0f6ceaac05 OpcodeDispatcher: Load only the element size from memory for VFMAScalarImpl 2025-07-09 00:52:59 +01:00
Billy Laws 635816c07f OpcodeDispatcher: Force ElementSize loads for UCOMISxOp 2025-07-09 00:52:59 +01:00
Billy Laws 2fac5c23ff OpcodeDispatcher: Always use 32-bit load/store for {LD,ST}MXCSR 2025-07-09 00:52:59 +01:00
Billy Laws ff4c1cf0d5 X86Tables: Fix (V)CV(T)TSD2SI size flags
This led to incorrect OOB handling for 32-bit dests.
2025-07-09 00:52:45 +01:00
Billy Laws c2e2d1e92f OpcodeDispatcher: Only read at most ElementSize in CVTFPR_To_GPR 2025-07-09 00:50:47 +01:00
Billy Laws 0ab56bad95 SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops 2025-07-08 23:45:32 +01:00
Billy Laws 407c5a0f78 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-08 23:38:43 +01:00
564 changed files with 37465 additions and 35062 deletions

No files matched your search

+2 -2
View File
@@ -32,7 +32,7 @@ AttributeMacros:
BinPackArguments: true
BinPackParameters: true
BitFieldColonSpacing: Both
BreakAfterAttributes: Always # clang 16 required
BreakAfterAttributes: Leave
BreakBeforeBraces: Attach
BreakBeforeBinaryOperators: None
BreakBeforeInlineASMColon: OnlyMultiline # clang 16 required
@@ -60,7 +60,7 @@ IndentRequires: false
IndentWidth: 2
InsertBraces: true
KeepEmptyLinesAtTheStartOfBlocks: true
LambdaBodyIndentation: OuterScope
LambdaBodyIndentation: Signature
LineEnding: LF # clang 16 required
MaxEmptyLinesToKeep: 2
NamespaceIndentation: Inner
-4
View File
@@ -1,8 +1,4 @@
# This file is used to ignore files and directories from clang-format
# Ignore all files in the External directory
External/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
+4
View File
@@ -16,3 +16,7 @@
# Reformat of CodeEmitter inl files
8760c593ece92d7e9fa94c40da0368fd367c9cad
# Whole-tree reformat with clang-format-19
5267cde60e7642852d18f20ae8568643bb5293d5
+4
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
build_plus_test:
@@ -136,6 +137,9 @@ jobs:
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
env:
# These tests require non-portable install due to thunks.
FEX_PORTABLE: 0
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
+1
View File
@@ -20,6 +20,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
glibc_fault_test:
+1
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
hostrunner_tests:
+4 -11
View File
@@ -40,11 +40,8 @@ jobs:
echo "Formatting files:"
echo "$CHANGED_FILES"
- name: Check for correct clang-format version
run: clang-format --version | grep -qF '16.0.6'
- name: Check git-clang-format-16 exists
run: which git-clang-format-16
- name: Check git-clang-format-19 exists
run: which git-clang-format-19
- name: Setup Python env
uses: actions/setup-python@v4
@@ -58,19 +55,15 @@ jobs:
- name: Run code formatter
env:
CLANG_FORMAT_PATH: 'git-clang-format-16'
CLANG_FORMAT_PATH: 'git-clang-format-19'
GITHUB_PR_NUMBER: ${{ github.event.pull_request.number }}
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
# TODO(pmatos): Once we adopt v18, we should be able
# to take advantage of the new --diff_from_common_commit option
# explicitly in code-format-helper.py and not have to diff starting at
# the merge base.
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
--issue-number $GITHUB_PR_NUMBER \
--start-rev $(git merge-base $START_REV $END_REV) \
--start-rev $START_REV \
--end-rev $END_REV \
--changed-files "$CHANGED_FILES"
+1
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
vixl_simulator:
+88
View File
@@ -0,0 +1,88 @@
name: Wine DLL artifacts
on:
push:
branches:
- main
env:
BUILD_TYPE: Release
jobs:
wine_dll_artifacts:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, mingw]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean install directory
run: |
rm -Rf ${{runner.workspace}}/build_install
mkdir ${{runner.workspace}}/build_install
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build_arm64ec
rm -Rf ${{runner.workspace}}/build_wow64
- name: Create Build Environment arm64ec
run: |
cmake -E make_directory ${{runner.workspace}}/build_arm64ec
cmake -E make_directory ${{runner.workspace}}/build_wow64
- name: Configure CMake arm64ec
shell: bash
working-directory: ${{runner.workspace}}/build_arm64ec
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Configure CMake wow64
shell: bash
working-directory: ${{runner.workspace}}/build_wow64
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Build arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Build wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: wine_dll_artifacts
path: ${{runner.workspace}}/build_install/usr/lib/wine/aarch64-windows/lib*.dll
retention-days: 60
compression-level: 9
+3
View File
@@ -46,3 +46,6 @@
[submodule "External/tracy"]
path = External/tracy
url = https://github.com/wolfpld/tracy
[submodule "External/range-v3"]
path = External/range-v3
url = https://github.com/ericniebler/range-v3.git
+19
View File
@@ -39,10 +39,16 @@ set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set (HOSTLIBS_DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
if (NOT DATA_DIRECTORY)
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
endif()
include(GNUInstallDirs)
if (NOT HOSTLIBS_DATA_DIRECTORY)
set(HOSTLIBS_DATA_DIRECTORY "${CMAKE_INSTALL_FULL_LIBDIR}/fex-emu")
endif()
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
message (STATUS "Mingw build")
@@ -348,6 +354,12 @@ if (NOT fmt_FOUND)
add_subdirectory(External/fmt/)
endif()
find_package(range-v3 QUIET)
if (NOT range-v3_FOUND)
add_subdirectory(External/range-v3/)
target_compile_definitions(range-v3 INTERFACE RANGES_DISABLE_DEPRECATED_WARNINGS)
endif()
add_subdirectory(External/tiny-json/)
include_directories(External/tiny-json/)
@@ -406,6 +418,13 @@ if (TUNE_CPU STREQUAL "native")
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/NeedDisabledSVE.py"
RESULT_VARIABLE NEEDS_SVE_DISABLED)
if (NEEDS_SVE_DISABLED)
message(STATUS "Platform has bugged SVE. Disabling")
set(AARCH64_CPU "cortex-a78")
endif()
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${AARCH64_CPU}")
+6 -6
View File
@@ -174,7 +174,7 @@ public:
// Logical immediate
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
and_(s, rd, rn, n, immr, imms);
}
@@ -185,7 +185,7 @@ public:
void ands(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
ands(s, rd, rn, n, immr, imms);
}
@@ -196,14 +196,14 @@ public:
void orr(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
orr(s, rd, rn, n, immr, imms);
}
void eor(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
eor(s, rd, rn, n, immr, imms);
}
@@ -333,7 +333,7 @@ public:
bfi(s, rd, Reg::zr, lsb, width);
}
void bfxil(ARMEmitter::Size s, Register rd, Register rn, uint32_t lsb, uint32_t width) {
[[maybe_unused]] const auto reg_size_bits = RegSizeInBits(s);
const auto reg_size_bits = RegSizeInBits(s);
const auto lsb_p_width = lsb + width;
LOGMAN_THROW_A_FMT(width >= 1, "bfxil needs width >= 1");
@@ -977,7 +977,7 @@ private:
}
void xbfiz_helper(bool is_signed, ARMEmitter::Size s, Register rd, Register rn, uint32_t lsb, uint32_t width) {
[[maybe_unused]] const auto lsb_p_width = lsb + width;
const auto lsb_p_width = lsb + width;
const auto reg_size_bits = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(lsb_p_width <= reg_size_bits, "lsb + width ({}) must be <= {}. lsb={}, width={}", lsb_p_width, reg_size_bits, lsb, width);
File diff suppressed because it is too large. Load diff
+9
View File
@@ -12,6 +12,7 @@
#include <CodeEmitter/Registers.h>
#include <array>
#include <bit>
#include <cstdint>
#include <utility>
#include <type_traits>
@@ -86,6 +87,14 @@ constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t {8} << FEXCore::ToUnderlying(size);
}
// Many floating point operations constrain their element sizes to the
// main three float sizes half, single, and double precision. This just
// combines all the checks together for brevity.
[[nodiscard]]
constexpr bool IsStandardFloatSize(SubRegSize size) {
return size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit;
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
+41 -48
View File
@@ -60,8 +60,7 @@ public:
}
void fcmla(SubRegSize size, ZRegister zda, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcmla can only use p0 to p7");
uint32_t Op = 0b0110'0100'0000'0000'0000'0000'0000'0000;
@@ -76,8 +75,7 @@ public:
}
void fcadd(SubRegSize size, ZRegister zd, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcadd can only use p0 to p7");
LOGMAN_THROW_A_FMT(rot == Rotation::ROTATE_90 || rot == Rotation::ROTATE_270, "fcadd rotation may only be 90 or 270 degrees");
LOGMAN_THROW_A_FMT(zd == zn, "fcadd zd and zn must be the same register");
@@ -815,16 +813,12 @@ public:
// SVE Integer Misc - Unpredicated
// SVE floating-point trig select coefficient
void ftssel(SubRegSize size, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "ftssel may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "ftssel may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b00, zm.Idx(), FEXCore::ToUnderlying(size), zd, zn);
}
// SVE floating-point exponential accelerator
void fexpa(SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "fexpa may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "fexpa may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b10, 0b00000, FEXCore::ToUnderlying(size), zd, zn);
}
// SVE constructive prefix (unpredicated)
@@ -1503,9 +1497,9 @@ public:
}
// SVE broadcast floating-point immediate (unpredicated)
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(size == ARMEmitter::SubRegSize::i16Bit || size == ARMEmitter::SubRegSize::i32Bit || size == ARMEmitter::SubRegSize::i64Bit,
"Unsupported fmov size");
void fdup(SubRegSize size, ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fmov size");
uint32_t Imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -1518,7 +1512,7 @@ public:
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
}
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
void fmov(SubRegSize size, ZRegister zd, float Value) {
fdup(size, zd, Value);
}
@@ -1547,7 +1541,7 @@ public:
void sqincp(SubRegSize size, XRegister rdn, PRegister pm) {
SVEIncDecPredicateCountScalar(0, 1, 0b10, 0b00, size, rdn, pm);
}
void sqincp(SubRegSize size, XRegister rdn, PRegister pm, [[maybe_unused]] WRegister wn) {
void sqincp(SubRegSize size, XRegister rdn, PRegister pm, WRegister wn) {
LOGMAN_THROW_A_FMT(rdn.Idx() == wn.Idx(), "rdn and wn must be the same");
SVEIncDecPredicateCountScalar(0, 1, 0b00, 0b00, size, rdn, pm);
}
@@ -1560,7 +1554,7 @@ public:
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm) {
SVEIncDecPredicateCountScalar(0, 1, 0b10, 0b10, size, rdn, pm);
}
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm, [[maybe_unused]] WRegister wn) {
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm, WRegister wn) {
LOGMAN_THROW_A_FMT(rdn.Idx() == wn.Idx(), "rdn and wn must be the same");
SVEIncDecPredicateCountScalar(0, 1, 0b00, 0b10, size, rdn, pm);
}
@@ -3302,7 +3296,7 @@ private:
const auto log2_size_bytes = FEXCore::ilog2(size_bytes);
// We can index up to 512-bit registers with dup
[[maybe_unused]] const auto max_index = (64U >> log2_size_bytes) - 1;
const auto max_index = (64U >> log2_size_bytes) - 1;
LOGMAN_THROW_A_FMT(Index <= max_index, "dup index ({}) too large. Must be within [0, {}].", Index, max_index);
// imm2:tsz make up a 7 bit wide field, with each increasing element size
@@ -3332,7 +3326,7 @@ private:
uint32_t shift = 0;
if (!is_uint8_imm) {
[[maybe_unused]] const bool is_uint16_imm = (imm >> 16) == 0;
const bool is_uint16_imm = (imm >> 16) == 0;
LOGMAN_THROW_A_FMT(is_uint16_imm, "Immediate ({}) must be a 16-bit value within [256, 65280]", imm);
LOGMAN_THROW_A_FMT((imm % 256) == 0, "Immediate ({}) must be a multiple of 256", imm);
@@ -3401,8 +3395,8 @@ private:
}
void SVEBroadcastFloatImmPredicated(SubRegSize size, ZRegister zd, PRegister pg, float value) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported fcpy/fmov "
"size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fcpy/fmov size");
uint32_t imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -3578,7 +3572,7 @@ private:
// SVE2 floating-point pairwise operations
void SVEFloatPairwiseArithmetic(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zd needs to equal zn");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0100'0001'0000'1000'0000'0000'0000;
@@ -3592,7 +3586,7 @@ private:
// SVE floating-point arithmetic (unpredicated)
void SVEFloatArithmeticUnpredicated(uint32_t opc, SubRegSize size, ZRegister zm, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
uint32_t Instr = 0b0110'0101'0000'0000'0000'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -3700,7 +3694,7 @@ private:
// SVE floating-point arithmetic (predicated)
void SVEFloatArithmeticPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zn needs to equal zd");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'1000'0000'0000'0000;
@@ -3728,9 +3722,7 @@ private:
}
void SVEFPRecursiveReduction(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "FP reduction operation can "
"only use 16-bit, 32-bit, "
"or 64-bit element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "FP reduction operation can only use 16/32/64-bit element sizes");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "FP reduction operation can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'0010'0000'0000'0000;
@@ -4112,7 +4104,7 @@ private:
// 0b111 - I - Current
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'0000'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4160,7 +4152,7 @@ private:
const auto& op_data = mem_op.MetaType.ScalarVectorType;
const bool is_scaled = op_data.scale != 0;
[[maybe_unused]] const auto msize_value = FEXCore::ToUnderlying(msize);
const auto msize_value = FEXCore::ToUnderlying(msize);
LOGMAN_THROW_A_FMT(op_data.scale == 0 || op_data.scale == msize_value, "scale may only be 0 or {}", msize_value);
@@ -4274,7 +4266,7 @@ private:
const auto msize_value = FEXCore::ToUnderlying(msize);
const auto msize_bytes = 1U << msize_value;
[[maybe_unused]] const auto imm_limit = (32U << msize_value) - msize_bytes;
const auto imm_limit = (32U << msize_value) - msize_bytes;
const auto imm = mem_op.MetaType.VectorImmType.Imm;
const auto imm_to_encode = imm >> msize_value;
@@ -4340,8 +4332,8 @@ private:
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT((imm % num_regs) == 0, "Offset must be a multiple of {}", num_regs);
[[maybe_unused]] const auto min_offset = -8 * num_regs;
[[maybe_unused]] const auto max_offset = 7 * num_regs;
const auto min_offset = -8 * num_regs;
const auto max_offset = 7 * num_regs;
LOGMAN_THROW_A_FMT(imm >= min_offset && imm <= max_offset,
"Invalid load/store offset ({}). Offset must be a multiple of {} and be within [{}, {}]", imm, num_regs, min_offset,
max_offset);
@@ -4448,8 +4440,8 @@ private:
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
const auto esize = static_cast<int>(16 << ssz);
[[maybe_unused]] const auto max_imm = (esize << 3) - esize;
[[maybe_unused]] const auto min_imm = -(max_imm + esize);
const auto max_imm = (esize << 3) - esize;
const auto min_imm = -(max_imm + esize);
LOGMAN_THROW_A_FMT((imm % esize) == 0, "imm ({}) must be a multiple of {}", imm, esize);
LOGMAN_THROW_A_FMT(imm >= min_imm && imm <= max_imm, "imm ({}) must be within [{}, {}]", imm, min_imm, max_imm);
@@ -4493,7 +4485,7 @@ private:
const auto msize_value = FEXCore::ToUnderlying(msize);
const auto data_size_bytes = 1U << msize_value;
[[maybe_unused]] const auto max_imm = (64U << msize_value) - data_size_bytes;
const auto max_imm = (64U << msize_value) - data_size_bytes;
LOGMAN_THROW_A_FMT((imm % data_size_bytes) == 0 && imm <= max_imm, "imm must be a multiple of {} and be within [0, {}]",
data_size_bytes, max_imm);
@@ -4721,7 +4713,7 @@ private:
void SVEFloatUnary(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'1100'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4809,8 +4801,7 @@ private:
}
void SVEFPUnaryOpsUnpredicated(uint32_t opc, SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
uint32_t Instr = 0b0110'0101'0000'1000'0011'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4821,8 +4812,7 @@ private:
}
void SVEFPSerialReductionPredicated(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, VRegister vn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(vd == vn, "vn must be the same as vd");
@@ -4836,8 +4826,7 @@ private:
}
void SVEFPCompareWithZero(uint32_t eqlt, uint32_t ne, SubRegSize size, PRegister pd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0001'0000'0010'0000'0000'0000;
@@ -4852,8 +4841,7 @@ private:
void SVEFPMultiplyAdd(uint32_t opc, SubRegSize size, ZRegister zd, PRegister pg, ZRegister zn, ZRegister zm) {
// NOTE: opc also includes the op0 bit (bit 15) like op0:opc, since the fields are adjacent
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0010'0000'0000'0000'0000'0000;
@@ -4867,14 +4855,13 @@ private:
}
void SVEFPMultiplyAddIndexed(uint32_t op, SubRegSize size, ZRegister zda, ZRegister zn, ZRegister zm, uint32_t index) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT((size <= SubRegSize::i32Bit && zm <= ZReg::z7) || (size == SubRegSize::i64Bit && zm <= ZReg::z15),
"16-bit and 32-bit indexed variants may only use Zm between z0-z7\n"
"64-bit variants may only use Zm between z0-z15");
const auto Underlying = FEXCore::ToUnderlying(size);
[[maybe_unused]] const uint32_t IndexMax = (16 / (1U << Underlying)) - 1;
const uint32_t IndexMax = (16 / (1U << Underlying)) - 1;
LOGMAN_THROW_A_FMT(index <= IndexMax, "Index must be within 0-{}", IndexMax);
// Can be bit 20 or 19 depending on whether or not the element size is 64-bit.
@@ -5130,12 +5117,13 @@ private:
requires (std::is_same_v<T, float> || std::is_same_v<T, double>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Determines if a floating-point value is capable of being converted
// into an 8-bit immediate. See pseudocode definition of VFPExpandImm
// in ARM A-profile reference manual for a general overview of how this was derived.
template<typename T>
requires (std::is_same_v<T, float> || std::is_same_v<T, double>)
[[nodiscard, maybe_unused]]
[[nodiscard]]
static bool IsValidFPValueForImm8(T value) {
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
@@ -5175,10 +5163,13 @@ private:
return true;
}
#endif
protected:
static uint32_t FP32ToImm8(float value) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint32_t>(value);
const auto sign = (bits & 0x80000000) >> 24;
@@ -5189,7 +5180,9 @@ protected:
}
static uint32_t FP64ToImm8(double value) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint64_t>(value);
const auto sign = (bits & 0x80000000'00000000) >> 56;
@@ -5215,7 +5208,7 @@ private:
uint32_t shift = 0;
if (!is_int8_imm) {
const int32_t imm16_limit = 32768;
[[maybe_unused]] const bool is_int16_imm = -imm16_limit <= imm && imm < imm16_limit;
const bool is_int16_imm = -imm16_limit <= imm && imm < imm16_limit;
LOGMAN_THROW_A_FMT(is_int16_imm, "Immediate ({}) must be a 16-bit value within [-32768, 32512]", imm);
LOGMAN_THROW_A_FMT((imm % 256) == 0, "Immediate ({}) must be a multiple of 256", imm);
+7 -9
View File
@@ -27,21 +27,19 @@ struct EmitterOps : Emitter {
public:
// Advanced SIMD scalar copy
void dup(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
constexpr uint32_t Op = 0b0101'1110'0000'0000'0000'01 << 10;
const uint32_t SizeImm = FEXCore::ToUnderlying(size);
const uint32_t IndexShift = SizeImm + 1;
const uint32_t ElementSize = 1U << SizeImm;
[[maybe_unused]] const uint32_t MaxIndex = 128U / (ElementSize * 8);
const uint32_t MaxIndex = 128U / (ElementSize * 8);
LOGMAN_THROW_A_FMT(Index < MaxIndex, "Index too large. Index={}, Max Index: {}", Index, MaxIndex);
const uint32_t imm5 = (Index << IndexShift) | ElementSize;
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
ASIMDScalarCopy(1, 1, imm5, 0b0000, rd, rn);
}
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
void mov(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
dup(size, rd, rn, Index);
}
@@ -1282,10 +1280,10 @@ public:
private:
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
void ASIMDScalarCopy(uint32_t Q, uint32_t b28, uint32_t imm5, uint32_t imm4, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0000'1110'0000'0000'0000'01U << 10;
Instr |= Q << 30;
Instr |= b28 << 28;
Instr |= imm5 << 16;
Instr |= imm4 << 11;
Instr |= Encode_rn(rn);
@@ -1383,7 +1381,7 @@ private:
void ASIMDScalarXIndexedElement(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i8Bit, "Scalar size must not be 8-bit");
[[maybe_unused]] const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
LOGMAN_THROW_A_FMT(index < invalid_bound, "Index ({}) must be within [0-{}]", index, invalid_bound - 1);
uint32_t Instr = 0b0101'1111'0000'0000'0000'0000'0000'0000;
+11 -59
View File
@@ -41,7 +41,6 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
[[maybe_unused]] constexpr auto kDRegSize = 64;
constexpr auto kWRegSize = 32;
constexpr auto kXRegSize = 64;
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) || (width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
@@ -129,8 +128,8 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// Compute the repeat distance d, and set up a bitmask covering the basic
// unit of repetition (i.e. a word with the bottom d bits set). Also, in all
// of these cases the N bit of the output will be zero.
clz_a = CountLeadingZeros(a, kXRegSize);
int clz_c = CountLeadingZeros(c, kXRegSize);
clz_a = std::countl_zero(a);
int clz_c = std::countl_zero(c);
d = clz_a - clz_c;
mask = ((UINT64_C(1) << d) - 1);
out_n = 0;
@@ -151,7 +150,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// of set bits in our word, meaning that we have the trivial case of
// d == 64 and only one 'repetition'. Set up all the same variables as in
// the general case above, and set the N bit in the output.
clz_a = CountLeadingZeros(a, kXRegSize);
clz_a = std::countl_zero(a);
d = 64;
mask = ~UINT64_C(0);
out_n = 1;
@@ -159,7 +158,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
}
// If the repeat period d is not a power of two, it can't be encoded.
if (!IsPowerOf2(d)) {
if (!std::has_single_bit(uint32_t(d))) {
return false;
}
@@ -179,7 +178,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
static const uint64_t multipliers[] = {
0x0000000000000001UL, 0x0000000100000001UL, 0x0001000100010001UL, 0x0101010101010101UL, 0x1111111111111111UL, 0x5555555555555555UL,
};
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
uint64_t multiplier = multipliers[std::countl_zero(uint64_t(d)) - 57];
uint64_t candidate = (b - a) * multiplier;
if (value != candidate) {
@@ -194,7 +193,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// Count the set bits in our basic stretch. The special case of clz(0) == -1
// makes the answer come out right for stretches that reach the very top of
// the word (e.g. numbers like 0xffffc00000000000).
int clz_b = (b == 0) ? -1 : CountLeadingZeros(b, kXRegSize);
int clz_b = (b == 0) ? -1 : std::countl_zero(b);
int s = clz_a - clz_b;
// Decide how many bits to rotate right by, to put the low bit of that basic
@@ -224,9 +223,13 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// 11110s 2 UInt(s)
//
// So we 'or' (2 * -d) with our computed s to form imms.
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
if (n != nullptr) {
*n = out_n;
}
if (imm_s != nullptr) {
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
}
if (imm_r != nullptr) {
*imm_r = r;
}
@@ -281,11 +284,6 @@ INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
private:
template<typename V>
static inline bool IsPowerOf2(V value) {
return (value != 0) && ((value & (value - 1)) == 0);
}
// Some compilers dislike negating unsigned integers,
// so we provide an equivalent.
template<typename T>
@@ -298,50 +296,4 @@ static inline uint64_t LowestSetBit(uint64_t value) {
return value & UnsignedNegate(value);
}
template<typename V>
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
#if COMPILER_HAS_BUILTIN_CLZ
if (width == 32) {
return (value == 0) ? 32 : __builtin_clz(static_cast<unsigned>(value));
} else if (width == 64) {
return (value == 0) ? 64 : __builtin_clzll(value);
}
#endif
return CountLeadingZerosFallBack(value, width);
}
static inline int CountLeadingZerosFallBack(uint64_t value, int width) {
LOGMAN_THROW_A_FMT(IsPowerOf2(width) && (width <= 64), "Invalid width");
if (value == 0) {
return width;
}
int count = 0;
value = value << (64 - width);
if ((value & UINT64_C(0xffffffff00000000)) == 0) {
count += 32;
value = value << 32;
}
if ((value & UINT64_C(0xffff000000000000)) == 0) {
count += 16;
value = value << 16;
}
if ((value & UINT64_C(0xff00000000000000)) == 0) {
count += 8;
value = value << 8;
}
if ((value & UINT64_C(0xf000000000000000)) == 0) {
count += 4;
value = value << 4;
}
if ((value & UINT64_C(0xc000000000000000)) == 0) {
count += 2;
value = value << 2;
}
if ((value & UINT64_C(0x8000000000000000)) == 0) {
count += 1;
}
count += (value == 0);
return count;
}
public:
+1
View File
@@ -4,6 +4,7 @@ set(CMAKE_RC_COMPILER ${MINGW_TRIPLE}-windres)
set(CMAKE_C_COMPILER ${MINGW_TRIPLE}-clang)
set(CMAKE_CXX_COMPILER ${MINGW_TRIPLE}-clang++)
set(CMAKE_DLLTOOL ${MINGW_TRIPLE}-dlltool)
set(CMAKE_AR ${MINGW_TRIPLE}-ar)
# Compile everything as static to avoid requiring the MinGW runtime libraries, force page aligned sections so that
# debug symbols work correctly, and disable loop alignment to workaround an LLVM bug
+1
View File
@@ -0,0 +1 @@
DisableFormat: true
+3 -8
View File
@@ -169,14 +169,9 @@ View the diff from {self.name} here.
class ClangFormatHelper(FormatHelper):
name = "clang-format"
name = "git-clang-format"
friendly_name = "C/C++ code formatter"
@property
def cformat_wrapper_path(self) -> str:
relpath = "../../Scripts/clang-format.py"
curpath = os.path.dirname(os.path.abspath(__file__))
return os.path.abspath(os.path.normpath(os.path.join(curpath, relpath)))
@property
def instructions(self) -> str:
@@ -199,7 +194,7 @@ class ClangFormatHelper(FormatHelper):
def clang_fmt_path(self) -> str:
if "CLANG_FORMAT_PATH" in os.environ:
return os.environ["CLANG_FORMAT_PATH"]
return "git-clang-format"
return "git-clang-format-19"
def has_tool(self) -> bool:
cmd = [self.clang_fmt_path, "-h"]
@@ -217,7 +212,7 @@ class ClangFormatHelper(FormatHelper):
cf_cmd = [
self.clang_fmt_path,
f"--binary={self.cformat_wrapper_path}",
"--binary=clang-format-19",
"--diff",
]
+378 -38
View File
@@ -1,52 +1,392 @@
#
# This file is autogenerated by pip-compile with Python 3.11
# This file is autogenerated by pip-compile with Python 3.13
# by the following command:
#
# pip-compile --output-file=llvm/utils/git/requirements_formatting.txt llvm/utils/git/requirements_formatting.txt.in
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==23.9.1
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
# via
# -r llvm/utils/git/requirements_formatting.txt.in
# -r requirements_formatting.txt.in
# darker
certifi==2023.7.22
# via requests
cffi==1.15.1
certifi==2025.7.14 \
--hash=sha256:6b31f564a415d79ee77df69d757bb49a5bb53bd9f756cbbe24394ffd6fc1f4b2 \
--hash=sha256:8ea99dbdfaaf2ba2f9bac77b9249ef62ec5218e7c2b2e903378ed5fccf765995
# via
# -r requirements_formatting.txt.in
# requests
cffi==1.15.1 \
--hash=sha256:00a9ed42e88df81ffae7a8ab6d9356b371399b91dbdf0c3cb1e84c03a13aceb5 \
--hash=sha256:03425bdae262c76aad70202debd780501fabeaca237cdfddc008987c0e0f59ef \
--hash=sha256:04ed324bda3cda42b9b695d51bb7d54b680b9719cfab04227cdd1e04e5de3104 \
--hash=sha256:0e2642fe3142e4cc4af0799748233ad6da94c62a8bec3a6648bf8ee68b1c7426 \
--hash=sha256:173379135477dc8cac4bc58f45db08ab45d228b3363adb7af79436135d028405 \
--hash=sha256:198caafb44239b60e252492445da556afafc7d1e3ab7a1fb3f0584ef6d742375 \
--hash=sha256:1e74c6b51a9ed6589199c787bf5f9875612ca4a8a0785fb2d4a84429badaf22a \
--hash=sha256:2012c72d854c2d03e45d06ae57f40d78e5770d252f195b93f581acf3ba44496e \
--hash=sha256:21157295583fe8943475029ed5abdcf71eb3911894724e360acff1d61c1d54bc \
--hash=sha256:2470043b93ff09bf8fb1d46d1cb756ce6132c54826661a32d4e4d132e1977adf \
--hash=sha256:285d29981935eb726a4399badae8f0ffdff4f5050eaa6d0cfc3f64b857b77185 \
--hash=sha256:30d78fbc8ebf9c92c9b7823ee18eb92f2e6ef79b45ac84db507f52fbe3ec4497 \
--hash=sha256:320dab6e7cb2eacdf0e658569d2575c4dad258c0fcc794f46215e1e39f90f2c3 \
--hash=sha256:33ab79603146aace82c2427da5ca6e58f2b3f2fb5da893ceac0c42218a40be35 \
--hash=sha256:3548db281cd7d2561c9ad9984681c95f7b0e38881201e157833a2342c30d5e8c \
--hash=sha256:3799aecf2e17cf585d977b780ce79ff0dc9b78d799fc694221ce814c2c19db83 \
--hash=sha256:39d39875251ca8f612b6f33e6b1195af86d1b3e60086068be9cc053aa4376e21 \
--hash=sha256:3b926aa83d1edb5aa5b427b4053dc420ec295a08e40911296b9eb1b6170f6cca \
--hash=sha256:3bcde07039e586f91b45c88f8583ea7cf7a0770df3a1649627bf598332cb6984 \
--hash=sha256:3d08afd128ddaa624a48cf2b859afef385b720bb4b43df214f85616922e6a5ac \
--hash=sha256:3eb6971dcff08619f8d91607cfc726518b6fa2a9eba42856be181c6d0d9515fd \
--hash=sha256:40f4774f5a9d4f5e344f31a32b5096977b5d48560c5592e2f3d2c4374bd543ee \
--hash=sha256:4289fc34b2f5316fbb762d75362931e351941fa95fa18789191b33fc4cf9504a \
--hash=sha256:470c103ae716238bbe698d67ad020e1db9d9dba34fa5a899b5e21577e6d52ed2 \
--hash=sha256:4f2c9f67e9821cad2e5f480bc8d83b8742896f1242dba247911072d4fa94c192 \
--hash=sha256:50a74364d85fd319352182ef59c5c790484a336f6db772c1a9231f1c3ed0cbd7 \
--hash=sha256:54a2db7b78338edd780e7ef7f9f6c442500fb0d41a5a4ea24fff1c929d5af585 \
--hash=sha256:5635bd9cb9731e6d4a1132a498dd34f764034a8ce60cef4f5319c0541159392f \
--hash=sha256:59c0b02d0a6c384d453fece7566d1c7e6b7bae4fc5874ef2ef46d56776d61c9e \
--hash=sha256:5d598b938678ebf3c67377cdd45e09d431369c3b1a5b331058c338e201f12b27 \
--hash=sha256:5df2768244d19ab7f60546d0c7c63ce1581f7af8b5de3eb3004b9b6fc8a9f84b \
--hash=sha256:5ef34d190326c3b1f822a5b7a45f6c4535e2f47ed06fec77d3d799c450b2651e \
--hash=sha256:6975a3fac6bc83c4a65c9f9fcab9e47019a11d3d2cf7f3c0d03431bf145a941e \
--hash=sha256:6c9a799e985904922a4d207a94eae35c78ebae90e128f0c4e521ce339396be9d \
--hash=sha256:70df4e3b545a17496c9b3f41f5115e69a4f2e77e94e1d2a8e1070bc0c38c8a3c \
--hash=sha256:7473e861101c9e72452f9bf8acb984947aa1661a7704553a9f6e4baa5ba64415 \
--hash=sha256:8102eaf27e1e448db915d08afa8b41d6c7ca7a04b7d73af6514df10a3e74bd82 \
--hash=sha256:87c450779d0914f2861b8526e035c5e6da0a3199d8f1add1a665e1cbc6fc6d02 \
--hash=sha256:8b7ee99e510d7b66cdb6c593f21c043c248537a32e0bedf02e01e9553a172314 \
--hash=sha256:91fc98adde3d7881af9b59ed0294046f3806221863722ba7d8d120c575314325 \
--hash=sha256:94411f22c3985acaec6f83c6df553f2dbe17b698cc7f8ae751ff2237d96b9e3c \
--hash=sha256:98d85c6a2bef81588d9227dde12db8a7f47f639f4a17c9ae08e773aa9c697bf3 \
--hash=sha256:9ad5db27f9cabae298d151c85cf2bad1d359a1b9c686a275df03385758e2f914 \
--hash=sha256:a0b71b1b8fbf2b96e41c4d990244165e2c9be83d54962a9a1d118fd8657d2045 \
--hash=sha256:a0f100c8912c114ff53e1202d0078b425bee3649ae34d7b070e9697f93c5d52d \
--hash=sha256:a591fe9e525846e4d154205572a029f653ada1a78b93697f3b5a8f1f2bc055b9 \
--hash=sha256:a5c84c68147988265e60416b57fc83425a78058853509c1b0629c180094904a5 \
--hash=sha256:a66d3508133af6e8548451b25058d5812812ec3798c886bf38ed24a98216fab2 \
--hash=sha256:a8c4917bd7ad33e8eb21e9a5bbba979b49d9a97acb3a803092cbc1133e20343c \
--hash=sha256:b3bbeb01c2b273cca1e1e0c5df57f12dce9a4dd331b4fa1635b8bec26350bde3 \
--hash=sha256:cba9d6b9a7d64d4bd46167096fc9d2f835e25d7e4c121fb2ddfc6528fb0413b2 \
--hash=sha256:cc4d65aeeaa04136a12677d3dd0b1c0c94dc43abac5860ab33cceb42b801c1e8 \
--hash=sha256:ce4bcc037df4fc5e3d184794f27bdaab018943698f4ca31630bc7f84a7b69c6d \
--hash=sha256:cec7d9412a9102bdc577382c3929b337320c4c4c4849f2c5cdd14d7368c5562d \
--hash=sha256:d400bfb9a37b1351253cb402671cea7e89bdecc294e8016a707f6d1d8ac934f9 \
--hash=sha256:d61f4695e6c866a23a21acab0509af1cdfd2c013cf256bbf5b6b5e2695827162 \
--hash=sha256:db0fbb9c62743ce59a9ff687eb5f4afbe77e5e8403d6697f7446e5f609976f76 \
--hash=sha256:dd86c085fae2efd48ac91dd7ccffcfc0571387fe1193d33b6394db7ef31fe2a4 \
--hash=sha256:e00b098126fd45523dd056d2efba6c5a63b71ffe9f2bbe1a4fe1716e1d0c331e \
--hash=sha256:e229a521186c75c8ad9490854fd8bbdd9a0c9aa3a524326b55be83b54d4e0ad9 \
--hash=sha256:e263d77ee3dd201c3a142934a086a4450861778baaeeb45db4591ef65550b0a6 \
--hash=sha256:ed9cb427ba5504c1dc15ede7d516b84757c3e3d7868ccc85121d9310d27eed0b \
--hash=sha256:fa6693661a4c91757f4412306191b6dc88c1703f780c8234035eac011922bc01 \
--hash=sha256:fcd131dd944808b5bdb38e6f5b53013c5aa4f334c5cad0c72742f6eba4b73db0
# via
# cryptography
# pynacl
charset-normalizer==3.2.0
charset-normalizer==3.2.0 \
--hash=sha256:04e57ab9fbf9607b77f7d057974694b4f6b142da9ed4a199859d9d4d5c63fe96 \
--hash=sha256:09393e1b2a9461950b1c9a45d5fd251dc7c6f228acab64da1c9c0165d9c7765c \
--hash=sha256:0b87549028f680ca955556e3bd57013ab47474c3124dc069faa0b6545b6c9710 \
--hash=sha256:1000fba1057b92a65daec275aec30586c3de2401ccdcd41f8a5c1e2c87078706 \
--hash=sha256:1249cbbf3d3b04902ff081ffbb33ce3377fa6e4c7356f759f3cd076cc138d020 \
--hash=sha256:1920d4ff15ce893210c1f0c0e9d19bfbecb7983c76b33f046c13a8ffbd570252 \
--hash=sha256:193cbc708ea3aca45e7221ae58f0fd63f933753a9bfb498a3b474878f12caaad \
--hash=sha256:1a100c6d595a7f316f1b6f01d20815d916e75ff98c27a01ae817439ea7726329 \
--hash=sha256:1f30b48dd7fa1474554b0b0f3fdfdd4c13b5c737a3c6284d3cdc424ec0ffff3a \
--hash=sha256:203f0c8871d5a7987be20c72442488a0b8cfd0f43b7973771640fc593f56321f \
--hash=sha256:246de67b99b6851627d945db38147d1b209a899311b1305dd84916f2b88526c6 \
--hash=sha256:2dee8e57f052ef5353cf608e0b4c871aee320dd1b87d351c28764fc0ca55f9f4 \
--hash=sha256:2efb1bd13885392adfda4614c33d3b68dee4921fd0ac1d3988f8cbb7d589e72a \
--hash=sha256:2f4ac36d8e2b4cc1aa71df3dd84ff8efbe3bfb97ac41242fbcfc053c67434f46 \
--hash=sha256:3170c9399da12c9dc66366e9d14da8bf7147e1e9d9ea566067bbce7bb74bd9c2 \
--hash=sha256:3b1613dd5aee995ec6d4c69f00378bbd07614702a315a2cf6c1d21461fe17c23 \
--hash=sha256:3bb3d25a8e6c0aedd251753a79ae98a093c7e7b471faa3aa9a93a81431987ace \
--hash=sha256:3bb7fda7260735efe66d5107fb7e6af6a7c04c7fce9b2514e04b7a74b06bf5dd \
--hash=sha256:41b25eaa7d15909cf3ac4c96088c1f266a9a93ec44f87f1d13d4a0e86c81b982 \
--hash=sha256:45de3f87179c1823e6d9e32156fb14c1927fcc9aba21433f088fdfb555b77c10 \
--hash=sha256:46fb8c61d794b78ec7134a715a3e564aafc8f6b5e338417cb19fe9f57a5a9bf2 \
--hash=sha256:48021783bdf96e3d6de03a6e39a1171ed5bd7e8bb93fc84cc649d11490f87cea \
--hash=sha256:4957669ef390f0e6719db3613ab3a7631e68424604a7b448f079bee145da6e09 \
--hash=sha256:5e86d77b090dbddbe78867a0275cb4df08ea195e660f1f7f13435a4649e954e5 \
--hash=sha256:6339d047dab2780cc6220f46306628e04d9750f02f983ddb37439ca47ced7149 \
--hash=sha256:681eb3d7e02e3c3655d1b16059fbfb605ac464c834a0c629048a30fad2b27489 \
--hash=sha256:6c409c0deba34f147f77efaa67b8e4bb83d2f11c8806405f76397ae5b8c0d1c9 \
--hash=sha256:7095f6fbfaa55defb6b733cfeb14efaae7a29f0b59d8cf213be4e7ca0b857b80 \
--hash=sha256:70c610f6cbe4b9fce272c407dd9d07e33e6bf7b4aa1b7ffb6f6ded8e634e3592 \
--hash=sha256:72814c01533f51d68702802d74f77ea026b5ec52793c791e2da806a3844a46c3 \
--hash=sha256:7a4826ad2bd6b07ca615c74ab91f32f6c96d08f6fcc3902ceeedaec8cdc3bcd6 \
--hash=sha256:7c70087bfee18a42b4040bb9ec1ca15a08242cf5867c58726530bdf3945672ed \
--hash=sha256:855eafa5d5a2034b4621c74925d89c5efef61418570e5ef9b37717d9c796419c \
--hash=sha256:8700f06d0ce6f128de3ccdbc1acaea1ee264d2caa9ca05daaf492fde7c2a7200 \
--hash=sha256:89f1b185a01fe560bc8ae5f619e924407efca2191b56ce749ec84982fc59a32a \
--hash=sha256:8b2c760cfc7042b27ebdb4a43a4453bd829a5742503599144d54a032c5dc7e9e \
--hash=sha256:8c2f5e83493748286002f9369f3e6607c565a6a90425a3a1fef5ae32a36d749d \
--hash=sha256:8e098148dd37b4ce3baca71fb394c81dc5d9c7728c95df695d2dca218edf40e6 \
--hash=sha256:94aea8eff76ee6d1cdacb07dd2123a68283cb5569e0250feab1240058f53b623 \
--hash=sha256:95eb302ff792e12aba9a8b8f8474ab229a83c103d74a750ec0bd1c1eea32e669 \
--hash=sha256:9bd9b3b31adcb054116447ea22caa61a285d92e94d710aa5ec97992ff5eb7cf3 \
--hash=sha256:9e608aafdb55eb9f255034709e20d5a83b6d60c054df0802fa9c9883d0a937aa \
--hash=sha256:a103b3a7069b62f5d4890ae1b8f0597618f628b286b03d4bc9195230b154bfa9 \
--hash=sha256:a386ebe437176aab38c041de1260cd3ea459c6ce5263594399880bbc398225b2 \
--hash=sha256:a38856a971c602f98472050165cea2cdc97709240373041b69030be15047691f \
--hash=sha256:a401b4598e5d3f4a9a811f3daf42ee2291790c7f9d74b18d75d6e21dda98a1a1 \
--hash=sha256:a7647ebdfb9682b7bb97e2a5e7cb6ae735b1c25008a70b906aecca294ee96cf4 \
--hash=sha256:aaf63899c94de41fe3cf934601b0f7ccb6b428c6e4eeb80da72c58eab077b19a \
--hash=sha256:b0dac0ff919ba34d4df1b6131f59ce95b08b9065233446be7e459f95554c0dc8 \
--hash=sha256:baacc6aee0b2ef6f3d308e197b5d7a81c0e70b06beae1f1fcacffdbd124fe0e3 \
--hash=sha256:bf420121d4c8dce6b889f0e8e4ec0ca34b7f40186203f06a946fa0276ba54029 \
--hash=sha256:c04a46716adde8d927adb9457bbe39cf473e1e2c2f5d0a16ceb837e5d841ad4f \
--hash=sha256:c0b21078a4b56965e2b12f247467b234734491897e99c1d51cee628da9786959 \
--hash=sha256:c1c76a1743432b4b60ab3358c937a3fe1341c828ae6194108a94c69028247f22 \
--hash=sha256:c4983bf937209c57240cff65906b18bb35e64ae872da6a0db937d7b4af845dd7 \
--hash=sha256:c4fb39a81950ec280984b3a44f5bd12819953dc5fa3a7e6fa7a80db5ee853952 \
--hash=sha256:c57921cda3a80d0f2b8aec7e25c8aa14479ea92b5b51b6876d975d925a2ea346 \
--hash=sha256:c8063cf17b19661471ecbdb3df1c84f24ad2e389e326ccaf89e3fb2484d8dd7e \
--hash=sha256:ccd16eb18a849fd8dcb23e23380e2f0a354e8daa0c984b8a732d9cfaba3a776d \
--hash=sha256:cd6dbe0238f7743d0efe563ab46294f54f9bc8f4b9bcf57c3c666cc5bc9d1299 \
--hash=sha256:d62e51710986674142526ab9f78663ca2b0726066ae26b78b22e0f5e571238dd \
--hash=sha256:db901e2ac34c931d73054d9797383d0f8009991e723dab15109740a63e7f902a \
--hash=sha256:e03b8895a6990c9ab2cdcd0f2fe44088ca1c65ae592b8f795c3294af00a461c3 \
--hash=sha256:e1c8a2f4c69e08e89632defbfabec2feb8a8d99edc9f89ce33c4b9e36ab63037 \
--hash=sha256:e4b749b9cc6ee664a3300bb3a273c1ca8068c46be705b6c31cf5d276f8628a94 \
--hash=sha256:e6a5bf2cba5ae1bb80b154ed68a3cfa2fa00fde979a7f50d6598d3e17d9ac20c \
--hash=sha256:e857a2232ba53ae940d3456f7533ce6ca98b81917d47adc3c7fd55dad8fab858 \
--hash=sha256:ee4006268ed33370957f55bf2e6f4d263eaf4dc3cfc473d1d90baff6ed36ce4a \
--hash=sha256:eef9df1eefada2c09a5e7a40991b9fc6ac6ef20b1372abd48d2794a316dc0449 \
--hash=sha256:f058f6963fd82eb143c692cecdc89e075fa0828db2e5b291070485390b2f1c9c \
--hash=sha256:f25c229a6ba38a35ae6e25ca1264621cc25d4d38dca2942a7fce0b67a4efe918 \
--hash=sha256:f2a1d0fd4242bd8643ce6f98927cf9c04540af6efa92323e9d3124f57727bfc1 \
--hash=sha256:f7560358a6811e52e9c4d142d497f1a6e10103d3a6881f18d04dbce3729c0e2c \
--hash=sha256:f779d3ad205f108d14e99bb3859aa7dd8e9c68874617c72354d7ecaec2a054ac \
--hash=sha256:f87f746ee241d30d6ed93969de31e5ffd09a2961a051e60ae6bddde9ec3583aa
# via requests
click==8.1.7
click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==41.0.3
# via pyjwt
darker==1.7.2
# via -r llvm/utils/git/requirements_formatting.txt.in
deprecated==1.2.14
cryptography==45.0.5 \
--hash=sha256:0027d566d65a38497bc37e0dd7c2f8ceda73597d2ac9ba93810204f56f52ebc7 \
--hash=sha256:101ee65078f6dd3e5a028d4f19c07ffa4dd22cce6a20eaa160f8b5219911e7d8 \
--hash=sha256:12e55281d993a793b0e883066f590c1ae1e802e3acb67f8b442e721e475e6463 \
--hash=sha256:14d96584701a887763384f3c47f0ca7c1cce322aa1c31172680eb596b890ec30 \
--hash=sha256:1e1da5accc0c750056c556a93c3e9cb828970206c68867712ca5805e46dc806f \
--hash=sha256:206210d03c1193f4e1ff681d22885181d47efa1ab3018766a7b32a7b3d6e6afd \
--hash=sha256:2089cc8f70a6e454601525e5bf2779e665d7865af002a5dec8d14e561002e135 \
--hash=sha256:3a264aae5f7fbb089dbc01e0242d3b67dffe3e6292e1f5182122bdf58e65215d \
--hash=sha256:3af26738f2db354aafe492fb3869e955b12b2ef2e16908c8b9cb928128d42c57 \
--hash=sha256:3fcfbefc4a7f332dece7272a88e410f611e79458fab97b5efe14e54fe476f4fd \
--hash=sha256:460f8c39ba66af7db0545a8c6f2eabcbc5a5528fc1cf6c3fa9a1e44cec33385e \
--hash=sha256:57c816dfbd1659a367831baca4b775b2a5b43c003daf52e9d57e1d30bc2e1b0e \
--hash=sha256:5aa1e32983d4443e310f726ee4b071ab7569f58eedfdd65e9675484a4eb67bd1 \
--hash=sha256:6ff8728d8d890b3dda5765276d1bc6fb099252915a2cd3aff960c4c195745dd0 \
--hash=sha256:7259038202a47fdecee7e62e0fd0b0738b6daa335354396c6ddebdbe1206af2a \
--hash=sha256:72e76caa004ab63accdf26023fccd1d087f6d90ec6048ff33ad0445abf7f605a \
--hash=sha256:7760c1c2e1a7084153a0f68fab76e754083b126a47d0117c9ed15e69e2103492 \
--hash=sha256:8c4a6ff8a30e9e3d38ac0539e9a9e02540ab3f827a3394f8852432f6b0ea152e \
--hash=sha256:9024beb59aca9d31d36fcdc1604dd9bbeed0a55bface9f1908df19178e2f116e \
--hash=sha256:90cb0a7bb35959f37e23303b7eed0a32280510030daba3f7fdfbb65defde6a97 \
--hash=sha256:91098f02ca81579c85f66df8a588c78f331ca19089763d733e34ad359f474174 \
--hash=sha256:926c3ea71a6043921050eaa639137e13dbe7b4ab25800932a8498364fc1abec9 \
--hash=sha256:982518cd64c54fcada9d7e5cf28eabd3ee76bd03ab18e08a48cad7e8b6f31b18 \
--hash=sha256:9b4cf6318915dccfe218e69bbec417fdd7c7185aa7aab139a2c0beb7468c89f0 \
--hash=sha256:ad0caded895a00261a5b4aa9af828baede54638754b51955a0ac75576b831b27 \
--hash=sha256:b85980d1e345fe769cfc57c57db2b59cff5464ee0c045d52c0df087e926fbe63 \
--hash=sha256:b8fa8b0a35a9982a3c60ec79905ba5bb090fc0b9addcfd3dc2dd04267e45f25e \
--hash=sha256:b9e38e0a83cd51e07f5a48ff9691cae95a79bea28fe4ded168a8e5c6c77e819d \
--hash=sha256:bd4c45986472694e5121084c6ebbd112aa919a25e783b87eb95953c9573906d6 \
--hash=sha256:be97d3a19c16a9be00edf79dca949c8fa7eff621763666a145f9f9535a5d7f42 \
--hash=sha256:c648025b6840fe62e57107e0a25f604db740e728bd67da4f6f060f03017d5097 \
--hash=sha256:d05a38884db2ba215218745f0781775806bde4f32e07b135348355fe8e4991d9 \
--hash=sha256:dd420e577921c8c2d31289536c386aaa30140b473835e97f83bc71ea9d2baf2d \
--hash=sha256:e357286c1b76403dd384d938f93c46b2b058ed4dfcdce64a770f0537ed3feb6f \
--hash=sha256:e6c00130ed423201c5bc5544c23359141660b07999ad82e34e7bb8f882bb78e0 \
--hash=sha256:e74d30ec9c7cb2f404af331d5b4099a9b322a8a6b25c4632755c8757345baac5 \
--hash=sha256:f3562c2f23c612f2e4a6964a61d942f891d29ee320edb62ff48ffb99f3de9ae8
# via
# -r requirements_formatting.txt.in
# pyjwt
darker==2.1.1 \
--hash=sha256:a6e6a682c0604e76fe9aec7650e96a944f517563c69b28fcc076db9d957d98ea \
--hash=sha256:ead701414c45359fc0312bc285614d3285fc135476d43f3bc08d989ee19d9020
# via -r requirements_formatting.txt.in
darkgraylib==1.2.1 \
--hash=sha256:60c59de69842367ce0c78c32c451fa8e9d29500e681312d9864a7416bcdb7792 \
--hash=sha256:a5dd6a2015a470d9047278cdd01a91ccb1d746675f8fd4562b3b5f6b8cbda930
# via
# darker
# graylint
deprecated==1.2.14 \
--hash=sha256:6fac8b097794a90302bdbb17b9b815e732d3c4720583ff1b198499d78470466c \
--hash=sha256:e5323eb936458dccc2582dc6f9c322c852a775a27065ff2b0c4970b9d53d01b3
# via pygithub
idna==3.4
# via requests
mypy-extensions==1.0.0
# via black
packaging==23.1
# via black
pathspec==0.11.2
# via black
platformdirs==3.10.0
# via black
pycparser==2.21
# via cffi
pygithub==1.59.1
# via -r llvm/utils/git/requirements_formatting.txt.in
pyjwt[crypto]==2.8.0
# via pygithub
pynacl==1.5.0
# via pygithub
requests==2.31.0
# via pygithub
toml==0.10.2
graylint==1.1.1 \
--hash=sha256:0fd8e02972ca03d0ef2bf0adea76b5343efcd492d7afb5f658f3e3a724f55a36 \
--hash=sha256:b7e0eab6c159684dbf5ef84e942c3340f6a6549b02a3d11b1a1763cc4f8f0593
# via darker
urllib3==2.0.4
# via requests
wrapt==1.15.0
idna==3.10 \
--hash=sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9 \
--hash=sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3
# via
# -r requirements_formatting.txt.in
# requests
mypy-extensions==1.0.0 \
--hash=sha256:4392f6c0eb8a5668a69e23d168ffa70f0be9ccfd32b5cc2d26a34ae5b844552d \
--hash=sha256:75dbf8955dc00442a438fc4d0666508a9a97b6bd41aa2f0ffe9d2f2725af0782
# via black
packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
--hash=sha256:d7c24979f292f916dc9cbf8648319032f551ea8c49a4c9bf2fb556a02070ec1d
# via black
pycparser==2.21 \
--hash=sha256:8ee45429555515e1f6b185e78100aea234072576aa43ab53aefcae078162fca9 \
--hash=sha256:e644fdec12f7872f86c58ff790da456218b10f863970249516d60a5eaca77206
# via cffi
pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pynacl==1.5.0 \
--hash=sha256:06b8f6fa7f5de8d5d2f7573fe8c863c051225a27b61e6860fd047b1775807858 \
--hash=sha256:0c84947a22519e013607c9be43706dd42513f9e6ae5d39d3613ca1e142fba44d \
--hash=sha256:20f42270d27e1b6a29f54032090b972d97f0a1b0948cc52392041ef7831fee93 \
--hash=sha256:401002a4aaa07c9414132aaed7f6836ff98f59277a234704ff66878c2ee4a0d1 \
--hash=sha256:52cb72a79269189d4e0dc537556f4740f7f0a9ec41c1322598799b0bdad4ef92 \
--hash=sha256:61f642bf2378713e2c2e1de73444a3778e5f0a38be6fee0fe532fe30060282ff \
--hash=sha256:8ac7448f09ab85811607bdd21ec2464495ac8b7c66d146bf545b0f08fb9220ba \
--hash=sha256:a36d4a9dda1f19ce6e03c9a784a2921a4b726b02e1c736600ca9c22029474394 \
--hash=sha256:a422368fc821589c228f4c49438a368831cb5bbc0eab5ebe1d7fac9dded6567b \
--hash=sha256:e46dae94e34b085175f8abb3b0aaa7da40767865ac82c928eeb9e57e1ea8a543
# via pygithub
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
# via
# -r requirements_formatting.txt.in
# pygithub
toml==0.10.2 \
--hash=sha256:806143ae5bfb6a3c6e736a764057db0e6a0e05e338b5630894a5f779cabb4f9b \
--hash=sha256:b3bda1d108d5dd99f4a20d24d9c348e91c4db7ab1b749200bded2f839ccbe68f
# via
# darker
# darkgraylib
typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.5.0 \
--hash=sha256:3fc47733c7e419d4bc3f6b3dc2b4f890bb743906a30d56ba4a5bfa4bbff92760 \
--hash=sha256:e6b01673c0fa6a13e374b50871808eb3bf7046c4b125b216f6bf1cc604cff0dc
# via
# -r requirements_formatting.txt.in
# pygithub
# requests
wrapt==1.15.0 \
--hash=sha256:02fce1852f755f44f95af51f69d22e45080102e9d00258053b79367d07af39c0 \
--hash=sha256:077ff0d1f9d9e4ce6476c1a924a3332452c1406e59d90a2cf24aeb29eeac9420 \
--hash=sha256:078e2a1a86544e644a68422f881c48b84fef6d18f8c7a957ffd3f2e0a74a0d4a \
--hash=sha256:0970ddb69bba00670e58955f8019bec4a42d1785db3faa043c33d81de2bf843c \
--hash=sha256:1286eb30261894e4c70d124d44b7fd07825340869945c79d05bda53a40caa079 \
--hash=sha256:21f6d9a0d5b3a207cdf7acf8e58d7d13d463e639f0c7e01d82cdb671e6cb7923 \
--hash=sha256:230ae493696a371f1dbffaad3dafbb742a4d27a0afd2b1aecebe52b740167e7f \
--hash=sha256:26458da5653aa5b3d8dc8b24192f574a58984c749401f98fff994d41d3f08da1 \
--hash=sha256:2cf56d0e237280baed46f0b5316661da892565ff58309d4d2ed7dba763d984b8 \
--hash=sha256:2e51de54d4fb8fb50d6ee8327f9828306a959ae394d3e01a1ba8b2f937747d86 \
--hash=sha256:2fbfbca668dd15b744418265a9607baa970c347eefd0db6a518aaf0cfbd153c0 \
--hash=sha256:38adf7198f8f154502883242f9fe7333ab05a5b02de7d83aa2d88ea621f13364 \
--hash=sha256:3a8564f283394634a7a7054b7983e47dbf39c07712d7b177b37e03f2467a024e \
--hash=sha256:3abbe948c3cbde2689370a262a8d04e32ec2dd4f27103669a45c6929bcdbfe7c \
--hash=sha256:3bbe623731d03b186b3d6b0d6f51865bf598587c38d6f7b0be2e27414f7f214e \
--hash=sha256:40737a081d7497efea35ab9304b829b857f21558acfc7b3272f908d33b0d9d4c \
--hash=sha256:41d07d029dd4157ae27beab04d22b8e261eddfc6ecd64ff7000b10dc8b3a5727 \
--hash=sha256:46ed616d5fb42f98630ed70c3529541408166c22cdfd4540b88d5f21006b0eff \
--hash=sha256:493d389a2b63c88ad56cdc35d0fa5752daac56ca755805b1b0c530f785767d5e \
--hash=sha256:4ff0d20f2e670800d3ed2b220d40984162089a6e2c9646fdb09b85e6f9a8fc29 \
--hash=sha256:54accd4b8bc202966bafafd16e69da9d5640ff92389d33d28555c5fd4f25ccb7 \
--hash=sha256:56374914b132c702aa9aa9959c550004b8847148f95e1b824772d453ac204a72 \
--hash=sha256:578383d740457fa790fdf85e6d346fda1416a40549fe8db08e5e9bd281c6a475 \
--hash=sha256:58d7a75d731e8c63614222bcb21dd992b4ab01a399f1f09dd82af17bbfc2368a \
--hash=sha256:5c5aa28df055697d7c37d2099a7bc09f559d5053c3349b1ad0c39000e611d317 \
--hash=sha256:5fc8e02f5984a55d2c653f5fea93531e9836abbd84342c1d1e17abc4a15084c2 \
--hash=sha256:63424c681923b9f3bfbc5e3205aafe790904053d42ddcc08542181a30a7a51bd \
--hash=sha256:64b1df0f83706b4ef4cfb4fb0e4c2669100fd7ecacfb59e091fad300d4e04640 \
--hash=sha256:74934ebd71950e3db69960a7da29204f89624dde411afbfb3b4858c1409b1e98 \
--hash=sha256:75669d77bb2c071333417617a235324a1618dba66f82a750362eccbe5b61d248 \
--hash=sha256:75760a47c06b5974aa5e01949bf7e66d2af4d08cb8c1d6516af5e39595397f5e \
--hash=sha256:76407ab327158c510f44ded207e2f76b657303e17cb7a572ffe2f5a8a48aa04d \
--hash=sha256:76e9c727a874b4856d11a32fb0b389afc61ce8aaf281ada613713ddeadd1cfec \
--hash=sha256:77d4c1b881076c3ba173484dfa53d3582c1c8ff1f914c6461ab70c8428b796c1 \
--hash=sha256:780c82a41dc493b62fc5884fb1d3a3b81106642c5c5c78d6a0d4cbe96d62ba7e \
--hash=sha256:7dc0713bf81287a00516ef43137273b23ee414fe41a3c14be10dd95ed98a2df9 \
--hash=sha256:7eebcdbe3677e58dd4c0e03b4f2cfa346ed4049687d839adad68cc38bb559c92 \
--hash=sha256:896689fddba4f23ef7c718279e42f8834041a21342d95e56922e1c10c0cc7afb \
--hash=sha256:96177eb5645b1c6985f5c11d03fc2dbda9ad24ec0f3a46dcce91445747e15094 \
--hash=sha256:96e25c8603a155559231c19c0349245eeb4ac0096fe3c1d0be5c47e075bd4f46 \
--hash=sha256:9d37ac69edc5614b90516807de32d08cb8e7b12260a285ee330955604ed9dd29 \
--hash=sha256:9ed6aa0726b9b60911f4aed8ec5b8dd7bf3491476015819f56473ffaef8959bd \
--hash=sha256:a487f72a25904e2b4bbc0817ce7a8de94363bd7e79890510174da9d901c38705 \
--hash=sha256:a4cbb9ff5795cd66f0066bdf5947f170f5d63a9274f99bdbca02fd973adcf2a8 \
--hash=sha256:a74d56552ddbde46c246b5b89199cb3fd182f9c346c784e1a93e4dc3f5ec9975 \
--hash=sha256:a89ce3fd220ff144bd9d54da333ec0de0399b52c9ac3d2ce34b569cf1a5748fb \
--hash=sha256:abd52a09d03adf9c763d706df707c343293d5d106aea53483e0ec8d9e310ad5e \
--hash=sha256:abd8f36c99512755b8456047b7be10372fca271bf1467a1caa88db991e7c421b \
--hash=sha256:af5bd9ccb188f6a5fdda9f1f09d9f4c86cc8a539bd48a0bfdc97723970348418 \
--hash=sha256:b02f21c1e2074943312d03d243ac4388319f2456576b2c6023041c4d57cd7019 \
--hash=sha256:b06fa97478a5f478fb05e1980980a7cdf2712015493b44d0c87606c1513ed5b1 \
--hash=sha256:b0724f05c396b0a4c36a3226c31648385deb6a65d8992644c12a4963c70326ba \
--hash=sha256:b130fe77361d6771ecf5a219d8e0817d61b236b7d8b37cc045172e574ed219e6 \
--hash=sha256:b56d5519e470d3f2fe4aa7585f0632b060d532d0696c5bdfb5e8319e1d0f69a2 \
--hash=sha256:b67b819628e3b748fd3c2192c15fb951f549d0f47c0449af0764d7647302fda3 \
--hash=sha256:ba1711cda2d30634a7e452fc79eabcadaffedf241ff206db2ee93dd2c89a60e7 \
--hash=sha256:bbeccb1aa40ab88cd29e6c7d8585582c99548f55f9b2581dfc5ba68c59a85752 \
--hash=sha256:bd84395aab8e4d36263cd1b9308cd504f6cf713b7d6d3ce25ea55670baec5416 \
--hash=sha256:c99f4309f5145b93eca6e35ac1a988f0dc0a7ccf9ccdcd78d3c0adf57224e62f \
--hash=sha256:ca1cccf838cd28d5a0883b342474c630ac48cac5df0ee6eacc9c7290f76b11c1 \
--hash=sha256:cd525e0e52a5ff16653a3fc9e3dd827981917d34996600bbc34c05d048ca35cc \
--hash=sha256:cdb4f085756c96a3af04e6eca7f08b1345e94b53af8921b25c72f096e704e145 \
--hash=sha256:ce42618f67741d4697684e501ef02f29e758a123aa2d669e2d964ff734ee00ee \
--hash=sha256:d06730c6aed78cee4126234cf2d071e01b44b915e725a6cb439a879ec9754a3a \
--hash=sha256:d5fe3e099cf07d0fb5a1e23d399e5d4d1ca3e6dfcbe5c8570ccff3e9208274f7 \
--hash=sha256:d6bcbfc99f55655c3d93feb7ef3800bd5bbe963a755687cbf1f490a71fb7794b \
--hash=sha256:d787272ed958a05b2c86311d3a4135d3c2aeea4fc655705f074130aa57d71653 \
--hash=sha256:e169e957c33576f47e21864cf3fc9ff47c223a4ebca8960079b8bd36cb014fd0 \
--hash=sha256:e20076a211cd6f9b44a6be58f7eeafa7ab5720eb796975d0c03f05b47d89eb90 \
--hash=sha256:e826aadda3cae59295b95343db8f3d965fb31059da7de01ee8d1c40a60398b29 \
--hash=sha256:eef4d64c650f33347c1f9266fa5ae001440b232ad9b98f1f43dfe7a79435c0a6 \
--hash=sha256:f2e69b3ed24544b0d3dbe2c5c0ba5153ce50dcebb576fdc4696d52aa22db6034 \
--hash=sha256:f87ec75864c37c4c6cb908d282e1969e79763e0d9becdfe9fe5473b7bb1e5f09 \
--hash=sha256:fbec11614dba0424ca72f4e8ba3c420dba07b4a7c206c8c8e4e73f2e98f4c559 \
--hash=sha256:fd69666217b62fa5d7c6aa88e507493a34dec4fa20c5bd925e4bc12fce586639
# via deprecated
@@ -0,0 +1,8 @@
black~=25.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=43.0.1
urllib3>=2.5.0
requests>=2.32.4
idna>=3.7
certifi>=2024.7.4
+1 -1
Vendored Submodule
+1
Submodule External/range-v3 added at ca1388fb9d.
+6 -6
View File
@@ -418,7 +418,7 @@ def print_parse_argloader_options(options):
if (value_type == "strenum"):
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key))
elif (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
@@ -447,13 +447,13 @@ def print_parse_envloader_options(options):
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key, op_key))
output_argloader.write("\tValue = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key))
output_argloader.write("}\n")
if ("ArgumentHandler" in op_vals):
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("\tValue = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
@@ -467,15 +467,15 @@ def print_parse_jsonloader_options(options):
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (KeyName == \"{0}\") {{\n".format(op_key))
output_argloader.write("\tSet(KeyOption, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View));\n".format(op_key, op_key, op_key))
output_argloader.write("\tSet(KeyOption, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View));\n".format(op_key, op_key))
output_argloader.write("}\n")
elif (value_type == "strarray"):
output_argloader.write("else if (KeyName == \"{0}\") {{\n".format(op_key))
output_argloader.write("\tAppendStrArrayValue(KeyOption, ConfigString);\n")
output_argloader.write("}\n")
assert op_key is not None, "No options found in JSONLOADER"
output_argloader.write("else {{\n".format(op_key))
output_argloader.write("Set(KeyOption, ConfigString);\n")
output_argloader.write("else {\n")
output_argloader.write("\tSet(KeyOption, ConfigString);\n")
output_argloader.write("}\n")
output_argloader.write("#endif\n")
+42 -15
View File
@@ -58,6 +58,7 @@ class OpDefinition:
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
Inline: list
Arguments: list
EmitValidation: list
Desc: list
@@ -278,6 +279,12 @@ def parse_ops(ops):
if "TiedSource" in op_val:
OpDef.TiedSource = op_val["TiedSource"]
# Pad Inline out to the argument count
OpDef.Inline = [''] * len(OpDef.Arguments)
if "Inline" in op_val:
Value = op_val["Inline"]
OpDef.Inline[0:len(Value)] = Value
# Do some fixups of the data here
if len(OpDef.EmitValidation) != 0:
for i in range(len(OpDef.EmitValidation)):
@@ -397,16 +404,16 @@ def print_ir_sizes():
// Make sure our array maps directly to the IROps enum
static_assert(IRSizes[IROps::OP_LAST] == -1ULL);
[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);
[[nodiscard]] inline size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const]] int8_t TiedSource(IROps Op);
#undef IROP_SIZES
#endif
@@ -588,13 +595,13 @@ def print_ir_arg_printer():
output_file.write("#endif\n")
def print_validation(op):
if op.EmitValidation != None:
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
if len(op.EmitValidation) != 0:
output_file.write("#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
Sanitized = Validation.replace("\"", "\\\"")
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("#endif\n")
# Print out IR allocator helpers
def print_ir_allocator_helpers():
@@ -773,9 +780,29 @@ def print_ir_allocator_helpers():
output_file.write(") {\n")
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
idx = 0
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
# Inline an immediate if we can
inline = op.Inline[idx]
idx += 1
if inline != '':
Sized = "Size" in [x.Name for x in op.Arguments]
P = ["Size" if Sized else "OpSize::i64Bit", arg.Name]
# A few cases need extra info plumbed.
if inline == "SubtractZero":
P += ["Src2"]
elif inline == "Mem":
P += ["OffsetType", "OffsetScale"]
elif inline == "Memtso":
P += ["OffsetType", "OffsetScale", "true /* TSO */"]
inline = "Mem"
output_file.write(f"\t\t{arg.Name} = Inline{inline}({', '.join(P)});\n")
output_file.write(f"\t\t{arg.Name}->AddUse();\n")
# Insert validation here. This is skipped for the
# OrderedNodeWrapper version because validation can depend on
-6
View File
@@ -1,4 +1,3 @@
include(GNUInstallDirs)
set (MAN_DIR share/man CACHE PATH "MAN_DIR")
set (FEXCORE_BASE_SRCS
@@ -24,9 +23,6 @@ set (SRCS
Interface/Core/Addressing.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
Interface/Core/OpcodeDispatcher/AVX_128.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
@@ -34,7 +30,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86Tables.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
@@ -66,7 +61,6 @@ set (SRCS
Interface/IR/IRDumper.cpp
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
+1 -3
View File
@@ -2,12 +2,10 @@
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <chrono>
#include <cstddef>
#include <cstdint>
#include <cstdio>
#include <memory>
#include <string_view>
namespace FEXCore {
+58 -2
View File
@@ -157,7 +157,30 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
/*
* FPREM is not an IEEE-754 remainder. From the spec:
* Check for invalid operation cases first - Intel FPREM sets Invalid Operation
* for several cases including infinity dividend and zero divisor.
*/
X80SoftFloat result = 0;
if (HandleInfinityOp(state, lhs, result)) {
return result;
} else if (lhs.Exponent == 0x7FFF && (lhs.Significand & 0x7FFFFFFFFFFFFFFFULL)) { // NaN
// propagate NaN
state->exceptionFlags |= softfloat_flag_invalid;
return lhs;
}
// Check for zero divisor - fprem(x, 0) is invalid operation
if (rhs.Exponent == 0 && rhs.Significand == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return result;
}
/*
* FPREM is not an IEEE-754 remainder. From the Intel spec:
*
* Computes the remainder obtained from dividing the value in the ST(0)
* register (the dividend) by the value in the ST(1) register (the divisor
@@ -271,7 +294,11 @@ struct FEX_PACKED X80SoftFloat {
FCMP(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(state, lhs, rhs);
*lt = extF80_lt(state, lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
// Use IEEE 754 semantics: unordered if neither <, =, nor > is true
// This is more reliable than custom NaN detection
bool gt = !(*eq) && !(*lt) && extF80_le(state, rhs, lhs);
*nan = !(*eq) && !(*lt) && !gt;
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
@@ -390,6 +417,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::tanl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -411,6 +443,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::sinl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -432,6 +469,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::cosl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -591,6 +633,20 @@ private:
static constexpr uint64_t IntegerBit = (1ULL << 63);
static constexpr uint64_t Bottom62Significand = ((1ULL << 62) - 1);
static constexpr uint32_t ExponentBias = 16383;
// Helper function to check for infinity and set invalid operation flag.
// Returns true if infinity is dealt with, false otherwise.
FEXCORE_PRESERVE_ALL_ATTR static bool HandleInfinityOp(softfloat_state* state, const X80SoftFloat& arg, X80SoftFloat& result) {
if (arg.Exponent == 0x7FFF && arg.Significand == 0x8000000000000000ULL) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN.
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return true;
}
return false;
}
};
#ifndef _WIN32
+11 -22
View File
@@ -7,69 +7,58 @@
#include <optional>
namespace FEXCore::StrConv {
[[maybe_unused]]
static bool Conv(std::string_view Value, bool* Result) {
inline bool Conv(std::string_view Value, bool* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint8_t* Result) {
inline bool Conv(std::string_view Value, uint8_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int8_t* Result) {
inline bool Conv(std::string_view Value, int8_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint16_t* Result) {
inline bool Conv(std::string_view Value, uint16_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int16_t* Result) {
inline bool Conv(std::string_view Value, int16_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint32_t* Result) {
inline bool Conv(std::string_view Value, uint32_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int32_t* Result) {
inline bool Conv(std::string_view Value, int32_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint64_t* Result) {
inline bool Conv(std::string_view Value, uint64_t* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int64_t* Result) {
inline bool Conv(std::string_view Value, int64_t* Result) {
*Result = std::strtoll(Value.data(), nullptr, 0);
return true;
}
template<typename T, typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]]
static bool Conv(std::string_view Value, T* Result) {
inline bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, fextl::string* Result) {
inline bool Conv(std::string_view Value, fextl::string* Result) {
*Result = Value;
return true;
}
+36 -36
View File
@@ -18,17 +18,6 @@
"Maximum number of instruction to store in a block"
]
},
"CacheObjectCodeCompilation": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE",
"TextDefault": "none",
"Choices": [ "none", "read", "readwrite" ],
"ArgumentHandler": "CacheObjectCodeHandler",
"Desc": [
"Cache JIT object code to drive.",
"Allows JIT code to be shared between applications"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -68,7 +57,11 @@
"ENABLESVEBITPERM": "enablesvebitperm",
"DISABLESVEBITPERM": "disablesvebitperm",
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi"
"DISABLEPRESERVEALLABI": "disablepreserveallabi",
"ENABLEWFXT": "enablewfxt",
"DISABLEWFXT": "disablewfxt",
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -89,7 +82,9 @@
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it"
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -470,32 +465,16 @@
"Desc": [
"Contrains the startup sleep to only apply to processes that match this name."
]
},
"MonoHacks": {
"Type": "bool",
"Default": "true",
"Desc": [
"Permits a hook-based SMC approach and smaller JIT blocks when mono is detected."
]
}
},
"Misc": {
"AOTIRCapture": {
"Type": "bool",
"Default": "false",
"Desc": [
"Captures IR and generates an AOT IR cache.",
"Captures both the loaded executable and libraries it loads."
]
},
"AOTIRGenerate": {
"Type": "bool",
"Default": "false",
"Desc": [
"Scans file for executable code and generates an AOT IR cache.",
"Does not run the executable."
]
},
"AOTIRLoad": {
"Type": "bool",
"Default": "false",
"Desc": [
"Loads an AOT IR cache for the loaded executable."
]
},
"ServerSocketPath": {
"Type": "str",
"Default": "",
@@ -509,6 +488,27 @@
"Desc": [
"Disables inline syscalls in order to support seccomp handling"
]
},
"ExtendedVolatileMetadata": {
"Type": "str",
"Default": "",
"Desc": [
"Configuration provided volatile metadata. Only implemented for WoW64/arm64ec.",
"Limited in its use but can be handy.",
"Extends on top of what Microsoft has for volatile metadata, but also supported for WoW64.",
"Colon delimited modules, then semi-colon delimited instructions, then comma delimited ranges",
"Default disables TSO in the module, unless instructions overlap the range",
"<module>;<offset begin>-<offset-end>,...;<instruction offset to force TSO>,...:<another>",
"examples:",
" * Disable TSO for a full module: Just provide the module name:",
" `hl2_linux`",
" * Disable TSO for a part of the module:",
" `hl2_linux;<offset begin>-<offset-end>`",
" * Disable TSO for a part of the module, but enable TSO for some instructions within the module",
" `hl2_linux;<offset begin>-<offset-end>;<instruction offset>,<instruction offset>`",
" * Disable TSO for multiple modules",
" `hl2_linux:libsdl2.so`"
]
}
}
},
+2 -8
View File
@@ -8,18 +8,12 @@
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Core/Thunks.h>
#include "FEXCore/Debug/InternalThreadState.h"
#include <string.h>
#include <utility>
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
X86Tables::InitializeInfoTables(Mode);
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext(const FEXCore::HostFeatures& Features) {
return fextl::make_unique<FEXCore::Context::ContextImpl>(Features);
}
+38 -58
View File
@@ -2,11 +2,12 @@
#pragma once
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/AOTIR.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
@@ -33,10 +34,6 @@ namespace FEXCore {
class CodeLoader;
class ThunkHandler;
namespace CodeSerialize {
class CodeObjectSerializeService;
}
namespace CPU {
class Arm64JITCore;
class Dispatcher;
@@ -50,8 +47,6 @@ namespace HLE {
} // namespace FEXCore
namespace FEXCore::IR {
struct IRListCopy;
class IRListView;
namespace Validation {
class IRValidation;
}
@@ -59,8 +54,9 @@ namespace Validation {
namespace FEXCore::Context {
struct FEX_PACKED ExitFunctionLinkData {
uint64_t HostBranch;
uint64_t HostCode;
uint64_t GuestRIP;
int64_t CallerOffset;
};
struct CustomIRResult {
@@ -89,6 +85,7 @@ public:
bool IsAddressInCurrentBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, uint64_t Size) override;
bool IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* Thread) override;
uint64_t GetGuestBlockEntry(FEXCore::Core::InternalThreadState* Thread) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, const uint64_t* HostGPRs, uint64_t PSTATE) override;
@@ -147,26 +144,12 @@ public:
FEXCore::IR::AOTIRCacheEntry* LoadAOTIRCacheEntry(const fextl::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry* Entry) override;
void SetAOTIRLoader(AOTIRLoaderCBFn CacheReader) override {
IRCaptureCache.SetAOTIRLoader(std::move(CacheReader));
}
void SetAOTIRWriter(AOTIRWriterCBFn CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(std::move(CacheWriter));
}
void SetAOTIRRenamer(AOTIRRenamerCBFn CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(std::move(CacheRenamer));
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(AOTIRCodeFileWriterFn Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void FinalizeAOTIRCache() override {}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -187,6 +170,12 @@ public:
void RemoveForceTSOInformation(uint64_t Address, uint64_t Size) override;
void MarkMonoDetected() override {
MonoDetected = true;
}
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
@@ -211,9 +200,6 @@ public:
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRGenerate, AOTIRGENERATE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
@@ -222,12 +208,12 @@ public:
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
@@ -246,32 +232,17 @@ public:
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
return Fn(Frame, Record);
}
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
// (safe as the invalidation mutex is locked) and manually invalidates the modified range. Allowing SMC to be detected
// even if faulting is disabled.
static void MonoBackpatcherWrite(FEXCore::Core::CpuStateFrame* Frame, uint8_t Size, uint64_t Address, uint64_t Value);
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
// NOTE: Other threads sharing the same CodeBuffer may reference
// invalidated data ranges through their L1/L2 caches. This is
// not currently a problem since FEX does not repurpose the
// invalidated CodeBuffer memory range currently.
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
void RemoveCustomIREntrypoint(FEXCore::Core::InternalThreadState* Thread, uintptr_t Entrypoint);
struct GenerateIRResult {
std::optional<IR::IRListView> IRView;
@@ -279,25 +250,23 @@ public:
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
bool NeedsAddGuestCodeRanges;
};
[[nodiscard]]
GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst);
struct CompileCodeResult {
void* CompiledCode;
CPU::CPUBackend::CompiledCode CompiledCode;
fextl::unique_ptr<FEXCore::Core::DebugData> DebugData;
uint64_t StartAddr;
uint64_t Length;
bool NeedsAddGuestCodeRanges;
};
[[nodiscard]]
CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
IR::OpSize GetGPROpSize() const {
return Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
FEXCore::JITSymbols Symbols;
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
@@ -332,6 +301,10 @@ public:
return ExitOnHLT;
}
bool AreMonoHacksActive() const {
return Config.MonoHacks && MonoDetected;
}
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -364,7 +337,6 @@ private:
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
@@ -377,8 +349,16 @@ private:
std::shared_mutex CustomIRMutex;
std::atomic<bool> HasCustomIRHandlers {};
fextl::unordered_map<uint64_t, std::tuple<CustomIREntrypointHandler, void*, void*>> CustomIRHandlers;
struct CustomIRHandlerEntry final {
CustomIREntrypointHandler Handler;
void *Creator;
void *Data;
};
fextl::unordered_map<uint64_t, CustomIRHandlerEntry> CustomIRHandlers;
IntervalList<uint64_t> ForceTSOValidRanges; // The ranges for which ForceTSOInstructions has populated data
fextl::set<uint64_t> ForceTSOInstructions;
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
};
} // namespace FEXCore::Context
+24 -38
View File
@@ -11,8 +11,7 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
Ref Tmp = A.Base;
if (A.Offset) {
Ref Offset = IREmit->_Constant(A.Offset);
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, Offset) : Offset;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Offset) : IREmit->Constant(A.Offset);
}
if (A.Index) {
@@ -22,10 +21,10 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
if (Tmp) {
Tmp = IREmit->_AddShift(GPRSize, Tmp, A.Index, ShiftType::LSL, Log2);
} else {
Tmp = IREmit->_Lshl(GPRSize, A.Index, IREmit->_Constant(Log2));
Tmp = IREmit->_Lshl(GPRSize, A.Index, IREmit->Constant(Log2));
}
} else {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Index) : A.Index;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Index) : A.Index;
}
}
@@ -41,32 +40,28 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
} else if (A.Offset) {
uint64_t X = A.Offset;
X &= (1ull << Bits) - 1;
Tmp = IREmit->_Constant(X);
Tmp = IREmit->Constant(X);
}
}
if (A.Segment && AddSegmentBase) {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Segment) : A.Segment;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Segment) : A.Segment;
}
return Tmp ?: IREmit->_Constant(0);
return Tmp ?: IREmit->Constant(0);
}
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
IR::OpSize AccessSize) {
auto SoftwareAddressCalculation = [IREmit, &A, GPRSize]() -> AddressMode {
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
};
const auto Is32Bit = GPRSize == OpSize::i32Bit;
const auto GPRSizeMatchesAddrSize = A.AddrSize == GPRSize;
const auto OffsetIndexToLargeFor32Bit = Is32Bit && (A.Offset <= -16384 || A.Offset >= 16384);
if (!GPRSizeMatchesAddrSize || OffsetIndexToLargeFor32Bit) {
// If address size doesn't match GPR size then no optimizations can occur.
return SoftwareAddressCalculation();
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
}
// Loadstore rules:
@@ -100,43 +95,31 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
const bool OffsetIsSIMM9 = A.Offset && A.Offset >= -256 && A.Offset <= 255;
const bool OffsetIsUnsignedScaled = A.Offset > 0 && (A.Offset & (AccessSizeAsImm - 1)) == 0 && (A.Offset / AccessSizeAsImm) <= 4095;
auto InlineImmOffsetLoadstore = [IREmit, &GPRSize](AddressMode A) -> AddressMode {
if ((AtomicTSO && !Vector && HostSupportsTSOImm9 && OffsetIsSIMM9) || (!AtomicTSO && (OffsetIsSIMM9 || OffsetIsUnsignedScaled))) {
// Peel off the offset
AddressMode B = A;
B.Offset = 0;
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->_Constant(A.Offset),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = 1,
};
};
}
auto ScaledRegisterLoadstore = [IREmit, GPRSize](AddressMode A) -> AddressMode {
if (AtomicTSO) {
// TODO: LRCPC3 support for vector Imm9.
} else if (!Is32Bit && A.Base && (A.Index || A.Segment) && !A.Offset && (A.IndexScale == 1 || A.IndexScale == AccessSizeAsImm)) {
// ScaledRegisterLoadstore
if (A.Index && A.Segment) {
A.Base = IREmit->_Add(GPRSize, A.Base, A.Segment);
A.Base = IREmit->Add(GPRSize, A.Base, A.Segment);
} else if (A.Segment) {
A.Index = A.Segment;
A.IndexScale = 1;
}
return A;
};
if (AtomicTSO) {
if (!Vector) {
if (HostSupportsTSOImm9 && OffsetIsSIMM9) {
return InlineImmOffsetLoadstore(A);
}
} else {
// TODO: LRCPC3 support for vector Imm9.
}
} else {
if (OffsetIsSIMM9 || OffsetIsUnsignedScaled) {
return InlineImmOffsetLoadstore(A);
} else if (!Is32Bit && A.Base && (A.Index || A.Segment) && !A.Offset && (A.IndexScale == 1 || A.IndexScale == AccessSizeAsImm)) {
return ScaledRegisterLoadstore(A);
}
return A;
}
if (Vector || !AtomicTSO) {
@@ -150,7 +133,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->_Constant(A.Offset),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = 1,
};
@@ -159,7 +142,10 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
}
// Fallback on software address calculation
return SoftwareAddressCalculation();
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
}
@@ -73,13 +73,13 @@ namespace x64 {
ARMEmitter::Reg::r8, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
};
constexpr std::array<ARMEmitter::Register, 8> RA = {
constexpr std::array<ARMEmitter::Register, 7> RA = {
// All these callee saved
ARMEmitter::Reg::r20, ARMEmitter::Reg::r21, ARMEmitter::Reg::r22, ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24, ARMEmitter::Reg::r25, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
ARMEmitter::Reg::r24, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
};
constexpr unsigned RAPairs = 6;
constexpr unsigned RAPairs = 4;
// Dynamic GPRs
constexpr std::array<ARMEmitter::Register, 2> PreserveAll_Dynamic = {
@@ -143,18 +143,18 @@ namespace x64 {
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r8,
};
constexpr std::array<ARMEmitter::Register, 7> RA = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15,
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
constexpr std::array<ARMEmitter::Register, 6> RA = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15, ARMEmitter::Reg::r16, ARMEmitter::Reg::r30,
};
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_Dynamic = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
};
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_Dynamic = {ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r16,
ARMEmitter::Reg::r17, ARMEmitter::Reg::r30};
constexpr std::array<ARMEmitter::Register, 7> NotPreserved_Dynamic = RA;
constexpr std::array<ARMEmitter::Register, 7> NotPreserved_Dynamic = {ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14,
ARMEmitter::Reg::r15, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
ARMEmitter::Reg::r30};
constexpr unsigned RAPairs = 6;
constexpr unsigned RAPairs = 4;
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
ARMEmitter::VReg::v0, ARMEmitter::VReg::v1, ARMEmitter::VReg::v2, ARMEmitter::VReg::v3,
@@ -245,14 +245,12 @@ namespace x32 {
REG_AF,
};
constexpr std::array<ARMEmitter::Register, 15> RA = {
constexpr std::array<ARMEmitter::Register, 14> RA = {
// All these callee saved
ARMEmitter::Reg::r20,
ARMEmitter::Reg::r21,
ARMEmitter::Reg::r22,
ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24,
ARMEmitter::Reg::r25,
// Registers only available on 32-bit
// All these are caller saved (except for r19).
@@ -265,6 +263,7 @@ namespace x32 {
ARMEmitter::Reg::r29,
ARMEmitter::Reg::r30,
ARMEmitter::Reg::r24,
ARMEmitter::Reg::r19,
};
@@ -273,7 +272,7 @@ namespace x32 {
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
};
constexpr unsigned RAPairs = 12;
constexpr unsigned RAPairs = 10;
// All are caller saved
constexpr std::array<ARMEmitter::VRegister, 8> SRAFPR = {
@@ -370,6 +369,8 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr
// Hardcode a 256-bit vector width if we are running in the simulator.
// Allow the user to override this.
Simulator.SetVectorLengthInBits(ForceSVEWidth() ? ForceSVEWidth() : 256);
// FEX doesn't support GCS.
Simulator.DisableGCSCheck();
#endif
#ifdef VIXL_DISASSEMBLER
// Only setup the disassembler if enabled.
@@ -495,7 +496,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
uint64_t AlignedPC = PC & ~0xFFFULL;
// Offset from aligned PC
int64_t AlignedOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(AlignedPC);
auto AlignedOffset = std::bit_cast<int64_t>(Constant - AlignedPC);
int NumMoves = 0;
@@ -511,7 +512,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
} else {
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
auto SmallOffset = std::bit_cast<int64_t>(Constant - PC);
if (ARMEmitter::Emitter::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
} else {
@@ -694,6 +695,8 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
@@ -709,7 +712,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
// Now handle PF/AF
if (PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
[[maybe_unused]] auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
LOGMAN_THROW_A_FMT(AFOffset == PFOffset + 4, "PF/AF are together");
@@ -789,6 +792,8 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(STATE, TmpReg, CPU_AREA_EMULATOR_DATA_OFFSET);
#endif
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
@@ -2,7 +2,7 @@
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "Interface/Core/JIT/Relocations.h"
#ifdef VIXL_DISASSEMBLER
#include <aarch64/disasm-aarch64.h>
@@ -43,6 +43,8 @@ constexpr bool TMP_ABIARGS = true;
constexpr auto REG_PF = ARMEmitter::Reg::r26;
constexpr auto REG_AF = ARMEmitter::Reg::r27;
constexpr auto REG_CALLRET_SP = ARMEmitter::XReg::x25;
// Vector temporaries
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
@@ -61,6 +63,8 @@ constexpr bool TMP_ABIARGS = false;
constexpr auto REG_PF = ARMEmitter::Reg::r9;
constexpr auto REG_AF = ARMEmitter::Reg::r24;
constexpr auto REG_CALLRET_SP = ARMEmitter::XReg::x17;
// Vector temporaries
constexpr auto VTMP1 = ARMEmitter::VReg::v16;
constexpr auto VTMP2 = ARMEmitter::VReg::v17;
@@ -84,7 +88,8 @@ constexpr uint64_t EC_CODE_BITMAP_MAX_ADDRESS = 1ULL << 47;
#endif
// Will force one single instruction block to be generated first if set when entering the JIT filling SRA.
constexpr auto ENTRY_FILL_SRA_SINGLE_INST_REG = TMP1;
// FillStaticRegs must preserve this
constexpr auto ENTRY_FILL_SRA_SINGLE_INST_REG = TMP2;
// Predicate to use in the X87 SVE optimization
constexpr ARMEmitter::PRegister PRED_X87_SVEOPT = ARMEmitter::PReg::p2;
+2 -3
View File
@@ -317,7 +317,7 @@ namespace CPU {
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer = CodeBuffers.StartLargerCodeBuffer();
RegisterForSignalHandler(PrevCodeBuffer);
RegisterForSignalHandler(std::move(PrevCodeBuffer));
return CurrentCodeBuffer.get();
}
@@ -326,14 +326,13 @@ namespace CPU {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Keep a reference to the old code buffer to delay deallocation
SignalHandlerCodeBuffers.push_back(CodeBuffer);
SignalHandlerCodeBuffers.push_back(std::move(CodeBuffer));
} else {
SignalHandlerCodeBuffers.clear();
}
}
fextl::shared_ptr<CodeBuffer> CPUBackend::CheckCodeBufferUpdate() {
fextl::shared_ptr<CodeBuffer> OldCodeBuffer;
auto NewCodeBuffer = CodeBuffers.GetLatest();
if (CurrentCodeBuffer != NewCodeBuffer) {
RegisterForSignalHandler(CurrentCodeBuffer);
+2 -9
View File
@@ -13,6 +13,7 @@ $end_info$
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/map.h>
#include <cstdint>
@@ -94,15 +95,7 @@ namespace CPU {
struct CompiledCode {
// Where this code block begins.
uint8_t* BlockBegin;
/**
* The function entrypoint to this codeblock.
*
* This may or may not equal `BlockBegin` above. Depending on the CPU backend, it may stick data
* prior to the BlockEntry.
*
* Is actually a function pointer of type `void (FEXCore::Core::ThreadState *Thread)`
*/
uint8_t* BlockEntry;
fextl::map<uint64_t, uint8_t*> EntryPoints;
// The total size of the codeblock from [BlockBegin, BlockBegin+Size).
size_t Size;
};
+62 -61
View File
@@ -626,6 +626,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
// Only enable EnhancedREPMOVS if atomic memcpy tso emulation isn't enabled.
const uint32_t SupportsEnhancedREPMOVS = CTX->IsMemcpyAtomicTSOEnabled() == false;
const uint32_t SupportsVPCLMULQDQ = CTX->HostFeatures.SupportsPMULL_128Bit && SupportsAVX();
const uint32_t SupportsWFXT = CTX->HostFeatures.SupportsWFXT;
// Number of subfunctions
Res.eax = 0x0;
@@ -645,39 +646,39 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(1 << 13) | // Deprecates FPU CS and DS
(0 << 14) | // Intel MPX
(0 << 15) | // Intel Resource Directory Technology Allocation
(0 << 16) | // Reserved
(0 << 17) | // Reserved
(0 << 16) | // AVX512-F
(0 << 17) | // AVX512-DQ
(CTX->HostFeatures.SupportsRAND << 18) | // RDSEED
(1 << 19) | // ADCX and ADOX instructions
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(0 << 21) | // AVX512-IFMA
(0 << 22) | // PCOMMIT (deprecated?)
(1 << 23) | // CLFLUSHOPT instruction
(1 << 24) | // CLWB instruction
(0 << 25) | // Intel processor trace
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 28) | // Reserved
(0 << 26) | // AVX512-PF
(0 << 27) | // AVX512-ER
(0 << 28) | // AVX512-CD
(Features.SHA << 29) | // SHA instructions
(0 << 30) | // Reserved
(0 << 31); // Reserved
(0 << 30) | // AVX512-BW
(0 << 31); // AVX512-VL
Res.ecx = (1 << 0) | // PREFETCHWT1
(0 << 1) | // AVX512VBMI
(0 << 2) | // Usermode instruction prevention
(0 << 3) | // Protection keys for user mode pages
(0 << 4) | // OS protection keys
(0 << 5) | // waitpkg
(0 << 6) | // AVX512_VBMI2
(SupportsWFXT << 5) | // waitpkg
(0 << 6) | // AVX512-VBMI2
(0 << 7) | // CET shadow stack
(0 << 8) | // GFNI
(CTX->HostFeatures.SupportsAES256 << 9) | // VAES
(SupportsVPCLMULQDQ << 10) | // VPCLMULQDQ
(0 << 11) | // AVX512_VNNI
(0 << 12) | // AVX512_BITALG
(0 << 11) | // AVX512-VNNI
(0 << 12) | // AVX512-BITALG
(0 << 13) | // Intel Total Memory Encryption
(0 << 14) | // AVX512_VPOPCNTDQ
(0 << 15) | // Reserved
(0 << 14) | // AVX512-VPOPCNTDQ
(0 << 15) | // FZM (TDX)
(0 << 16) | // 5 Level page tables
(0 << 17) | // MPX MAWAU
(0 << 18) | // MPX MAWAU
@@ -685,28 +686,28 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(0 << 23) | // Reserved
(0 << 24) | // Reserved
(0 << 23) | // AES Key Locker
(1 << 24) | // bus-lock-detect
(0 << 25) | // CLDEMOTE
(0 << 26) | // Reserved
(0 << 26) | // MPRR (TDX)
(0 << 27) | // MOVDIRI
(0 << 28) | // MOVDIR64B
(0 << 29) | // Reserved
(0 << 29) | // ENQCMD
(0 << 30) | // SGX Launch configuration
(0 << 31); // Reserved
(0 << 31); // PKS
Res.edx = (0 << 0) | // Reserved
(0 << 1) | // Reserved
(0 << 2) | // AVX512_4VNNIW
(0 << 3) | // AVX512_4FMAPS
Res.edx = (0 << 0) | // SGX-TEM (TDX)
(0 << 1) | // SGX-KEYS
(0 << 2) | // AVX512-4VNNIW
(0 << 3) | // AVX512-4FMAPS
(1 << 4) | // Fast Short Rep Mov
(0 << 5) | // Reserved
(0 << 5) | // UINTR
(0 << 6) | // Reserved
(0 << 7) | // Reserved
(0 << 8) | // AVX512_VP2INTERSECT
(0 << 8) | // AVX512-VP2INTERSECT
(0 << 9) | // SRBDS_CTRL (Special Register Buffer Data Sampling Mitigations)
(0 << 10) | // VERW clears CPU buffers
(0 << 11) | // Reserved
(0 << 11) | // rtm-always-abort
(0 << 12) | // Reserved
(0 << 13) | // TSX Force Abort (TSX will force abort if attempted)
(0 << 14) | // SERIALIZE instruction
@@ -718,7 +719,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 20) | // Intel CET
(0 << 21) | // Reserved
(0 << 22) | // AMX-BF16 - Tile computation on bfloat16
(0 << 23) | // AVX512_FP16 - FP16 AVX512 instructions
(0 << 23) | // AVX512-FP16 - FP16 AVX512 instructions
(0 << 24) | // AMX-tile - If AMX is implemented
(0 << 25) | // AMX-int8 - AMX on 8-bit integers
(0 << 26) | // IBRS_IBPB - Speculation control
@@ -754,7 +755,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
// XFeatureSupportedMask[63:32]
Res.edx = 0; // Upper 32-bits of XFeatureSupportedMask
} else if (Leaf == 1) {
Res.eax = (0 << 0) | // XSAVEOPT
Res.eax = (1 << 0) | // XSAVEOPT
(0 << 1) | // XSAVEC (and XRSTOR)
(0 << 2) | // XGETBV - XGETBV with ECX=1 supported
(0 << 3); // XSAVES - XSAVES, XRSTORS, and IA32_XSS supported
@@ -924,38 +925,38 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual mode extensions
(1 << 2) | // Debugging extensions
(1 << 3) | // Page size extensions
(1 << 4) | // TSC
(1 << 5) | // MSR support
(1 << 6) | // PAE
(1 << 7) | // Machine Check Exception
(1 << 8) | // CMPXCHG8B
(1 << 9) | // APIC
(0 << 10) | // Reserved
(1 << 11) | // SYSCALL/SYSRET
(1 << 12) | // MTRR
(1 << 13) | // Page global extension
(1 << 14) | // Machine Check architecture
(1 << 15) | // CMOV
(1 << 16) | // Page attribute table
(1 << 17) | // Page-size extensions
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(1 << 20) | // NX
(0 << 21) | // Reserved
(1 << 22) | // MMXExt
(1 << 23) | // MMX
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(1 << 30) | // 3DNow! Extensions
(1 << 31); // 3DNow!
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual mode extensions
(1 << 2) | // Debugging extensions
(1 << 3) | // Page size extensions
(1 << 4) | // TSC
(1 << 5) | // MSR support
(1 << 6) | // PAE
(1 << 7) | // Machine Check Exception
(1 << 8) | // CMPXCHG8B
(1 << 9) | // APIC
(0 << 10) | // Reserved
(1 << 11) | // SYSCALL/SYSRET
(1 << 12) | // MTRR
(1 << 13) | // Page global extension
(1 << 14) | // Machine Check architecture
(1 << 15) | // CMOV
(1 << 16) | // Page attribute table
(1 << 17) | // Page-size extensions
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(1 << 20) | // NX
(0 << 21) | // Reserved
(1 << 22) | // MMXExt
(1 << 23) | // MMX
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(CTX->HostFeatures.Supports3DNow << 30) | // 3DNow! Extensions
(CTX->HostFeatures.Supports3DNow << 31); // 3DNow!
return Res;
}
+151 -160
View File
@@ -14,7 +14,6 @@ $end_info$
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -46,13 +45,13 @@ $end_info$
#include "FEXCore/Utils/SignalScopeGuards.h"
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SHMStats.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TodoDefines.h>
#include <algorithm>
#include <array>
@@ -79,9 +78,6 @@ ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
: HostFeatures {Features}
, CPUID {this}
, IRCaptureCache {this} {
if (Config.CacheObjectCodeCompilation() != FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
CodeObjectCacheService = fextl::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
}
if (!Config.Is64BitMode()) {
// When operating in 32-bit mode, the virtual memory we care about is only the lower 32-bits.
Config.VirtualMemSize = 1ULL << 32;
@@ -105,14 +101,6 @@ ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
UpdateAtomicTSOEmulationConfig();
}
ContextImpl::~ContextImpl() {
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
}
}
}
struct GetFrameBlockInfoResult {
const CPU::CPUBackend::JITCodeHeader* InlineHeader;
const CPU::CPUBackend::JITCodeTail* InlineTail;
@@ -139,6 +127,11 @@ bool ContextImpl::IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* T
return InlineTail && InlineTail->SingleInst;
}
uint64_t ContextImpl::GetGuestBlockEntry(FEXCore::Core::InternalThreadState* Thread) {
auto [_, InlineTail] = GetFrameBlockInfo(Thread->CurrentFrame);
return InlineTail ? InlineTail->RIP : 0;
}
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
@@ -349,36 +342,9 @@ bool ContextImpl::InitCore() {
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
// Set up the SignalDelegator config since core is initialized.
FEXCore::SignalDelegator::SignalDelegatorConfig SignalConfig {
.DispatcherBegin = Dispatcher->Start,
.DispatcherEnd = Dispatcher->End,
SignalDelegation->SetConfig(Dispatcher->MakeSignalDelegatorConfig());
.AbsoluteLoopTopAddress = Dispatcher->AbsoluteLoopTopAddress,
.AbsoluteLoopTopAddressFillSRA = Dispatcher->AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = Dispatcher->SignalHandlerReturnAddressRT,
.PauseReturnInstruction = Dispatcher->PauseReturnInstruction,
.ThreadPauseHandlerAddressSpillSRA = Dispatcher->ThreadPauseHandlerAddressSpillSRA,
.ThreadPauseHandlerAddress = Dispatcher->ThreadPauseHandlerAddress,
// Stop handlers.
.ThreadStopHandlerAddressSpillSRA = Dispatcher->ThreadStopHandlerAddressSpillSRA,
.ThreadStopHandlerAddress = Dispatcher->ThreadStopHandlerAddress,
// SRA information.
.SRAGPRCount = Dispatcher->GetSRAGPRCount(),
.SRAFPRCount = Dispatcher->GetSRAFPRCount(),
};
Dispatcher->GetSRAGPRMapping(SignalConfig.SRAGPRMapping);
Dispatcher->GetSRAFPRMapping(SignalConfig.SRAFPRMapping);
// Give this configuration to the SignalDelegator.
SignalDelegation->SetConfig(SignalConfig);
#ifndef _WIN32
#elif !defined(_M_ARM_64EC)
#if defined(_WIN32) && !defined(_M_ARM_64EC)
// WOW64 always needs the interrupt fault check to be enabled.
Config.NeedsPendingInterruptFaultCheck = true;
#endif
@@ -398,21 +364,15 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
// TODO: This doesn't make sense when the parent thread doesn't outlive its children
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->CurrentFrame->Pointers.Common.L1Pointer = Thread->LookupCache->GetL1Pointer();
@@ -504,12 +464,6 @@ void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
if (NewCodeBuffer) {
// Allocate new CodeBuffer + L3 LookupCache and clear L1+L2 caches
Thread->CPUBackend->ClearCache();
@@ -517,6 +471,7 @@ void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, boo
// Clear L1+L2 cache of this thread, and clear L3 cache across any threads using it
Thread->LookupCache->ClearCache();
}
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter* IREmitter, uint64_t GuestRIP) {
@@ -545,7 +500,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
TotalInstructionsLength = 1;
std::get<0>(Handler->second)(GuestRIP, Thread->OpDispatcher.get());
Handler->second.Handler(GuestRIP, Thread->OpDispatcher.get());
HasCustomIR = true;
}
}
@@ -557,19 +512,15 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
bool HadDispatchError {false};
bool HadInvalidInst {false};
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, MaxInst,
[Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Thread, Start, Length);
}
});
Thread->FrontendDecoder->DecodeInstructionsAtEntry(Thread, GuestCode, GuestRIP, MaxInst);
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
auto CodeBlocks = &BlockInfo->Blocks;
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount);
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
AreMonoHacksActive() && MonoBackpatcherBlock.load(std::memory_order_relaxed) == GuestRIP);
const auto GPRSize = GetGPROpSize();
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
@@ -595,7 +546,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (InstsInBlock == 0) {
// Special case for an empty instruction block.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry - GuestRIP));
}
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -625,11 +576,12 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->_GuestOpcode(InstAddress - GuestRIP);
}
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(ExistingCodePtr[0], ExistingCodePtr[1],
(uintptr_t)ExistingCodePtr - GuestRIP, DecodedInfo->InstSize);
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL || Block.ForceFullSMCDetection) {
auto ExistingCodePtr = reinterpret_cast<uint8_t*>(Block.Entry + BlockInstructionsLength);
auto InstAddressReg = Thread->OpDispatcher->_EntrypointOffset(GPRSize, InstAddress - GuestRIP);
std::array<uint8_t, 0x10> CodeOriginal;
memcpy(CodeOriginal.data(), ExistingCodePtr, DecodedInfo->InstSize);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(CodeOriginal, InstAddressReg, DecodedInfo->InstSize);
auto InvalidateCodeCond = Thread->OpDispatcher->CondJump(CodeChanged);
@@ -639,7 +591,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, InstAddress - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -647,15 +599,21 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
}
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
if (TableInfo && TableInfo->OpcodeDispatcher.OpDispatch) {
auto Fn = TableInfo->OpcodeDispatcher.OpDispatch;
Thread->OpDispatcher->ResetHandledLock();
Thread->OpDispatcher->ResetDecodeFailure();
IR::ForceTSOMode ForceTSO =
BlockInForceTSOValidRange ?
(InstForceTSOIt != ForceTSOInstructions.end() && *InstForceTSOIt == InstAddress ? IR::ForceTSOMode::ForceEnabled :
IR::ForceTSOMode::ForceDisabled) :
IR::ForceTSOMode::NoOverride;
IR::ForceTSOMode ForceTSO = IR::ForceTSOMode::NoOverride;
if (BlockInForceTSOValidRange) {
if (InstForceTSOIt != ForceTSOInstructions.end() && *InstForceTSOIt == InstAddress) {
ForceTSO = IR::ForceTSOMode::ForceEnabled;
} else {
ForceTSO = IR::ForceTSOMode::ForceDisabled;
}
} else if (DecodedInfo->ForceTSO) {
ForceTSO = IR::ForceTSOMode::ForceEnabled;
}
Thread->OpDispatcher->SetForceTSO(ForceTSO);
std::invoke(Fn, Thread->OpDispatcher, DecodedInfo);
if (Thread->OpDispatcher->HadDecodeFailure()) {
@@ -683,7 +641,11 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
Thread->OpDispatcher->InvalidOp(DecodedInfo);
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::NOEXEC_INST) {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
} else {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
}
}
HadInvalidInst = true;
@@ -701,7 +663,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (NeedsBlockEnd) {
// We had some instructions. Early exit
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
break;
}
@@ -739,37 +702,23 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
.TotalInstructionsLength = TotalInstructionsLength,
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
.Length = Thread->FrontendDecoder->DecodedMaxAddress - Thread->FrontendDecoder->DecodedMinAddress,
.NeedsAddGuestCodeRanges = !HasCustomIR,
};
}
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
// JIT Code object cache lookup
if (CodeObjectCacheService) {
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
if (CodeCacheEntry) {
auto CompiledCode = Thread->CPUBackend->RelocateJITObjectCode(GuestRIP, CodeCacheEntry);
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.StartAddr = 0, // Unused
.Length = 0, // Unused
};
}
}
}
if (SourcecodeResolver && Config.GDBSymbols()) {
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (AOTIRCacheEntry.Entry && !AOTIRCacheEntry.Entry->ContainsCode) {
if (AOTIRCacheEntry.Entry) {
AOTIRCacheEntry.Entry->SourcecodeMap = SourcecodeResolver->GenerateMap(AOTIRCacheEntry.Entry->Filename, AOTIRCacheEntry.Entry->FileId);
}
}
// Generate IR + Meta Info
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length] = GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
return {nullptr, nullptr, 0, 0};
return {{}, nullptr, 0, 0, false};
}
// Attempt to get the CPU backend to compile this code
@@ -780,7 +729,11 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(GuestRIP)) {
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = reinterpret_cast<uint8_t*>(Block), .DebugData = nullptr, .StartAddr = 0, .Length = 0};
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
.StartAddr = 0,
.Length = 0,
.NeedsAddGuestCodeRanges = false};
}
}
@@ -795,13 +748,11 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
Thread->OpDispatcher->DelayedDisownBuffer();
return {
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = CompiledCode.BlockEntry,
.CompiledCode = std::move(CompiledCode),
.DebugData = std::move(DebugData),
.StartAddr = StartAddr,
.Length = Length,
.NeedsAddGuestCodeRanges = NeedsAddGuestCodeRanges,
};
}
@@ -821,7 +772,8 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return HostCode;
}
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
auto [CompiledCode, DebugData, StartAddr, Length, NeedsAddGuestCodeRanges] = CompileCode(Thread, GuestRIP, MaxInst);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
return 0;
} else if (!DebugData) {
@@ -831,7 +783,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
// The core managed to compile the code.
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = reinterpret_cast<uint8_t*>(CodePtr);
auto FragmentBasePtr = CompiledCode.BlockBegin;
if (DebugData) {
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
@@ -840,7 +792,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
for (auto& Subblock : DebugData->Subblocks) {
auto BlockBasePtr = FragmentBasePtr + Subblock.HostCodeOffset;
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename,
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, GuestRIP, Subblock.HostCodeSize);
@@ -848,41 +800,38 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
} else {
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename,
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, DebugData->HostCodeSize);
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, CompiledCode.Size);
}
}
}
}
// Tell the object cache service to serialize the code if enabled
if (CodeObjectCacheService && Config.CacheObjectCodeCompilation == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_READWRITE && DebugData) {
CodeObjectCacheService->AsyncAddSerializationJob(
fextl::make_unique<CodeSerialize::AsyncJobHandler::SerializationJobData>(CodeSerialize::AsyncJobHandler::SerializationJobData {
.GuestRIP = GuestRIP,
.GuestCodeLength = Length,
.GuestCodeHash = 0,
.HostCodeBegin = CodePtr,
.HostCodeLength = DebugData->HostCodeSize,
.HostCodeHash = 0,
.ThreadJobRefCount = &Thread->ObjectCacheRefCounter,
.Relocations = std::move(*DebugData->Relocations),
}));
}
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, {}, DebugData.get(), false)) {
if (IRCaptureCache.PostCompileCode(Thread, CompiledCode.BlockBegin, GuestRIP, StartAddr, Length, DebugData.get())) {
// Early exit
return (uintptr_t)CodePtr;
}
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
// Insert to lookup cache
// Pages containing this block are added via AddBlockExecutableRange before each page gets accessed in the frontend
Thread->LookupCache->AddBlockMapping(GuestRIP, CodePtr);
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(GuestAddr, HostAddr);
}
return (uintptr_t)CodePtr;
}
@@ -896,7 +845,8 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, 1);
auto [CompiledCode, DebugData, StartAddr, Length, _] = CompileCode(Thread, GuestRIP, 1);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
return 0;
}
@@ -907,22 +857,40 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
auto lk = Thread->LookupCache->AcquireLock();
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lower = Thread->LookupCache->CodePages.lower_bound(Start >> 12);
auto upper = Thread->LookupCache->CodePages.upper_bound((Start + Length - 1) >> 12);
auto lk = Thread->LookupCache->AcquireLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (auto Address : it->second) {
ContextImpl::ThreadRemoveCodeEntry(Thread, Address);
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry)) {
InvalidatedAnyEntries = true;
}
}
it->second.clear();
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
@@ -942,11 +910,15 @@ void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
}
}
void ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
std::optional<CustomIRResult>
@@ -955,7 +927,7 @@ ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandl
std::unique_lock lk(CustomIRMutex);
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, std::tuple(Handler, Creator, Data));
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, CustomIRHandlerEntry {Handler, Creator, Data});
HasCustomIRHandlers = true;
if (!InsertedIterator.second) {
@@ -979,21 +951,21 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
auto Result = AddCustomIREntrypoint(
Entrypoint,
[this, GuestThunkEntrypoint](uintptr_t Entrypoint, FEXCore::IR::IREmitter* emit) {
auto IRHeader = emit->_IRHeader(emit->Invalid(), Entrypoint, 0, 0, 0, 0);
auto Block = emit->CreateCodeNode();
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
auto IRHeader = emit->_IRHeader(emit->Invalid(), Entrypoint, 0, 0, 0, 0);
auto Block = emit->CreateCodeNode(true, 0);
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
const auto GPRSize = GetGPROpSize();
const auto GPRSize = this->Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->_Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::GPRFixedClass, X86State::REG_R11).Raw;
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->_Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(IR::OpSize::i64Bit, emit->_Constant(GuestThunkEntrypoint));
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::GPRFixedClass, X86State::REG_R11).Raw;
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(IR::OpSize::i64Bit, emit->Constant(GuestThunkEntrypoint), IR::BranchHint::None, emit->Invalid(), emit->Invalid());
},
ThunkHandler, (void*)GuestThunkEntrypoint);
@@ -1022,15 +994,36 @@ void ContextImpl::RemoveForceTSOInformation(uint64_t Address, uint64_t Size) {
ForceTSOInstructions.erase(ForceTSOInstructions.lower_bound(Address), ForceTSOInstructions.upper_bound(Address + Size));
}
void ContextImpl::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
void ContextImpl::MarkMonoBackpatcherBlock(uint64_t BlockEntry) {
MonoBackpatcherBlock.store(BlockEntry, std::memory_order_relaxed);
}
void ContextImpl::RemoveCustomIREntrypoint(FEXCore::Core::InternalThreadState* Thread, uintptr_t Entrypoint) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(nullptr, Entrypoint, 1);
CustomIRHandlers.erase(Entrypoint);
HasCustomIRHandlers = !CustomIRHandlers.empty();
SyscallHandler->InvalidateGuestCodeRange(Thread, Entrypoint, 1);
}
void ContextImpl::MonoBackpatcherWrite(FEXCore::Core::CpuStateFrame* Frame, uint8_t Size, uint64_t Address, uint64_t Value) {
auto Thread = Frame->Thread;
auto CTX = static_cast<ContextImpl*>(Thread->CTX);
{
auto lk = GuardSignalDeferringSection(CTX->CodeInvalidationMutex, Thread);
if (Size == 8) {
*reinterpret_cast<uint64_t *>(Address) = Value;
} else if (Size == 4) {
*reinterpret_cast<uint32_t *>(Address) = Value;
} else {
ERROR_AND_DIE_FMT("Unexpected write size for backpatcher: {}", Size);
}
}
CTX->SyscallHandler->InvalidateGuestCodeRange(Thread, Address, Size);
}
IR::AOTIRCacheEntry* ContextImpl::LoadAOTIRCacheEntry(const fextl::string& filename) {
@@ -1038,9 +1031,7 @@ IR::AOTIRCacheEntry* ContextImpl::LoadAOTIRCacheEntry(const fextl::string& filen
return rv;
}
void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry* Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
}
void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry* Entry) {}
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
@@ -2,6 +2,7 @@
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
@@ -19,11 +20,16 @@
#include <CodeEmitter/Emitter.h>
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#endif
#include <array>
#include <atomic>
#include <bit>
#include <condition_variable>
#include <csignal>
#include <cstring>
#include <signal.h>
namespace FEXCore::CPU {
@@ -81,11 +87,14 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::rsp, 0);
str(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
ARMEmitter::ForwardLabel CompileSingleStep;
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
FillStaticRegs();
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
ARMEmitter::BiDirectionalLabel LoopTop {};
ARMEmitter::ForwardLabel CompileSingleStep;
#ifdef _M_ARM_64EC
b(&LoopTop);
@@ -111,15 +120,26 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
ldr(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
FillSpecialRegs(TMP1, TMP2, false, true);
// As ARM64EC uses this as an entrypoint for both guest calls and host returns, opportunistically try to return
// using the call-ret stack to avoid unbalancing it.
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
ret(TMP2);
// Enter JIT
#endif
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
ARMEmitter::BiDirectionalLabel FullLookup {};
ARMEmitter::BiDirectionalLabel CallBlock {};
Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
@@ -127,23 +147,33 @@ void Dispatcher::EmitDispatcher() {
// Load in our RIP
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
#ifdef _M_ARM_64EC
// Clobbers TMP1/2
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::ForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
#endif
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
// L1 Cache
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP1, TMP1, 0);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(TMP4);
// L1C check failed, do a full lookup
Bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
@@ -271,38 +301,10 @@ void Dispatcher::EmitDispatcher() {
br(TMP1);
}
#ifdef _M_ARM_64EC
// Clobbers TMP1/2
auto EmitECExitCheck = [&]() {
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::ForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
};
#endif
// Need to create the block
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
EmitECExitCheck();
#endif
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -337,10 +339,6 @@ void Dispatcher::EmitDispatcher() {
{
Bind(&CompileSingleStep);
#ifdef _M_ARM_64EC
EmitECExitCheck();
#endif
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -498,6 +496,7 @@ void Dispatcher::EmitDispatcher() {
// load static regs
FillStaticRegs();
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
// Now go back to the regular dispatcher loop
b(&LoopTop);
@@ -550,8 +549,8 @@ void Dispatcher::EmitDispatcher() {
FABI_F80_I16_I32_PTR,
FABI_F32_I16_F80_PTR,
FABI_F64_I16_F80_PTR,
FABI_F64_I16_F64_PTR,
FABI_F64_I16_F64_F64_PTR,
FABI_F64_F64_PTR,
FABI_F64_F64_F64_PTR,
FABI_I16_I16_F80_PTR,
FABI_I32_I16_F80_PTR,
FABI_I64_I16_F80_PTR,
@@ -559,7 +558,7 @@ void Dispatcher::EmitDispatcher() {
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_F64x2_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
}};
@@ -607,6 +606,7 @@ void Dispatcher::EmitDispatcher() {
#ifdef VIXL_SIMULATOR
void Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.WriteXRegister(1, 0);
Simulator.RunFrom(reinterpret_cast< const vixl::aarch64::Instruction*>(DispatchPtr));
}
@@ -757,7 +757,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
mov(VABI1.Q(), VTMP1.Q());
}
mov(ARMEmitter::XReg::x1, STATE);
@@ -790,7 +790,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
FillF64Result();
} break;
case FABI_F64_I16_F64_PTR: {
case FABI_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -800,18 +800,17 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, uint64_t>(FallbackPointerReg);
GenerateIndirectRuntimeCall<double, double, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
FillF64Result();
} break;
case FABI_F64_I16_F64_F64_PTR: {
case FABI_F64_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -824,10 +823,9 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
fmov(VABI2.D(), VTMP2.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, double, uint64_t>(FallbackPointerReg);
GenerateIndirectRuntimeCall<double, double, double, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
@@ -987,7 +985,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
FillF80x2Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
case FABI_F64x2_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -996,14 +994,13 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
SpillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, TMP3, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
// GenerateIndirectRuntimeCall<FEXCore::VectorScalarF64Pair, uint16_t, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
// GenerateIndirectRuntimeCall<FEXCore::VectorScalarF64Pair, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
@@ -1105,6 +1102,53 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
}
}
SignalDelegatorConfig Dispatcher::MakeSignalDelegatorConfig() const {
// PF/AF are the final two SRA registers. We only want GPRs
const auto GPRCount = uint16_t(StaticRegisters.size() - 2);
const auto FPRCount = uint16_t(StaticFPRegisters.size());
const auto GetSRAGPRMapping = [GPRCount, this] {
SignalDelegatorConfig::SRAIndexMapping Mapping {};
for (size_t i = 0; i < GPRCount; ++i) {
Mapping[i] = StaticRegisters[i].Idx();
}
return Mapping;
};
const auto GetSRAFPRMapping = [FPRCount, this] {
SignalDelegatorConfig::SRAIndexMapping Mapping {};
for (size_t i = 0; i < FPRCount; ++i) {
Mapping[i] = StaticFPRegisters[i].Idx();
}
return Mapping;
};
return FEXCore::SignalDelegatorConfig {
.DispatcherBegin = Start,
.DispatcherEnd = End,
.AbsoluteLoopTopAddress = AbsoluteLoopTopAddress,
.AbsoluteLoopTopAddressFillSRA = AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = SignalHandlerReturnAddressRT,
.PauseReturnInstruction = PauseReturnInstruction,
.ThreadPauseHandlerAddressSpillSRA = ThreadPauseHandlerAddressSpillSRA,
.ThreadPauseHandlerAddress = ThreadPauseHandlerAddress,
// Stop handlers.
.ThreadStopHandlerAddressSpillSRA = ThreadStopHandlerAddressSpillSRA,
.ThreadStopHandlerAddress = ThreadStopHandlerAddress,
// SRA information.
.SRAGPRCount = GPRCount,
.SRAFPRCount = FPRCount,
.SRAGPRMapping = GetSRAGPRMapping(),
.SRAFPRMapping = GetSRAFPRMapping(),
};
}
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl* CTX) {
return fextl::make_unique<Dispatcher>(CTX);
}
@@ -2,25 +2,18 @@
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/fextl/memory.h>
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#endif
#include <array>
#include <cstddef>
#include <cstdint>
#include <signal.h>
#include <stddef.h>
#include <stack>
#include <tuple>
namespace FEXCore {
struct GuestSigAction;
}
struct SignalDelegatorConfig;
} // namespace FEXCore
namespace FEXCore::Core {
struct CpuStateFrame;
@@ -42,6 +35,32 @@ public:
Dispatcher(FEXCore::Context::ContextImpl* ctx);
~Dispatcher();
void InitThreadPointers(FEXCore::Core::InternalThreadState* Thread);
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame);
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame, false);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
#endif
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame, bool SingleInst);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
private:
/**
* @name Dispatch Helper functions
* @{ */
@@ -59,62 +78,14 @@ public:
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
uint64_t IntCallbackReturnAddress {};
uint64_t PauseReturnInstruction {};
std::array<uint64_t, FallbackABI::FABI_UNKNOWN> ABIPointers {};
/** @} */
uint64_t Start {};
uint64_t End {};
void InitThreadPointers(FEXCore::Core::InternalThreadState* Thread);
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame);
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
#endif
uint16_t GetSRAGPRCount() const {
// PF/AF are the final two SRA registers.
// Only return the SRA for GPRs.
return StaticRegisters.size() - 2;
}
uint16_t GetSRAFPRCount() const {
return StaticFPRegisters.size();
}
void GetSRAGPRMapping(uint8_t Mapping[16]) const {
for (size_t i = 0; i < StaticRegisters.size() - 2; ++i) {
Mapping[i] = StaticRegisters[i].Idx();
}
}
void GetSRAFPRMapping(uint8_t Mapping[16]) const {
for (size_t i = 0; i < StaticFPRegisters.size(); ++i) {
Mapping[i] = StaticFPRegisters[i].Idx();
}
}
protected:
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
private:
// Long division helpers
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
+283 -92
View File
@@ -9,6 +9,8 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
#include <algorithm>
@@ -21,6 +23,7 @@ $end_info$
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/fextl/set.h>
namespace FEXCore::Frontend {
@@ -64,29 +67,90 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
}
}
Decoder::Decoder(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, OSABI {ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {}
Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
: Thread {Thread}
, CTX {static_cast<FEXCore::Context::ContextImpl*>(Thread->CTX)}
, OSABI {CTX->SyscallHandler ? CTX->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {CTX->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
FEX_CONFIG_OPT(ReducedPrecision, X87REDUCEDPRECISION);
if (ReducedPrecision) {
X87Table = &FEXCore::X86Tables::X87F64Ops;
} else {
X87Table = &FEXCore::X86Tables::X87F80Ops;
}
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
VEXTable = &FEXCore::X86Tables::VEXTableOps;
VEXTableGroup = &FEXCore::X86Tables::VEXTableGroupOps;
} else if (CTX->HostFeatures.SupportsAVX) {
VEXTable = &FEXCore::X86Tables::VEXTableOps_AVX128;
VEXTableGroup = &FEXCore::X86Tables::VEXTableGroupOps_AVX128;
}
}
uint8_t Decoder::PeekByte(uint8_t Offset) const {
uint8_t Byte = InstStream[InstructionSize + Offset];
return Byte;
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
ExecutableRangeEnd = RangeInfo.Base + RangeInfo.Size;
ExecutableRangeWritable = RangeInfo.Writable;
if (RangeInfo.Size == 0) {
return false;
}
uint64_t RangeRemainingSize = ExecutableRangeEnd - Address;
if (Size > RangeRemainingSize) {
Size -= RangeRemainingSize;
Address += RangeRemainingSize;
}
}
return true;
}
uint8_t Decoder::ReadByte() {
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
std::optional<uint8_t> Byte = PeekByte(0);
if (!Byte) {
HitNonExecutableRange = true;
// Pretend we read 0, the main decode loop will see HitNonExecutableRange and rollback the instruction.
return 0;
}
Instruction[InstructionSize] = *Byte;
InstructionSize++;
return *Byte;
}
std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream + InstructionSize + Offset);
if (CheckRangeExecutable(ByteAddress, 1)) {
return InstStream[InstructionSize + Offset];
} else {
return std::nullopt;
}
}
uint64_t Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
uint64_t Address = reinterpret_cast<uint64_t>(InstStream + InstructionSize);
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream[InstructionSize], Size);
} else {
HitNonExecutableRange = true;
// See PeekByte, this specific case may cause some executable memory to read as 0 but it doesn't matter as the entire instruction will be rolled back anyway.
Res = 0;
}
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
for (size_t i = 0; i < Size; ++i) {
@@ -267,6 +331,13 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
// TODO: Move this in to `NormalOpHeader`, Dispatch tables have a bug currently where some subtables don't inherit flags correctly.
// Can be seen by running FEX asm tests if this is removed.
return NormalOp(&Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0], Op);
}
DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
@@ -286,7 +357,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
uint8_t DestSize {};
const bool HasWideningDisplacement =
(FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST) != 0 ||
(Options.w && CTX->Config.Is64BitMode);
(Options.w && BlockInfo.Is64BitMode);
const bool HasNarrowingDisplacement =
(FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
@@ -304,7 +375,21 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const bool HasMODRM = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM);
const bool HasREX = !!(DecodeInst->Flags & DecodeFlags::FLAG_REX_PREFIX);
const bool Has16BitAddressing = !CTX->Config.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
if (Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_0)) {
return false;
} else if (!Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_1)) {
return false;
}
if (Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_0)) {
return false;
} else if (!Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_1)) {
return false;
}
const bool UseVEXL = Options.L && !(Info->Flags & InstFlags::FLAGS_VEX_L_IGNORE);
// This is used for ModRM register modification
// For both modrm.reg and modrm.rm(when mod == 0b11) when value is >= 0b100
@@ -335,7 +420,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
DestSize = 2;
} else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
if (Options.L) {
if (UseVEXL) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_256BIT);
DestSize = 32;
} else {
@@ -351,7 +436,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
DestSize = 2;
} else if ((HasXMMDst || HasMMDst || CTX->Config.Is64BitMode) &&
} else if ((HasXMMDst || HasMMDst || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_64BIT);
@@ -368,7 +453,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
} else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_16BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
} else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
if (Options.L) {
if (UseVEXL) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
@@ -380,7 +465,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// See table 1-2. Operand-Size Overrides for this decoding
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
} else if ((HasXMMSrc || HasMMSrc || CTX->Config.Is64BitMode) &&
} else if ((HasXMMSrc || HasMMSrc || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_64BIT);
@@ -480,6 +565,9 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
size_t CurrentSrc = 0;
const auto VEXOperand = Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_SRC_MASK;
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_NO_OPERAND && Options.vvvv) {
return false;
}
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
@@ -562,7 +650,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
}
bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
DecodeInst->OP = Op;
DecodeInst->OPRaw = DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
@@ -578,6 +666,9 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
// A normal instruction is the most likely.
if (Info->Type == FEXCore::X86Tables::TYPE_INST) [[likely]] {
return NormalOp(Info, Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
return NormalOp(&Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0], Op);
} else if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
@@ -615,7 +706,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
ModRM.Hex = DecodeInst->ModRM;
uint16_t LocalOp = OPD(Info->Type, PrefixType, ModRM.reg);
FEXCore::X86Tables::X86InstInfo* LocalInfo = &SecondInstGroupOps[LocalOp];
const FEXCore::X86Tables::X86InstInfo* LocalInfo = &SecondInstGroupOps[LocalOp];
#undef OPD
if (LocalInfo->Type == FEXCore::X86Tables::TYPE_SECOND_GROUP_MODRM && ModRM.mod == 0b11) {
// Everything in this group is privileged instructions aside from XGETBV
@@ -639,15 +730,20 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodeInst->DecodedModRM = true;
uint16_t X87Op = ((Op - 0xD8) << 8) | ModRMByte;
return NormalOp(&X87Ops[X87Op], X87Op);
return NormalOp(&(*X87Table)[X87Op], X87Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
if (!VEXTable) {
// AVX not enabled.
return false;
}
uint16_t map_select = 1;
uint16_t pp = 0;
const uint8_t Byte1 = ReadByte();
DecodedHeader options {};
if ((Byte1 & 0b10000000) == 0) {
if (!CTX->Config.Is64BitMode) {
if (!BlockInfo.Is64BitMode) {
return false;
}
@@ -666,19 +762,18 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
if (!CTX->Config.Is64BitMode) {
if (!BlockInfo.Is64BitMode) {
return false;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
if (BlockInfo.Is64BitMode && (Byte1 & 0b00100000) == 0) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
if (options.w) {
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
return false;
}
}
@@ -688,7 +783,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
Op = OPD(map_select, pp, VEXOp);
#undef OPD
FEXCore::X86Tables::X86InstInfo* LocalInfo = &VEXTableOps[Op];
const FEXCore::X86Tables::X86InstInfo* LocalInfo = &(*VEXTable)[Op];
if (LocalInfo->Type >= FEXCore::X86Tables::TYPE_VEX_GROUP_12 && LocalInfo->Type <= FEXCore::X86Tables::TYPE_VEX_GROUP_17) {
// We have ModRM
@@ -702,7 +797,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
Op = OPD(LocalInfo->Type, pp, ModRM.reg);
#undef OPD
return NormalOp(&VEXTableGroupOps[Op], Op, options);
return NormalOp(&(*VEXTableGroup)[Op], Op, options);
} else {
return NormalOp(LocalInfo, Op, options);
}
@@ -716,7 +811,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstruction(uint64_t PC) {
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
Instruction.fill(0);
@@ -744,7 +839,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
const bool Has16BitAddressing = !CTX->Config.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
@@ -853,29 +948,25 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->Flags |= DecodeFlags::FLAG_ADDRESS_SIZE;
break;
case 0x26: // ES legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_ES_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_ES_PREFIX;
}
break;
case 0x2E: // CS legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_CS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_CS_PREFIX;
}
break;
case 0x36: // SS legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_SS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_SS_PREFIX;
}
break;
case 0x3E: // DS legacy prefix
// Annoyingly GCC generates NOP ops with these prefixes
// Just ignore them for now
// eg. 66 2e 0f 1f 84 00 00 00 00 00 nop WORD PTR cs:[rax+rax*1+0x0]
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_DS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_DS_PREFIX;
}
break;
break;
case 0xF0: // LOCK prefix
DecodeInst->Flags |= DecodeFlags::FLAG_LOCK;
break;
@@ -888,19 +979,19 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->LastEscapePrefix = Op;
break;
case 0x64: // FS prefix
DecodeInst->Flags |= DecodeFlags::FLAG_FS_PREFIX;
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_FS_PREFIX;
break;
case 0x65: // GS prefix
DecodeInst->Flags |= DecodeFlags::FLAG_GS_PREFIX;
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_GS_PREFIX;
break;
default:
[[likely]] { // Default base table
auto Info = &FEXCore::X86Tables::BaseOps[Op];
const X86InstInfo* Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) {
Info = &Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0];
}
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
if (!CTX->Config.Is64BitMode) {
return false;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
@@ -931,6 +1022,47 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
}
}
if (DecodeInst->Dest.IsGPR()) {
return false;
}
return true;
}
Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
return ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST : DecodedBlockStatus::NOEXEC_INST;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
}
if (CTX->AreMonoHacksActive()) {
// Unity uses a standard SPSC ringbuffer with cached read/write pointers and thread waiting flags at the following
// offsets, which are consistent between 32-bit and 64-bit Unity versions from 2015 onwards.
auto IsKnownAtomicDisplacement = [](uint64_t Displacement) {
return Displacement == 0x80 || Displacement == 0x84 || Displacement == 0xC0 || Displacement == 0xC4;
};
if (DecodeInst->OP == 0x8b && DecodeInst->Src[0].IsGPRIndirect() &&
IsKnownAtomicDisplacement(DecodeInst->Src[0].Data.GPRIndirect.Displacement)) {
DecodeInst->ForceTSO = true;
}
if (DecodeInst->OP == 0x89 && DecodeInst->Dest.IsGPRIndirect() && IsKnownAtomicDisplacement(DecodeInst->Dest.Data.GPRIndirect.Displacement)) {
DecodeInst->ForceTSO = true;
}
}
return DecodedBlockStatus::SUCCESS;
}
void Decoder::BranchTargetInMultiblockRange() {
@@ -940,16 +1072,30 @@ void Decoder::BranchTargetInMultiblockRange() {
// If the RIP setting is conditional AND within our symbol range then it can be considered for multiblock
uint64_t TargetRIP = 0;
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
bool Conditional = true;
const auto InstEnd = DecodeInst->PC + DecodeInst->InstSize;
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_CALL) {
if (ExecutableRangeWritable && CTX->AreMonoHacksActive()) {
// Mono generated code often contains noreturn calls with garbage following them, and calls are always backpatched
// after CIL compilation leading to n recompiles for a multiblock with n calls. Choose to minimize stutters over
// raw performance and disable tracking past calls for mono generated code.
return;
}
AddBranchTarget(InstEnd);
BlockInfo.EntryPoints.emplace(InstEnd);
return;
}
// Calls are handled above
switch (DecodeInst->OP) {
case 0x70 ... 0x7F: // Conditional JUMP
case 0x80 ... 0x8F: { // More conditional
// Source is a literal
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// auto RIPTargetConst = Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
break;
@@ -959,11 +1105,6 @@ void Decoder::BranchTargetInMultiblockRange() {
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
Conditional = false;
break;
case 0xE8: // Call - Immediate target, We don't want to inline calls
if (ExternalBranches) {
ExternalBranches->insert(InstEnd);
}
[[fallthrough]];
case 0xC2: // RET imm
case 0xC3: // RET
default: return; break;
@@ -974,10 +1115,16 @@ void Decoder::BranchTargetInMultiblockRange() {
TargetRIP &= 0xFFFFFFFFU;
}
if (Conditional) {
// If we are conditional then a target can be the instruction past the conditional instruction
AddBranchTarget(InstEnd);
}
// If the target RIP is x86 code within the symbol ranges then we are golden
// Forbid cross-page branches to both avoid massive (range-wise) code blocks in highly fragmented code and trying to decode unmapped branch targets
bool ValidMultiblockMember =
TargetRIP >= SymbolMinAddress && TargetRIP < std::min(FEXCore::AlignUp(InstEnd, FEXCore::Utils::FEX_PAGE_SIZE), SymbolMaxAddress);
// Forbid distant branches to have the cost code better match the guest code layout, avoiding massive (range-wise) code
// blocks in highly fragmented guest code. Such branches are often not-taken branches to garbage in obfuscated code.
constexpr uint64_t MAX_FORWARD_BRANCH_DIST = FEXCore::Utils::FEX_PAGE_SIZE * 4;
bool ValidMultiblockMember = TargetRIP >= SymbolMinAddress && TargetRIP < std::min(InstEnd + MAX_FORWARD_BRANCH_DIST, SymbolMaxAddress);
#ifdef _M_ARM_64EC
ValidMultiblockMember = ValidMultiblockMember && !RtlIsEcCode(TargetRIP);
@@ -988,9 +1135,6 @@ void Decoder::BranchTargetInMultiblockRange() {
if (Conditional) {
MaxCondBranchForward = std::max(MaxCondBranchForward, TargetRIP);
MaxCondBranchBackwards = std::min(MaxCondBranchBackwards, TargetRIP);
// If we are conditional then a target can be the instruction past the conditional instruction
AddBranchTarget(InstEnd);
}
AddBranchTarget(TargetRIP);
@@ -1001,6 +1145,60 @@ void Decoder::BranchTargetInMultiblockRange() {
}
}
bool Decoder::IsBranchMonoTailcall(uint64_t NumInstructions) const {
// While the mono call backpatching block can easily be detected due it being the only one to contain SMC-faulting
// atomics, that can't be said for the tailcall jump backpatcher which has changed several times across versions and
// can be partially inlined. To work around this, instead detect the tailcall site itself and force full non-signal-based
// SMC detection for that single block.
if (!ExecutableRangeWritable) {
// We only care about jitted code
return false;
}
// See mini-{amd64,x86}.c in the mono codebase, specifically where METHOD_JUMP patches are emitted.
if (GetGPROpSize() == IR::OpSize::i32Bit) {
// Matches:
// LEAVE
// <none> / NOP / MOV EAX, EAX / LEA EBP, [EBP+0]
// JMP imm32
if (DecodeInst->OP != 0xE9 || NumInstructions < 2) {
return false;
}
auto PrevInst = std::prev(DecodeInst);
if (PrevInst->OP == 0xC9) {
return true;
}
if (NumInstructions < 3 || std::prev(PrevInst)->OP != 0xC9) {
return false;
}
return PrevInst->OP == 0x90 || (PrevInst->OP == 0x8B && PrevInst->ModRM == 0xC0) ||
(PrevInst->OP == 0x8D && PrevInst->ModRM == 0x6D && PrevInst->Src[1].IsLiteral() && PrevInst->Src[1].Literal() == 0);
} else {
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
if (DecodeInst->OPRaw == 0xFF && ModRM.reg == 4 && DecodeInst->Src[0].IsGPR()) {
if (DecodeInst->Src[0].Data.GPR.GPR == FEXCore::X86State::REG_RAX) {
// Found in versions of mono from 2024 onwards - matches:
// REX.W JMP rax
return (DecodeInst->Flags & (DecodeFlags::FLAG_REX_PREFIX | DecodeFlags::FLAG_REX_WIDENING | DecodeFlags::FLAG_REX_XGPR_B |
DecodeFlags::FLAG_REX_XGPR_X | DecodeFlags::FLAG_REX_XGPR_R)) ==
(DecodeFlags::FLAG_REX_PREFIX | DecodeFlags::FLAG_REX_WIDENING);
} else if (NumInstructions > 1 && DecodeInst->Src[0].Data.GPR.GPR == FEXCore::X86State::REG_R11) {
// Found in older versions of mono - match:
// MOV r11, imm64
// JMP r11
auto PrevInst = std::prev(DecodeInst);
return PrevInst->OP == 0xBB && PrevInst->Dest.IsGPR() && PrevInst->Dest.Data.GPR.GPR == FEXCore::X86State::REG_R11;
}
}
}
return false;
}
bool Decoder::InstCanContinue() const {
if (DecodeInst->PC + DecodeInst->InstSize == NextBlockStartAddress) {
return false;
@@ -1011,7 +1209,7 @@ bool Decoder::InstCanContinue() const {
}
uint64_t TargetRIP = 0;
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
if (DecodeInst->OP == 0xE8) { // Call - immediate target
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
@@ -1063,7 +1261,7 @@ void Decoder::AddBranchTarget(uint64_t Target) {
.Size = BlockIt->Size - SplitOffset,
.NumInstructions = BlockIt->NumInstructions - SplitIdx,
.DecodedInstructions = BlockIt->DecodedInstructions + SplitIdx,
.HasInvalidInstruction = BlockIt->HasInvalidInstruction,
.BlockStatus = BlockIt->BlockStatus,
};
BlockIt->Size = SplitOffset;
@@ -1103,12 +1301,10 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst,
std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage) {
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
BlocksToDecode.clear();
VisitedBlocks.clear();
// Reset internal state management
DecodedSize = 0;
@@ -1116,9 +1312,15 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// Decode operating mode from thread's CS segment.
const auto CSSegment = Core::CPUState::GetSegmentFromIndex(Thread->CurrentFrame->State, Thread->CurrentFrame->State.cs_idx);
BlockInfo.Is64BitMode = CSSegment->L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
// XXX: Load symbol data
SymbolAvailable = false;
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
@@ -1134,13 +1336,11 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
DecodedMaxAddress = EntryPoint;
// Entry is a jump target
BlocksToDecode.emplace(PC);
BlocksToDecode = {PC};
uint64_t CurrentCodePage = PC & FEXCore::Utils::FEX_PAGE_MASK;
fextl::set<uint64_t> CodePages = {CurrentCodePage};
AddContainedCodePage(PC, CurrentCodePage, FEXCore::Utils::FEX_PAGE_SIZE);
BlockInfo.CodePages = {CurrentCodePage};
if (MaxInst == 0) {
MaxInst = CTX->Config.MaxInstPerBlock;
@@ -1175,6 +1375,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockIt->Entry = RIPToDecode;
BlockIt->Size = 0;
BlockIt->IsEntryPoint = EntryBlock;
uint64_t PCOffset = 0;
uint64_t BlockStartOffset = DecodedSize;
@@ -1196,7 +1397,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
auto OpMinPage = OpAddress & FEXCore::Utils::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FEXCore::Utils::FEX_PAGE_MASK;
if (!EntryBlock && OpMinPage == OpMaxPage && PeekByte(0) == 0 && PeekByte(1) == 0) [[unlikely]] {
if (!EntryBlock && OpMinPage == OpMaxPage && PeekByte(0).value_or(0) == 0 && PeekByte(1).value_or(0) == 0) [[unlikely]] {
// End the multiblock early if we hit 2 consecutive null bytes (add [rax], al) in the same page with the
// assumption we are most likely trying to explore garbage code.
break;
@@ -1204,31 +1405,17 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
CodePages.insert(CurrentCodePage);
BlockInfo.CodePages.insert(CurrentCodePage);
}
if (OpMaxPage != CurrentCodePage) {
CurrentCodePage = OpMaxPage;
CodePages.insert(CurrentCodePage);
BlockInfo.CodePages.insert(CurrentCodePage);
}
bool ErrorDuringDecoding = !DecodeInstruction(OpAddress);
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
if (ErrorDuringDecoding) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
BlockIt->HasInvalidInstruction = true;
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
} else {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
auto TableInfo = DecodeInst->TableInfo;
if (!TableInfo || !TableInfo->OpcodeDispatcher) {
BlockIt->HasInvalidInstruction = true;
}
}
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
DecodedMaxAddress = std::max(DecodedMaxAddress, OpEndAddress);
@@ -1244,7 +1431,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockIt->Size += DecodeInst->InstSize;
// Can not continue this block at all on invalid instruction
if (BlockIt->HasInvalidInstruction) [[unlikely]] {
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
@@ -1253,6 +1440,9 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
DecodedSize = BlockStartOffset;
InstStream -= PCOffset;
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" : "NoExec", OpAddress);
}
break;
}
@@ -1269,6 +1459,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
// NOTE: This will invalidate BlockIt, this is fine as we immediately break from the loop and EraseBlock cannot be true
BlockIt->ForceFullSMCDetection = CTX->AreMonoHacksActive() && IsBranchMonoTailcall(BlockIt->NumInstructions);
BranchTargetInMultiblockRange();
}
@@ -1292,8 +1483,8 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockInfo.TotalInstructionCount = TotalInstructions;
for (auto CodePage : CodePages) {
AddContainedCodePage(PC, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
for (auto& Block : BlockInfo.Blocks) {
Block.IsEntryPoint = BlockInfo.EntryPoints.contains(Block.Entry);
}
}
+45 -9
View File
@@ -2,40 +2,54 @@
#pragma once
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <cstddef>
#include <cstdint>
#include <stddef.h>
#include <optional>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::HLE {
enum class SyscallOSABI;
}
namespace FEXCore::Frontend {
class Decoder final {
public:
enum class DecodedBlockStatus {
SUCCESS,
INVALID_INST,
NOEXEC_INST,
};
// New Frontend decoding
struct DecodedBlocks final {
uint64_t Entry {};
uint64_t Size {};
uint64_t NumInstructions {};
FEXCore::X86Tables::DecodedInst* DecodedInstructions;
bool HasInvalidInstruction {};
DecodedBlockStatus BlockStatus;
bool IsEntryPoint {};
bool ForceFullSMCDetection {};
};
struct DecodedBlockInformation final {
uint64_t TotalInstructionCount;
bool Is64BitMode {};
fextl::vector<DecodedBlocks> Blocks;
fextl::set<uint64_t> EntryPoints;
fextl::set<uint64_t> CodePages; // Start addresses of all pages touching the block
};
Decoder(FEXCore::Context::ContextImpl* ctx);
void DecodeInstructionsAtEntry(const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst,
std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
Decoder(FEXCore::Core::InternalThreadState* Thread);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
return &BlockInfo;
@@ -55,6 +69,10 @@ public:
PoolObject.DelayedDisownBuffer();
}
void ResetExecutableRangeCache() {
ExecutableRangeBase = ExecutableRangeEnd = 0;
}
private:
// To pass any information from instruction prefixes
// down into the actual instruction handling machinery.
@@ -64,18 +82,23 @@ private:
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Core::InternalThreadState* Thread;
FEXCore::Context::ContextImpl* CTX;
const FEXCore::HLE::SyscallOSABI OSABI {};
bool DecodeInstruction(uint64_t PC);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
bool IsBranchMonoTailcall(uint64_t NumInstructions) const;
bool InstCanContinue() const;
void AddBranchTarget(uint64_t Target);
bool CheckRangeExecutable(uint64_t Address, uint64_t Size);
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset) const;
std::optional<uint8_t> PeekByte(uint8_t Offset);
uint64_t ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) {
InstructionSize += Size;
@@ -89,7 +112,15 @@ private:
Utils::PoolBufferWithTimedRetirement<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint64_t ExecutableRangeBase {};
uint64_t ExecutableRangeEnd {};
bool ExecutableRangeWritable {};
bool HitNonExecutableRange {};
const uint8_t* InstStream {};
IR::OpSize GetGPROpSize() const {
return BlockInfo.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
static constexpr size_t MAX_INST_SIZE = 15;
uint8_t InstructionSize {};
@@ -122,6 +153,11 @@ private:
&FEXCore::Frontend::Decoder::DecodeModRM_16,
};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_X87_TABLE_SIZE>* X87Table;
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_TABLE_SIZE>* VEXTable {};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_GROUP_TABLE_SIZE>* VEXTableGroup {};
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
};
} // namespace FEXCore::Frontend
@@ -6,7 +6,7 @@
#include "Interface/IR/IR.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SHMStats.h>
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW, bool Force80BitPrecision = false) {
@@ -36,18 +36,48 @@ FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t
return State;
}
FEXCORE_PRESERVE_ALL_ATTR static void HandleX87Exception(const softfloat_state& State, FEXCore::Core::CpuStateFrame* Frame) {
// Check for Invalid Operation exception (bit 0 of X87 status word)
if (State.exceptionFlags & softfloat_flag_invalid) {
Frame->State.flags[FEXCore::X86State::X87FLAG_IE_LOC] = 1;
}
}
// Wrapper for SoftFloat state to handle X87 exceptions
class ScopedSoftFloatState {
public:
FEXCORE_PRESERVE_ALL_ATTR ScopedSoftFloatState(uint16_t FCW, FEXCore::Core::CpuStateFrame* Frame, bool Force80BitPrecision = false)
: State(SoftFloatStateFromFCW(FCW, Force80BitPrecision))
, Frame(Frame) {}
FEXCORE_PRESERVE_ALL_ATTR ~ScopedSoftFloatState() {
HandleX87Exception(State, Frame);
}
// Disable copy and move to ensure RAII semantics
ScopedSoftFloatState(const ScopedSoftFloatState&) = delete;
ScopedSoftFloatState& operator=(const ScopedSoftFloatState&) = delete;
ScopedSoftFloatState(ScopedSoftFloatState&&) = delete;
ScopedSoftFloatState& operator=(ScopedSoftFloatState&&) = delete;
softfloat_state State;
private:
FEXCore::Core::CpuStateFrame* Frame;
};
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle4(uint16_t FCW, float src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(&State.State, src);
}
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle8(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(&State.State, src);
}
};
@@ -55,12 +85,12 @@ template<>
struct OpHandlers<IR::OP_F80CMP> {
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
ScopedSoftFloatState State {FCW, Frame};
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(&State, Src1, Src2, &eq, &lt, &nan);
X80SoftFloat::FCMP(&State.State, Src1, Src2, &eq, &lt, &nan);
if (lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
@@ -78,14 +108,14 @@ template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToF32(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToF32(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToF64(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToF64(&State.State);
}
};
@@ -93,26 +123,26 @@ template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI16(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI16(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI32(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI32(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI64(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI64(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
auto rv = extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
auto rv = extF80_to_i32(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
@@ -124,14 +154,14 @@ struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
return extF80_to_i32(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i64(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
return extF80_to_i64(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
}
};
@@ -152,8 +182,8 @@ template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FRNDINT(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FRNDINT(&State.State, Src1);
}
};
@@ -161,8 +191,8 @@ template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::F2XM1(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::F2XM1(&State.State, Src1);
}
};
@@ -170,8 +200,8 @@ template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FTAN(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FTAN(&State.State, Src1);
}
};
@@ -179,8 +209,8 @@ template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSQRT(&State, Src1);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FSQRT(&State.State, Src1);
}
};
@@ -188,8 +218,8 @@ template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSIN(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FSIN(&State.State, Src1);
}
};
@@ -197,8 +227,8 @@ template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FCOS(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FCOS(&State.State, Src1);
}
};
@@ -206,8 +236,8 @@ template<>
struct OpHandlers<IR::OP_F80SINCOS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegPairType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return FEXCore::MakeVectorRegPair(X80SoftFloat::FSIN(&State, Src1), X80SoftFloat::FCOS(&State, Src1));
ScopedSoftFloatState State {FCW, Frame, true};
return FEXCore::MakeVectorRegPair(X80SoftFloat::FSIN(&State.State, Src1), X80SoftFloat::FCOS(&State.State, Src1));
}
};
@@ -231,8 +261,8 @@ template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FADD(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FADD(&State.State, Src1, Src2);
}
};
@@ -240,8 +270,8 @@ template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSUB(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FSUB(&State.State, Src1, Src2);
}
};
@@ -249,8 +279,8 @@ template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FMUL(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FMUL(&State.State, Src1, Src2);
}
};
@@ -258,8 +288,8 @@ template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FDIV(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FDIV(&State.State, Src1, Src2);
}
};
@@ -267,8 +297,8 @@ template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FYL2X(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FYL2X(&State.State, Src1, Src2);
}
};
@@ -276,8 +306,8 @@ template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FATAN(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FATAN(&State.State, Src1, Src2);
}
};
@@ -285,8 +315,8 @@ template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM1(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FREM1(&State.State, Src1, Src2);
}
};
@@ -294,8 +324,8 @@ template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FREM(&State.State, Src1, Src2);
}
};
@@ -303,14 +333,14 @@ template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSCALE(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FSCALE(&State.State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return sin(src);
}
@@ -318,7 +348,7 @@ struct OpHandlers<IR::OP_F64SIN> {
template<>
struct OpHandlers<IR::OP_F64COS> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return cos(src);
}
@@ -326,7 +356,7 @@ struct OpHandlers<IR::OP_F64COS> {
template<>
struct OpHandlers<IR::OP_F64SINCOS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorScalarF64Pair handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static VectorScalarF64Pair handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
double sin, cos;
#ifdef _WIN32
@@ -341,7 +371,7 @@ struct OpHandlers<IR::OP_F64SINCOS> {
template<>
struct OpHandlers<IR::OP_F64TAN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return tan(src);
}
@@ -349,7 +379,7 @@ struct OpHandlers<IR::OP_F64TAN> {
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return exp2(src) - 1.0;
}
@@ -357,7 +387,7 @@ struct OpHandlers<IR::OP_F64F2XM1> {
template<>
struct OpHandlers<IR::OP_F64ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return atan2(src1, src2);
}
@@ -365,7 +395,7 @@ struct OpHandlers<IR::OP_F64ATAN> {
template<>
struct OpHandlers<IR::OP_F64FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return fmod(src1, src2);
}
@@ -373,7 +403,7 @@ struct OpHandlers<IR::OP_F64FPREM> {
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return remainder(src1, src2);
}
@@ -381,7 +411,7 @@ struct OpHandlers<IR::OP_F64FPREM1> {
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(src1);
}
@@ -389,7 +419,7 @@ struct OpHandlers<IR::OP_F64FYL2X> {
template<>
struct OpHandlers<IR::OP_F64SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
if (src1 == 0.0) { // src1 might be +/- zero
return src1; // this will return negative or positive zero if when appropriate
@@ -404,18 +434,17 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1q, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
X80SoftFloat Src1 = Src1q;
softfloat_state State = SoftFloatStateFromFCW(FCW);
ScopedSoftFloatState State {FCW, Frame};
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(&State, Src1);
Src1 = X80SoftFloat::FRNDINT(&State.State, Src1);
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1.ToI64(&State);
uint64_t Tmp = Src1.ToI64(&State.State);
X80SoftFloat Rv;
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
for (size_t i = 0; i < 9; ++i) {
if (Tmp == 0) {
@@ -82,24 +82,24 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80SCALE>::handle)};
// Double Precision Unary
Info[Core::OPINDEX_F64SIN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle)};
Info[Core::OPINDEX_F64COS] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle)};
Info[Core::OPINDEX_F64SINCOS] = {ABIHandlers[FABI_F64x2_I16_F64_PTR],
Info[Core::OPINDEX_F64SIN] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle)};
Info[Core::OPINDEX_F64COS] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle)};
Info[Core::OPINDEX_F64SINCOS] = {ABIHandlers[FABI_F64x2_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SINCOS>::handle)};
Info[Core::OPINDEX_F64TAN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle)};
Info[Core::OPINDEX_F64F2XM1] = {ABIHandlers[FABI_F64_I16_F64_PTR],
Info[Core::OPINDEX_F64TAN] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle)};
Info[Core::OPINDEX_F64F2XM1] = {ABIHandlers[FABI_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle)};
// Double Precision Binary
Info[Core::OPINDEX_F64ATAN] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64ATAN] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64ATAN>::handle)};
Info[Core::OPINDEX_F64FPREM] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64FPREM] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM>::handle)};
Info[Core::OPINDEX_F64FPREM1] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64FPREM1] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle)};
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle)};
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle)};
// SSE4.2 string instructions
@@ -222,18 +222,18 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
#define COMMON_UNARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
*Info = {FABI_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_UNARYPAIR_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64x2_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
*Info = {FABI_F64x2_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_BINARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_I16_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
*Info = {FABI_F64_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
@@ -19,8 +19,8 @@ enum FallbackABI {
FABI_F80_I16_I32_PTR,
FABI_F32_I16_F80_PTR,
FABI_F64_I16_F80_PTR,
FABI_F64_I16_F64_PTR,
FABI_F64_I16_F64_F64_PTR,
FABI_F64_F64_PTR,
FABI_F64_F64_F64_PTR,
FABI_I16_I16_F80_PTR,
FABI_I32_I16_F80_PTR,
FABI_I64_I16_F80_PTR,
@@ -28,7 +28,7 @@ enum FallbackABI {
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_F64x2_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
FABI_UNKNOWN,
+55 -28
View File
@@ -515,6 +515,12 @@ DEF_OP(AndWithFlags) {
}
}
DEF_OP(AndShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
and_(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1), GetReg(Op->Src2), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
DEF_OP(XorShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
@@ -1038,34 +1044,55 @@ DEF_OP(Popcount) {
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src);
switch (OpSize) {
case IR::OpSize::i8Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
// only use lowest byte
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i16Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// only count two lowest bytes
addp(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i32Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i64Bit:
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
default: LOGMAN_MSG_A_FMT("Unsupported Popcount size: {}", OpSize);
if (CTX->HostFeatures.SupportsCSSC) {
switch (OpSize) {
case IR::OpSize::i8Bit:
uxtb(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i16Bit:
uxth(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i32Bit:
cnt(ARMEmitter::Size::i32Bit, Dst, Src);
break;
case IR::OpSize::i64Bit:
cnt(ARMEmitter::Size::i64Bit, Dst, Src);
break;
default: LOGMAN_MSG_A_FMT("Unsupported Popcount size: {}", OpSize);
}
}
else {
switch (OpSize) {
case IR::OpSize::i8Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
// only use lowest byte
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i16Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// only count two lowest bytes
addp(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i32Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
case IR::OpSize::i64Bit:
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), Src);
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
// fmov has zero extended, unused bytes are zero
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
break;
default: LOGMAN_MSG_A_FMT("Unsupported Popcount size: {}", OpSize);
}
umov<ARMEmitter::SubRegSize::i8Bit>(Dst, VTMP1, 0);
umov<ARMEmitter::SubRegSize::i8Bit>(Dst, VTMP1, 0);
}
}
DEF_OP(FindLSB) {
@@ -1347,12 +1374,12 @@ DEF_OP(VExtractToGPR) {
const auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto OpSize = IROp->Size;
[[maybe_unused]] constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto SSERegBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto ElementSizeBits = IR::OpSizeAsBits(Op->Header.ElementSize);
const auto Offset = ElementSizeBits * Op->Index;
[[maybe_unused]] const auto Is256Bit = Offset >= SSERegBitSize;
const auto Is256Bit = Offset >= SSERegBitSize;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetReg(Node);
@@ -33,7 +33,7 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, false);
Relocations.emplace_back(MoveABI);
}
@@ -77,7 +77,7 @@ void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constan
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
+20 -124
View File
@@ -138,108 +138,6 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
staddl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
add(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
staddl(SubEmitSize, TMP2, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
sub(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
stclrl(SubEmitSize, TMP2, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
and_(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicCLR) {
auto Op = IROp->C<IR::IROp_AtomicCLR>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
stclrl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
bic(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
stsetl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
orr(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
@@ -260,21 +158,6 @@ DEF_OP(AtomicXor) {
}
}
DEF_OP(AtomicNeg) {
auto Op = IROp->C<IR::IROp_AtomicNeg>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
@@ -439,13 +322,26 @@ DEF_OP(AtomicFetchNeg) {
auto MemSrc = GetReg(Op->Addr);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
if (CTX->HostFeatures.SupportsAtomics) {
// Use a CAS loop to avoid needing to emulate unaligned LLSC atomics
ldr(SubEmitSize, TMP2, MemSrc);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
mov(EmitSize, TMP4, TMP2);
neg(EmitSize, TMP3, TMP2);
casal(SubEmitSize, TMP2, TMP3, MemSrc);
sub(EmitSize, TMP3, TMP2, TMP4);
cbnz(EmitSize, TMP3, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
DEF_OP(TelemetrySetValue) {
+168 -61
View File
@@ -58,31 +58,121 @@ DEF_OP(ExitFunction) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef _M_ARM_64EC
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
} else {
#endif
// Align to 16 byte to allow atomic patching of the following 16 byte
// of code (excluding the RIP data) on platforms that support LSE2
Align16B();
// In order to support direct branches without constantly hitting the L1 cache, we emit a call to a block linker,
// this will compile the branch target block when it is hit and replace the branch to the linker at the callsite
// with a direct branch to the destination block. Upon invalidation of the target block the backpatch is undone.
//
// In addition, to avoid needing to lookup in the cache for returns and any indirect branch prediction penalty,
// a shadow stack of <GuestReturnRIP, HostReturnPC> pairs is maintained, acting as a first level cache for any
// return operations. As the guest may not balance calls and returns exactly, an exception handler is expected to
// be installed by the frontend, to reset the shadow stack to the middle of its valid bounds on overflow/underflow.
// This shadow stack is also cleared on block invalidation operations or codebuffer switches, to ensure all pointed-to
// host code is always valid.
// This code will be backpatched by Arm64JITCore_ExitFunctionLink, below is an enumeration of all the possible cases.
// Jump thunks are emitted in JIT.cpp after compilation of the entire multiblock.
//
// Call with known return block - unlinked
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl JmpThunk00
// JmpThunk00:
// 00: b 0x8
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode
// 18: GuestRIP
// 20: CallerOffset
//
// Call with known return block after backpatching - linked in branch immediate range
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl HostCode - MODIFIED
//
// Call with known return block after backpatching - linked out of range
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl JmpThunk00
// JmpThunk00:
// 00: ldr TMP1, 0x10 - MODIFIED 2nd
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode - MODIFIED 1st
// 18: GuestRIP
// 20: CallerOffset
//
// Jump - unlinked
// 00: b JmpThunk00
// JmpThunk00:
// 00: b 0x8
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode
// 18: GuestRIP
// 20: CallerOffset
//
// Jump after backpatching - linked in branch immediate range
// 00: b HostCode - MODIFIED
//
// Jump after backpatching - linked out of range
// 00: b JmpThunk00
// JmpThunk00:
// 00: ldr TMP1, 0x10 - MODIFIED 2nd
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode - MODIFIED 1st
// 18: GuestRIP
// 20: CallerOffset
ARMEmitter::ForwardLabel l_BranchHost;
ldr(TMP1, &l_BranchHost);
blr(TMP1);
ARMEmitter::ForwardLabel l_CallReturn;
if (Op->Hint == IR::BranchHint::Call) {
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
} else if (Op->Hint == IR::BranchHint::CheckTF) {
ARMEmitter::ForwardLabel TFUnset;
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, NewRIP);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
blr(TMP2);
Bind(&TFUnset);
}
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
dc64(NewRIP);
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
Bind(&l_CallReturn);
#ifdef _M_ARM_64EC
}
#endif
} else {
ARMEmitter::ForwardLabel FullLookup;
ARMEmitter::ForwardLabel SkipFullLookup;
auto RipReg = GetReg(Op->NewRIP);
if (Op->Hint == IR::BranchHint::Return) {
// First try to pop from the call-ret stack, otherwise follow the normal path (but ending in a ret)
ldp<ARMEmitter::IndexType::POST>(TMP1, TMP2, REG_CALLRET_SP, 0x10);
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
}
// L1 Cache
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
@@ -93,36 +183,51 @@ DEF_OP(ExitFunction) {
ubfiz(ARMEmitter::Size::i64Bit, TMP4, RipReg, 4, 20);
add(TMP1, TMP1, TMP4);
// Note: sub+cbnz used over cmp+br to preserve flags.
ldp<ARMEmitter::IndexType::OFFSET>(TMP2, TMP1, TMP1, 0);
sub(TMP1, TMP1, RipReg.X());
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(TMP2);
Bind(&FullLookup);
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
// Note: sub+cbnz used over cmp+br to preserve flags.
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
br(TMP1);
Bind(&SkipFullLookup);
if (Op->Hint == IR::BranchHint::Call) {
ARMEmitter::ForwardLabel l_CallReturn;
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
blr(TMP2);
Bind(&l_CallReturn);
} else if (Op->Hint == IR::BranchHint::Return) {
ret(TMP2);
} else {
br(TMP2);
}
}
}
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto Target = Op->TargetBlock;
PendingTargetLabel = &JumpTargets.try_emplace(Target.ID()).first->second;
PendingTargetLabel = JumpTarget(Op->TargetBlock);
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
auto TrueTargetLabel = JumpTarget(Op->TrueBlock);
if (Op->FromNZCV) {
b(MapCC(Op->Cond), TrueTargetLabel);
} else {
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
uint64_t Const;
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
auto Reg = GetReg(Op->Cmp1);
const auto Size = Op->CompareSize == IR::OpSize::i32Bit ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
@@ -147,7 +252,7 @@ DEF_OP(CondJump) {
}
}
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
PendingTargetLabel = JumpTarget(Op->FalseBlock);
}
DEF_OP(Syscall) {
@@ -341,48 +446,50 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
const auto* OldCode = (const uint8_t*)&Op->CodeOriginalLow;
auto OldCode = Op->CodeOriginal.data();
auto Base = GetReg(Op->Header.Args[0]).X();
int len = Op->CodeLength;
int idx = 0;
LoadConstant(ARMEmitter::Size::i64Bit, GetReg(Node), 0);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Entry + Op->Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, 1);
int Offset = 0;
ARMEmitter::ForwardLabel Fail;
const auto Dst = GetReg(Node);
while (len >= 8) {
ldr(ARMEmitter::XReg::x2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint64_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 8;
idx += 8;
}
while (len >= 4) {
ldr(ARMEmitter::WReg::w2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 4;
idx += 4;
}
while (len >= 2) {
ldrh(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint16_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 2;
idx += 2;
}
while (len >= 1) {
ldrb(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint8_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 1;
idx += 1;
}
auto EmitCheck = [&](size_t Size, auto&& LoadData) {
while (len >= Size) {
LoadData();
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &Fail);
len -= Size;
Offset += Size;
}
};
EmitCheck(8, [&]() {
ldr(TMP1, Base, Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, *(const uint64_t*)(OldCode + Offset));
});
EmitCheck(4, [&]() {
ldr(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint32_t*)(OldCode + Offset));
});
EmitCheck(2, [&]() {
ldrh(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint16_t*)(OldCode + Offset));
});
EmitCheck(1, [&]() {
ldrb(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint8_t*)(OldCode + Offset));
});
ARMEmitter::ForwardLabel End;
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 0);
b(&End);
Bind(&Fail);
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 1);
Bind(&End);
}
DEF_OP(ThreadRemoveCodeEntry) {
@@ -16,7 +16,7 @@ DEF_OP(VAESImc) {
DEF_OP(VAESEnc) {
const auto Op = IROp->C<IR::IROp_VAESEnc>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -41,7 +41,7 @@ DEF_OP(VAESEnc) {
DEF_OP(VAESEncLast) {
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -64,7 +64,7 @@ DEF_OP(VAESEncLast) {
DEF_OP(VAESDec) {
const auto Op = IROp->C<IR::IROp_VAESDec>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -89,7 +89,7 @@ DEF_OP(VAESDec) {
DEF_OP(VAESDecLast) {
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -322,7 +322,7 @@ DEF_OP(VSha256U1) {
DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
+190 -72
View File
@@ -222,7 +222,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
FillF64Result();
} break;
case FABI_F64_I16_F64_PTR: {
case FABI_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -239,7 +239,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
FillF64Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
case FABI_F64x2_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -264,7 +264,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
FillF64x2Result(DstLo, DstHi);
} break;
case FABI_F64_I16_F64_F64_PTR: {
case FABI_F64_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -495,58 +495,116 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
// Emit new 16 bytes of code to a temporary patch, then atomically apply it
__uint128_t Patch;
ARMEmitter::Emitter emit((uint8_t*)&Patch, sizeof(Patch));
emit.ldr(TMP1, 8); // PC-relative value pointing to constant after blr
emit.blr(TMP1);
emit.dc64(Frame->Pointers.Common.ExitFunctionLinker);
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
auto branch = reinterpret_cast<__uint128_t*>((uintptr_t)Record - 8);
std::atomic_ref<__uint128_t>(*branch).store(Patch, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache((void*)branch, sizeof(*branch));
// Replace the patched callsite with a branch to the jump thunk.
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
if (Call) {
BranchEmit.bl(BranchOffset);
} else {
BranchEmit.b(BranchOffset);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
auto Lock = Thread->LookupCache->AcquireLock();
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
// No need to reset HostCode here as the exit linker pointer is stored separately, and if the block is relinked it will be updated.
}
uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
bool TFSet = Thread->CurrentFrame->State.flags[X86State::RFLAG_TF_RAW_LOC];
uintptr_t HostCode {};
auto GuestRip = Record->GuestRIP;
if (!TFSet) {
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (TFSet || !HostCode) {
if (TFSet) {
// If TF is set, the cache must be skipped as different code needs to be generated.
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
} else {
{
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (!HostCode) {
// Hold a reference to the code buffer, to avoid linking unmapped code if compilation triggers a recreation.
auto CodeBuffer = static_cast<Arm64JITCore*>(Thread->CPUBackend.get())->CurrentCodeBuffer;
HostCode = static_cast<Context::ContextImpl*>(Thread->CTX)->CompileBlock(Frame, GuestRip, 0);
if (Thread->LookupCache->Shared != CodeBuffer->LookupCache.get()) {
return HostCode;
}
}
}
uintptr_t branch = (uintptr_t)(Record)-8;
LOGMAN_THROW_A_FMT((branch % 16) == 0, "Incorrect alignment for block linking record");
// See ExitFunction in BranchOps.cpp for an assembly level view of the handled cases.
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = HostCode / 4 - CallerAddress / 4;
auto offset = HostCode / 4 - branch / 4;
if (ARMEmitter::Emitter::IsInt26(offset)) {
// This is the optimal case, where the target can be encoded in a single instruction.
// Atomically patch the code with a relative branch.
const uint32_t Patch = (0b0001'01 << 26) | (offset & ((1u << 26) - 1));
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(branch)).store(Patch, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache((void*)branch, 4);
uint32_t ExpectedKnownCallMarkerInst = 0;
ARMEmitter::Emitter ExpectedKnownCallMarkerEmit(reinterpret_cast<uint8_t*>(&ExpectedKnownCallMarkerInst), 4);
ExpectedKnownCallMarkerEmit.adr(TMP1, 0xC);
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
// Lock here is necessary to prevent simultaneous linking and delinking
auto lk = Thread->LookupCache->AcquireLock();
// For non-calls, this would extend into the block's code, however that's fine as an out-of-range adr would never
// be generated avoiding any false positives.
uintptr_t KnownCallMarkerAddr = CallerAddress - 0x8;
uint32_t KnownCallMarkerInst = *reinterpret_cast<uint32_t*>(KnownCallMarkerAddr);
if (ARMEmitter::Emitter::IsInt26(BranchOffset)) {
// Directly patch the callsite with the appropriate branch instruction.
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, true);
});
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
});
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
} else {
// fallback case - do a soft-er link by patching the pointer
std::atomic_ref<uint64_t>(Record->HostBranch).store(HostCode, std::memory_order::seq_cst);
// This case is common between calls and jumps as the thunk callsite can be left untouched.
std::atomic_ref<uint64_t>(Record->HostCode).store(HostCode, std::memory_order::seq_cst);
#ifdef _M_ARM_64
// Make memory write visible to other threads reading the same location
asm volatile("dc cvau, %0; dsb ish" : : "r"(Record->HostBranch) :);
asm volatile("dc cvau, %0; dsb ish" : : "r"(Record->HostCode) :);
#endif
}
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
uint32_t LdrInst = 0;
ARMEmitter::Emitter LdrEmit(reinterpret_cast<uint8_t*>(&LdrInst), 4);
LdrEmit.ldr(TMP1, reinterpret_cast<uint64_t>(&Record->HostCode) - JumpThunkStartAddress);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(LdrInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
Thread->LookupCache->AddBlockLink(GuestRip, Record, IndirectBlockDelinker);
}
return HostCode;
}
@@ -581,6 +639,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -598,7 +657,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = PMF.GetVTableEntry(CTX->SyscallHandler);
}
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore::ExitFunctionLink);
// Platform Specific
auto& AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
@@ -741,19 +800,43 @@ void Arm64JITCore::EmitInterruptChecks(bool CheckTF) {
#endif
}
void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF) {
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &HeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
EmitInterruptChecks(CheckTF);
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
CallReturnTargets.clear();
PendingJumpThunks.clear();
uint32_t SSACount = IR->GetSSACount();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = 0x100 + SSACount * 24;
uint32_t BufferRange = 0x1000 + SSACount * 24;
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
@@ -768,9 +851,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
JITCodeHeader* CodeHeader = GetCursorAddress<JITCodeHeader*>();
CursorIncrement(sizeof(JITCodeHeader));
#ifdef VIXL_DISASSEMBLER
const auto DisasmBegin = GetCursorAddress<const vixl::aarch64::Instruction*>();
#endif
auto CodeBegin = GetCursorAddress<uint8_t*>();
// AAPCS64
// r30 = LR
@@ -792,49 +873,57 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// X1-X3 = Temp
// X4-r18 = RA
CodeData.BlockEntry = GetCursorAddress<uint8_t*>();
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &JITCodeHeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
EmitInterruptChecks(CheckTF);
SpillSlots = IR->SpillSlots();
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
PendingTargetLabel = nullptr;
PendingCallReturnTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = GetCursorAddress<uint8_t*>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
const auto Target = &JumpTargets[BlockIROp->ID];
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
if (PendingTargetLabel && PendingTargetLabel != Target) {
b(PendingTargetLabel);
PendingTargetLabel = nullptr;
}
if (BlockIROp->EntryPoint) {
uint64_t BlockStartRIP = Entry + BlockIROp->GuestEntryOffset;
const auto IsReturnTarget = CallReturnTargets.try_emplace(Node).first;
if (PendingTargetLabel) {
// If there is a fallthrough branch to this block, skip over the entrypoint code.
b(Target);
} else if (PendingCallReturnTargetLabel && PendingCallReturnTargetLabel != &IsReturnTarget->second) {
// If we just emitted a call, but the block we're now emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
}
PendingCallReturnTargetLabel = nullptr;
Bind(&IsReturnTarget->second);
CodeData.EntryPoints.emplace(BlockStartRIP, GetCursorAddress<uint8_t*>());
DebugData->GuestOpcodes.push_back({BlockIROp->GuestEntryOffset, GetCursorAddress<uint8_t*>() - CodeData.BlockBegin});
EmitEntryPoint(JITCodeHeaderLabel, CheckTF);
}
if (PendingCallReturnTargetLabel) {
// If there is still a pending call return target, then the block we're emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
PendingCallReturnTargetLabel = nullptr;
}
PendingTargetLabel = nullptr;
Bind(&IsTarget->second);
Bind(Target);
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
@@ -852,7 +941,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
DebugData->Subblocks.push_back({static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockEntry),
DebugData->Subblocks.push_back({static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockBegin),
static_cast<uint32_t>(GetCursorAddress<uint8_t*>() - BlockStartHostCode)});
}
@@ -862,8 +951,32 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
PendingTargetLabel = nullptr;
// CodeSize not including the tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
ARMEmitter::ForwardLabel l_ExitLink;
for (auto& PendingJumpThunk : PendingJumpThunks) {
// Align as 64-bit atomics are used on the HostCode field.
Align(8);
ARMEmitter::ForwardLabel l_DoLink;
uint64_t ThunkAddress = GetCursorAddress<uint64_t>();
Bind(&PendingJumpThunk.Label);
b(&l_DoLink);
br(TMP1);
Bind(&l_DoLink);
ldr(TMP1, &l_ExitLink);
blr(TMP1);
// This is a ExitFunctionLinkData struct
Bind(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
Bind(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
Align(alignof(JITCodeTail));
@@ -942,6 +1055,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
LOGMAN_THROW_A_FMT(CurrentCodeBuffer->LookupCache.get() == ThreadState->LookupCache->Shared, "INVARIANT VIOLATED: SharedLookupCache "
"doesn't match up!\n");
if (auto Prev = CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(ThreadState->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache);
}
@@ -960,7 +1074,10 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Adjust host addresses
const auto Delta = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
CodeData.BlockBegin += Delta;
CodeData.BlockEntry += Delta;
for (auto& EntryPoint : CodeData.EntryPoints) {
EntryPoint.second += Delta;
}
CodeBegin += Delta;
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
@@ -971,7 +1088,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
TempAllocator.DelayedDisownBuffer();
ClearICache(CodeData.BlockBegin, CodeOnlySize);
ClearICache(CodeBegin, CodeOnlySize);
#ifdef VIXL_DISASSEMBLER
if (Disassemble() & FEXCore::Config::Disassemble::STATS) {
@@ -985,7 +1102,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
if (Disassemble() & FEXCore::Config::Disassemble::BLOCKS) {
const auto DisasmEnd = reinterpret_cast<const vixl::aarch64::Instruction*>(JITBlockTailLocation);
const auto DisasmBegin = reinterpret_cast<const vixl::aarch64::Instruction*>(CodeBegin);
const auto DisasmEnd = reinterpret_cast<const vixl::aarch64::Instruction*>(CodeBegin + CodeOnlySize);
LogMan::Msg::IFmt("Disassemble Begin");
for (auto PCToDecode = DisasmBegin; PCToDecode < DisasmEnd; PCToDecode += 4) {
DisasmDecoder->Decode(PCToDecode);
@@ -1001,7 +1119,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
this->IR = nullptr;
return CodeData;
return std::move(CodeData);
}
void Arm64JITCore::ResetStack() {
+40 -5
View File
@@ -24,6 +24,7 @@ $end_info$
#include <array>
#include <cstdint>
#include <functional>
#include <utility>
#include <variant>
@@ -31,6 +32,10 @@ namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Context {
struct ExitFunctionLinkData;
}
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
@@ -57,15 +62,32 @@ private:
const bool HostSupportsAFP {};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
FEXCore::Context::ContextImpl* CTX {};
const FEXCore::IR::IRListView* IR {};
uint64_t Entry {};
CPUBackend::CompiledCode CodeData {};
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
fextl::vector<ARMEmitter::BiDirectionalLabel> JumpTargets;
ARMEmitter::BiDirectionalLabel* JumpTarget(IR::OrderedNodeWrapper Node) {
auto Block = IR->GetOp<IR::IROp_CodeBlock>(Node);
return &JumpTargets[Block->ID];
}
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> CallReturnTargets;
struct PendingJumpThunk {
uint64_t CallerAddress;
uint64_t GuestRIP;
ARMEmitter::ForwardLabel Label;
};
fextl::vector<PendingJumpThunk> PendingJumpThunks;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempAllocator;
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
[[nodiscard]]
ARMEmitter::Register GetReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
@@ -223,10 +245,10 @@ private:
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_FU:
case FEXCore::IR::COND_VS: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU:
case FEXCore::IR::COND_VC: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
@@ -290,6 +312,17 @@ private:
uint32_t End;
};
void EmitLinkedBranch(uint64_t GuestRIP, bool Call) {
PendingJumpThunks.push_back({GetCursorAddress<uint64_t>(), GuestRIP, {}});
auto& Thunk = PendingJumpThunks.back();
Bind(&Thunk.Label);
if (Call) {
bl(&Thunk.Label);
} else {
b(&Thunk.Label);
}
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass {};
@@ -375,6 +408,8 @@ private:
void EmitInterruptChecks(bool CheckTF);
void EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
@@ -140,7 +140,7 @@ DEF_OP(LoadRegister) {
mov(GetReg(Node).X(), StaticRegisters[Op->Reg].X());
} else if (Op->Class == IR::FPRClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "out of range reg");
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
@@ -181,7 +181,7 @@ DEF_OP(StoreRegister) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
mov(ARMEmitter::Size::i64Bit, GetReg(Reg), GetReg(Op->Value));
} else if (Reg.Class == IR::FPRFixedClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
const auto guest = GetVReg(Reg);
@@ -742,7 +742,7 @@ DEF_OP(LoadMemTSO) {
const auto Dst = GetReg(Node);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
[[maybe_unused]] bool IsInline = IsInlineConstant(Op->Offset, &Offset);
bool IsInline = IsInlineConstant(Op->Offset, &Offset);
LOGMAN_THROW_A_FMT(IsInline, "expected immediate");
}
@@ -1750,7 +1750,7 @@ DEF_OP(StoreMemTSO) {
const auto Src = GetZeroableReg(Op->Value);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
[[maybe_unused]] bool IsInline = IsInlineConstant(Op->Offset, &Offset);
bool IsInline = IsInlineConstant(Op->Offset, &Offset);
LOGMAN_THROW_A_FMT(IsInline, "expected immediate");
}
@@ -2586,7 +2586,7 @@ DEF_OP(VStoreNonTemporalPair) {
const auto Op = IROp->C<IR::IROp_VStoreNonTemporalPair>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Is128Bit, "This IR operation only operates at 128-bit wide");
const auto ValueLow = GetVReg(Op->ValueLow);
@@ -17,6 +17,26 @@ $end_info$
namespace FEXCore::CPU {
DEF_OP(WFET) {
auto Op = IROp->C<IR::IROp_WFET>();
const auto Lower = GetReg(Op->Lower);
const auto Upper = GetReg(Op->Upper);
// Combine registers.
mov(ARMEmitter::Size::i64Bit, TMP1, Lower);
bfi(ARMEmitter::Size::i64Bit, TMP1, Upper, 32, 32);
if (CTX->Config.TSCScale) {
// Scale back to ARM64 TSC scale if necessary
lsr(ARMEmitter::Size::i64Bit, TMP1, TMP1, CTX->Config.TSCScale);
}
// Clear the exclusive monitor so it can't spuriously wake up with that event.
clrex();
// Execute wfet to wait until the TSC.
wfet(TMP1);
}
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
@@ -266,4 +286,43 @@ DEF_OP(Yield) {
yield();
}
DEF_OP(MonoBackpatcherWrite) {
auto Op = IROp->C<IR::IROp_MonoBackpatcherWrite>();
mov(ARMEmitter::Size::i64Bit, TMP3, GetReg(Op->Addr));
mov(ARMEmitter::Size::i64Bit, TMP4, GetReg(Op->Value));
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, IR::OpSizeToSize(Op->Size));
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, TMP3);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, TMP4);
}
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 1);
strb(TMP1.W(), TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
ldr(ARMEmitter::XReg::x4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.MonoBackpatcherWrite));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, uint8_t, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4);
}
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
FillStaticRegs();
PopDynamicRegs();
}
} // namespace FEXCore::CPU
+32 -32
View File
@@ -193,29 +193,29 @@ namespace FEXCore::CPU {
VFScalarOperation(IROp->Size, ElementSize, Op->ZeroUpperBits, ScalarEmit, Dst, Vector1, Vector2); \
}
#define DEF_FMAOP_SCALAR_INSERT(FEXOp, ARMOp) \
DEF_OP(FEXOp) { \
const auto Op = IROp->C<IR::IROp_##FEXOp>(); \
const auto ElementSize = Op->Header.ElementSize; \
\
auto ScalarEmit = \
[this, ElementSize](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3) { \
if (ElementSize == IR::OpSize::i16Bit) { \
ARMOp(Dst.H(), Src1.H(), Src2.H(), Src3.H()); \
} else if (ElementSize == IR::OpSize::i32Bit) { \
ARMOp(Dst.S(), Src1.S(), Src2.S(), Src3.S()); \
} else if (ElementSize == IR::OpSize::i64Bit) { \
ARMOp(Dst.D(), Src1.D(), Src2.D(), Src3.D()); \
} \
}; \
\
const auto Dst = GetVReg(Node); \
const auto Upper = GetVReg(Op->Upper); \
const auto Vector1 = GetVReg(Op->Vector1); \
const auto Vector2 = GetVReg(Op->Vector2); \
const auto Addend = GetVReg(Op->Addend); \
\
VFScalarFMAOperation(IROp->Size, ElementSize, ScalarEmit, Dst, Upper, Vector1, Vector2, Addend); \
#define DEF_FMAOP_SCALAR_INSERT(FEXOp, ARMOp) \
DEF_OP(FEXOp) { \
const auto Op = IROp->C<IR::IROp_##FEXOp>(); \
const auto ElementSize = Op->Header.ElementSize; \
\
auto ScalarEmit = [this, ElementSize](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, \
ARMEmitter::VRegister Src3) { \
if (ElementSize == IR::OpSize::i16Bit) { \
ARMOp(Dst.H(), Src1.H(), Src2.H(), Src3.H()); \
} else if (ElementSize == IR::OpSize::i32Bit) { \
ARMOp(Dst.S(), Src1.S(), Src2.S(), Src3.S()); \
} else if (ElementSize == IR::OpSize::i64Bit) { \
ARMOp(Dst.D(), Src1.D(), Src2.D(), Src3.D()); \
} \
}; \
\
const auto Dst = GetVReg(Node); \
const auto Upper = GetVReg(Op->Upper); \
const auto Vector1 = GetVReg(Op->Vector1); \
const auto Vector2 = GetVReg(Op->Vector2); \
const auto Addend = GetVReg(Op->Addend); \
\
VFScalarFMAOperation(IROp->Size, ElementSize, ScalarEmit, Dst, Upper, Vector1, Vector2, Addend); \
}
DEF_UNOP(VAbs, abs, true)
@@ -803,8 +803,8 @@ DEF_OP(VFCMPScalarInsert) {
default: break;
}
};
auto ScalarEmitUNO =
[this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2) {
auto ScalarEmitUNO = [this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1,
ARMEmitter::VRegister Src2) {
switch (SubRegSize.Scalar) {
case ARMEmitter::ScalarRegSize::i16Bit: {
fcmge(VTMP1.H(), Src1.H(), Src2.H());
@@ -838,8 +838,8 @@ DEF_OP(VFCMPScalarInsert) {
}
}
};
auto ScalarEmitNEQ =
[this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2) {
auto ScalarEmitNEQ = [this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1,
ARMEmitter::VRegister Src2) {
switch (SubRegSize.Scalar) {
case ARMEmitter::ScalarRegSize::i16Bit: {
fcmeq(VTMP1.H(), Src2.H(), Src1.H());
@@ -868,8 +868,8 @@ DEF_OP(VFCMPScalarInsert) {
}
}
};
auto ScalarEmitORD =
[this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2) {
auto ScalarEmitORD = [this, SubRegSize, ZeroUpperBits, Is256Bit](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1,
ARMEmitter::VRegister Src2) {
switch (SubRegSize.Scalar) {
case ARMEmitter::ScalarRegSize::i16Bit: {
fcmge(VTMP1.H(), Src1.H(), Src2.H());
@@ -1115,7 +1115,7 @@ DEF_OP(VAddP) {
}
DEF_OP(VFAddV) {
const auto Op = IROp->C<IR::IROp_VAddV>();
const auto Op = IROp->C<IR::IROp_VFAddV>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -1352,7 +1352,7 @@ DEF_OP(VFMin) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubRegSize = ConvertSubRegSize248(IROp);
[[maybe_unused]] const auto IsScalar = ElementSize == OpSize;
const auto IsScalar = ElementSize == OpSize;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -1425,7 +1425,7 @@ DEF_OP(VFMax) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubRegSize = ConvertSubRegSize248(IROp);
[[maybe_unused]] const auto IsScalar = ElementSize == OpSize;
const auto IsScalar = ElementSize == OpSize;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
+24 -19
View File
@@ -56,6 +56,8 @@ struct GuestToHostMap {
fextl::robin_map<uint64_t, uint64_t> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
@@ -76,7 +78,7 @@ struct GuestToHostMap {
return HostCode->second;
}
void Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LockToken&) {
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
@@ -85,7 +87,7 @@ struct GuestToHostMap {
}
// Remove from BlockList
BlockList.erase(Address);
return BlockList.erase(Address) != 0;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -93,6 +95,18 @@ struct GuestToHostMap {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
}
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length, const LockToken&) {
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto& CodePage = CodePages[CurrentPage];
rv |= CodePage.empty();
CodePage.insert(CodePage.end(), Addresses.begin(), Addresses.end());
}
return rv;
}
void ClearCache(const LockToken&);
};
@@ -155,22 +169,11 @@ public:
GuestToHostMap* Shared = nullptr;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
// Appends Block {Address} to CodePages [Start, Start + Length)
// Appends a list of Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireLock();
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto& CodePage = CodePages[CurrentPage];
rv |= CodePage.empty();
CodePage.push_back(Address);
}
return rv;
return Shared->AddBlockExecutableRange(Addresses, Start, Length, lk);
}
// Adds to Guest -> Host code mapping
@@ -189,15 +192,16 @@ public:
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
void Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
auto lk = Shared->AcquireLock();
Shared->Erase(Frame, Address, lk);
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -212,13 +216,14 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return;
return ErasedAny;
}
// Page exists, just set the offset to zero
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
return true;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink, const FEXCore::Context::BlockDelinkerFunc& delinker) {
@@ -1,86 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct FEX_PACKED CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie {};
// Instructions per block configuration
int32_t MaxInstPerBlock {};
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Multiblock enabled
unsigned MultiBlock : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// TSO enabled
unsigned TSOEnabled : 1;
// ABI local flag unsafe optimization
unsigned ABILocalFlags : 1;
// Paranoid TSO mode enabled
unsigned ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
unsigned Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
unsigned x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 19;
bool operator==(const CodeObjectSerializationConfig& other) const {
return Cookie == other.Cookie && MaxInstPerBlock == other.MaxInstPerBlock && Arch == other.Arch && MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled && TSOEnabled == other.TSOEnabled && ABILocalFlags == other.ABILocalFlags &&
ParanoidTSO == other.ParanoidTSO && Is64BitMode == other.Is64BitMode && SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(const CodeObjectSerializationConfig& other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash {};
Hash <<= 32;
Hash |= other.MaxInstPerBlock;
Hash <<= 1;
Hash |= other.Arch;
Hash <<= 1;
Hash |= other.MultiBlock;
Hash <<= 1;
Hash |= other.HardwareTSOEnabled;
Hash <<= 1;
Hash |= other.TSOEnabled;
Hash <<= 1;
Hash |= other.ABILocalFlags;
Hash <<= 1;
Hash |= other.ParanoidTSO;
Hash <<= 1;
Hash |= other.Is64BitMode;
Hash <<= 2;
Hash |= other.SMCChecks;
Hash <<= 1;
Hash |= other.x87ReducedPrecision;
return Hash;
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is "
"generated!");
} // namespace FEXCore::CodeSerialize
@@ -1,121 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <fcntl.h>
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
#ifndef _WIN32
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
const fextl::string BaseFilename = FHU::Filesystem::GetFilename(filename);
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = fextl::make_unique<CodeRegionEntry>(Base, Size, Offset, filename, NamedRegionHandler->DefaultCodeHeader(Base, Offset));
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
} else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
#ifndef _WIN32
// Removing a named region through the job system
// We need to find the entry that we are deleting first
fextl::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
// Take the pointer from the map
EntryPointer = std::move(it->second);
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Remove this from the entry map
EntryMap.erase(it);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
}
} else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
} // namespace FEXCore::CodeSerialize
@@ -1,71 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename,
const fextl::string& filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion*>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion*>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
}
}
}
} // namespace FEXCore::CodeSerialize
@@ -1,85 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/Threads.h>
namespace {
static void* ThreadHandler(void* Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService* This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
} // namespace
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, AsyncHandler {&NamedRegionHandler, this}
, NamedRegionHandler {ctx} {
Initialize();
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
WorkerThreadShuttingDown = true;
// Kick the working thread
WorkAvailable.NotifyAll();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, fextl::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
const CodeObjectFileSection* CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
FEXCore::Threads::SetThreadName("ObjectCodeSeri\0");
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto& it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
} // namespace FEXCore::CodeSerialize
@@ -1,457 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "Interface/Core/ObjectCache/CodeObjectSerializationConfig.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/queue.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <shared_mutex>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData* Data;
const char* HostCode;
uint64_t NumRelocations;
const char* Relocations;
};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase {};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset {};
// Total amount of code that should be in this file
uint64_t TotalCodeSize {};
// Used to reserve the TSL map
uint64_t NumCodeEntries {};
// The number of relocations that point to this section
uint64_t NumRelocationsTo {};
// Total relocations in this file
uint64_t TotalRelocationsCount {};
};
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base {};
// Size of this code entry
uint64_t Size {};
// The offset inside the file that is mapped to Base
uint64_t Offset {};
// Filename of the object
fextl::string Filename {};
CodeObjectSerializationHeader EntryHeader {};
/** @} */
// The filename of the object cache for this entry
fextl::string ObjectEntrySourceFilename {};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char* CodeData {};
size_t FileSize {};
fextl::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
fextl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap {};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base, uint64_t Size, uint64_t Offset, const fextl::string& Filename, const CodeObjectSerializationHeader& DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = fextl::map<uint64_t, fextl::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = fextl::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
void* HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex* ThreadJobRefCount;
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex* ObjectJobRefCountMutexPtr;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
AsyncJobHandler(NamedRegionObjectHandler* NamedRegionHandler, CodeObjectSerializeService* CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const {
return Type;
}
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry} {}
const fextl::string BaseFilename;
const fextl::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, fextl::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
fextl::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler* NamedRegionHandler;
CodeObjectSerializeService* CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx);
void HandleNamedRegionObjectJobs();
const CodeObjectSerializationConfig& GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemAddNamedRegion>(base, filename, executable, entry));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion>(Base, Size, std::move(Entry)));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs {};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex {};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
fextl::queue<fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue {};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename, const fextl::string& filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx);
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(fextl::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex* ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
const CodeObjectFileSection* FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it);
CodeSerializationMutex& GetEntryMapMutex() {
return EntryMapMutex;
}
CodeSerializationMutex& GetUnrelocatedEntryMapMutex() {
return EntryMapMutex;
}
CodeRegionMapType& GetEntryMap() {
return AddressToEntryMap;
}
CodeRegionPtrMapType& GetUnrelocatedEntryMap() {
return UnrelocatedAddressToEntryMap;
}
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() {
WorkAvailable.NotifyOne();
}
private:
FEXCore::Context::ContextImpl* CTX;
Event WorkAvailable {};
fextl::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
};
} // namespace FEXCore::CodeSerialize
File diff suppressed because it is too large. Load diff
+194 -158
View File
@@ -27,8 +27,6 @@
#include <xxhash.h>
namespace FEXCore::IR {
class Pass;
class PassManager;
enum class MemoryAccessType {
// Choose TSO or Non-TSO depending on access type
@@ -80,10 +78,13 @@ struct LoadSourceOptions {
bool AllowUpperGarbage = false;
};
class OpDispatchBuilder final : public IREmitter {
friend class FEXCore::IR::Pass;
friend class FEXCore::IR::PassManager;
struct DispatchTableEntry {
uint16_t Op;
uint8_t Count;
X86Tables::OpDispatchPtr Ptr;
};
class OpDispatchBuilder final : public IREmitter {
public:
Ref GetNewJumpBlock(uint64_t RIP) {
auto it = JumpTargets.find(RIP);
@@ -160,9 +161,13 @@ public:
auto InlineConst = _InlineConstant(Bit);
return _CondJump(Src, InlineConst, InvalidNode, InvalidNode, {Set ? COND_TSTNZ : COND_TSTZ}, OpSize::iInvalid, false);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP) {
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint = BranchHint::None) {
FlushRegisterCache();
return _ExitFunction(GetOpSize(NewRIP), NewRIP);
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, InvalidNode, InvalidNode);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
FlushRegisterCache();
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, CallReturnAddress, CallReturnBlock);
}
IRPair<IROp_Break> Break(BreakDefinition Reason) {
FlushRegisterCache();
@@ -188,11 +193,10 @@ public:
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end()) {
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
// If we don't have a jump target to a new block then we have to leave
// Set the RIP to the next instruction and leave
auto RelocatedNextRIP = _EntrypointOffset(GPRSize, NextRIP - Entry);
ExitFunction(RelocatedNextRIP);
ExitFunction(_InlineEntrypointOffset(GPRSize, NextRIP - Entry));
} else if (it != JumpTargets.end()) {
Jump(it->second.BlockEntry);
return true;
@@ -232,7 +236,7 @@ public:
template<typename F>
void ForeachDirection(F&& Routine) {
// Otherwise, prepare to branch.
auto Zero = _Constant(0);
auto Zero = Constant(0);
// If the shift is zero, do not touch the flags.
auto ForwardBlock = CreateNewCodeBlockAfter(GetCurrentBlock());
@@ -256,7 +260,6 @@ public:
}
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator& Allocator);
void ResetWorkingList();
void ResetDecodeFailure() {
@@ -290,7 +293,7 @@ public:
return ShouldDump;
}
void BeginFunction(uint64_t RIP, const fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks>* Blocks, uint32_t NumInstructions);
void BeginFunction(uint64_t RIP, const fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks>* Blocks, uint32_t NumInstructions, bool Is64BitMode, bool MonoBackpatcherBlock);
void Finalize();
// Dispatch builder functions
@@ -340,6 +343,9 @@ public:
void LoopOp(OpcodeArgs);
void JUMPOp(OpcodeArgs);
void JUMPAbsoluteOp(OpcodeArgs);
void JUMPFARIndirectOp(OpcodeArgs);
void CALLFARIndirectOp(OpcodeArgs);
void RETFARIndirectOp(OpcodeArgs);
void TESTOp(OpcodeArgs, uint32_t SrcIndex);
void MOVSXDOp(OpcodeArgs);
void MOVSXOp(OpcodeArgs);
@@ -702,7 +708,7 @@ public:
Ref ReconstructX87StateFromFSW_Helper(Ref FSW);
void FLD(OpcodeArgs, IR::OpSize Width);
void FLDFromStack(OpcodeArgs);
void FLD_Const(OpcodeArgs, NamedVectorConstant Constant);
void FLD_Const(OpcodeArgs, NamedVectorConstant K);
void FBLD(OpcodeArgs);
void FBSTP(OpcodeArgs);
@@ -725,6 +731,7 @@ public:
void FADD(OpcodeArgs, IR::OpSize Width, bool Integer, OpResult ResInST0);
void FDIV(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpResult ResInST0);
void FMUL(OpcodeArgs, IR::OpSize Width, bool Integer, OpResult ResInST0);
void FNCLEX(OpcodeArgs);
void FNINIT(OpcodeArgs);
void FSUB(OpcodeArgs, IR::OpSize Width, bool Integer, bool Reverse, OpResult ResInST0);
void FTST(OpcodeArgs);
@@ -824,11 +831,13 @@ public:
void PHADDS(OpcodeArgs);
void PHSUBS(OpcodeArgs);
void CLWB(OpcodeArgs);
void CLWBOrTPause(OpcodeArgs);
void CLFLUSHOPT(OpcodeArgs);
void LoadFenceOrXRSTOR(OpcodeArgs);
void MemFenceOrXSAVEOPT(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void UMonitorOrCLRSSBSY(OpcodeArgs);
void UMWaitOp(OpcodeArgs);
void CLZeroOp(OpcodeArgs);
void RDTSCPOp(OpcodeArgs);
void RDPIDOp(OpcodeArgs);
@@ -837,8 +846,6 @@ public:
void PSADBW(OpcodeArgs);
Ref BitwiseAtLeastTwo(Ref A, Ref B, Ref C);
void SHA1NEXTEOp(OpcodeArgs);
void SHA1MSG1Op(OpcodeArgs);
void SHA1MSG2Op(OpcodeArgs);
@@ -928,7 +935,6 @@ public:
RefVSIB AVX128_LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags, bool NeedsHigh);
void AVX128_StoreResult_WithOpSize(FEXCore::X86Tables::DecodedOp Op, const FEXCore::X86Tables::DecodedOperand& Operand, const RefPair Src,
MemoryAccessType AccessType = MemoryAccessType::DEFAULT);
void InstallAVX128Handlers();
void AVX128_VMOVScalarImpl(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VectorALU(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVX128_VectorUnary(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
@@ -957,39 +963,26 @@ public:
void AVX128_VMOVDDUP(OpcodeArgs);
void AVX128_VMOVSLDUP(OpcodeArgs);
void AVX128_VMOVSHDUP(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VBROADCAST(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VPUNPCKL(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VPUNPCKH(OpcodeArgs);
void AVX128_VBROADCAST(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPUNPCKL(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPUNPCKH(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_MOVVectorUnaligned(OpcodeArgs);
template<IR::OpSize DstElementSize>
void AVX128_InsertCVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void AVX128_CVTFPR_To_GPR(OpcodeArgs);
void AVX128_InsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
void AVX128_CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void AVX128_VANDN(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VPACKSS(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VPACKUS(OpcodeArgs);
void AVX128_VPACKSS(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPACKUS(OpcodeArgs, IR::OpSize ElementSize);
Ref AVX128_PSIGNImpl(IR::OpSize ElementSize, Ref Src1, Ref Src2);
template<IR::OpSize ElementSize>
void AVX128_VPSIGN(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_UCOMISx(OpcodeArgs);
void AVX128_VPSIGN(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VectorScalarInsertALU(OpcodeArgs, FEXCore::IR::IROps IROp, IR::OpSize ElementSize);
Ref AVX128_VFCMPImpl(IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t CompType);
template<IR::OpSize ElementSize>
void AVX128_VFCMP(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_InsertScalarFCMP(OpcodeArgs);
void AVX128_VFCMP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_InsertScalarFCMP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_MOVBetweenGPR_FPR(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_PExtr(OpcodeArgs);
void AVX128_PExtr(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_ExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
template<IR::OpSize ElementSize>
void AVX128_MOVMSK(OpcodeArgs);
void AVX128_MOVMSK(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_MOVMSKB(OpcodeArgs);
void AVX128_PINSRImpl(OpcodeArgs, IR::OpSize ElementSize, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op, const X86Tables::DecodedOperand& Imm);
@@ -1003,33 +996,25 @@ public:
void AVX128_VINSERTPS(OpcodeArgs);
Ref AVX128_PHSUBImpl(Ref Src1, Ref Src2, size_t ElementSize);
template<IR::OpSize ElementSize>
void AVX128_VPHSUB(OpcodeArgs);
void AVX128_VPHSUB(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPHSUBSW(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VADDSUBP(OpcodeArgs);
void AVX128_VADDSUBP(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize, bool Signed>
void AVX128_VPMULL(OpcodeArgs);
void AVX128_VPMULL(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
void AVX128_VPMULHRSW(OpcodeArgs);
template<bool Signed>
void AVX128_VPMULHW(OpcodeArgs);
void AVX128_VPMULHW(OpcodeArgs, bool Signed);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVX128_InsertScalar_CVT_Float_To_Float(OpcodeArgs);
void AVX128_InsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVX128_Vector_CVT_Float_To_Float(OpcodeArgs);
void AVX128_Vector_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void AVX128_Vector_CVT_Float_To_Int(OpcodeArgs);
void AVX128_Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
template<IR::OpSize SrcElementSize, bool Widen>
void AVX128_Vector_CVT_Int_To_Float(OpcodeArgs);
void AVX128_Vector_CVT_Int_To_Float(OpcodeArgs, IR::OpSize SrcElementSize, bool Widen);
void AVX128_VEXTRACT128(OpcodeArgs);
void AVX128_VAESImc(OpcodeArgs);
@@ -1046,36 +1031,28 @@ public:
void AVX128_PHMINPOSUW(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VectorRound(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_InsertScalarRound(OpcodeArgs);
void AVX128_VectorRound(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_InsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVX128_VDPP(OpcodeArgs);
void AVX128_VDPP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPERMQ(OpcodeArgs);
void AVX128_VPSHUFW(OpcodeArgs, bool Low);
template<IR::OpSize ElementSize>
void AVX128_VSHUF(OpcodeArgs);
void AVX128_VSHUF(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVX128_VPERMILImm(OpcodeArgs);
void AVX128_VPERMILImm(OpcodeArgs, IR::OpSize ElementSize);
template<IROps IROp, IR::OpSize ElementSize>
void AVX128_VHADDP(OpcodeArgs);
void AVX128_VHADDP(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVX128_VPHADDSW(OpcodeArgs);
void AVX128_VPMADDUBSW(OpcodeArgs);
void AVX128_VPMADDWD(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VBLEND(OpcodeArgs);
void AVX128_VBLEND(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVX128_VHSUBP(OpcodeArgs);
void AVX128_VHSUBP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPSHUFB(OpcodeArgs);
void AVX128_VPSADBW(OpcodeArgs);
@@ -1086,28 +1063,23 @@ public:
void AVX128_VMASKMOVImpl(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstSize, bool IsStore, const X86Tables::DecodedOperand& MaskOp,
const X86Tables::DecodedOperand& DataOp);
template<bool IsStore>
void AVX128_VPMASKMOV(OpcodeArgs);
void AVX128_VPMASKMOV(OpcodeArgs, bool IsStore);
template<IR::OpSize ElementSize, bool IsStore>
void AVX128_VMASKMOV(OpcodeArgs);
void AVX128_VMASKMOV(OpcodeArgs, IR::OpSize ElementSize, bool IsStore);
void AVX128_MASKMOV(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VectorVariableBlend(OpcodeArgs);
void AVX128_VectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_SaveAVXState(Ref MemBase);
void AVX128_RestoreAVXState(Ref MemBase);
void AVX128_DefaultAVXState();
void AVX128_VPERM2(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VTESTP(OpcodeArgs);
void AVX128_VTESTP(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_PTest(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVX128_VPERMILReg(OpcodeArgs);
void AVX128_VPERMILReg(OpcodeArgs, IR::OpSize ElementSize);
void AVX128_VPERMD(OpcodeArgs);
@@ -1120,8 +1092,7 @@ public:
RefPair AVX128_VPGatherQPSImpl(Ref Dest, Ref Mask, RefVSIB VSIB);
RefPair AVX128_VPGatherImpl(OpSize Size, OpSize ElementLoadSize, OpSize AddrElementSize, RefPair Dest, RefPair Mask, RefVSIB VSIB);
template<OpSize AddrElementSize>
void AVX128_VPGATHER(OpcodeArgs);
void AVX128_VPGATHER(OpcodeArgs, OpSize AddrElementSize);
void AVX128_VCVTPH2PS(OpcodeArgs);
void AVX128_VCVTPS2PH(OpcodeArgs);
@@ -1155,9 +1126,14 @@ public:
StoreXMMRegister(XMM, Value);
}
void AVXVectorALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorUnaryOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
// End of AVX 256-bit implementation
void InvalidOp(OpcodeArgs);
void NoExecOp(OpcodeArgs);
void SetPackedRFLAG(bool Lower8, Ref Src);
Ref GetPackedRFLAG(uint32_t FlagsMask = ~0U);
@@ -1176,15 +1152,44 @@ public:
}
}
void FlushRegisterCache(bool SRAOnly = false) {
void StoreContextHelper(IR::OpSize Size, RegisterClassType Class, Ref Value, uint32_t Offset) {
// For i128Bit, we won't see a normal Constant to inline, but as a special
// case we can replace with a 2x64-bit store which can use inline zeroes.
if (Size == OpSize::i128Bit) {
auto Header = GetOpHeader(WrapNode(Value));
const auto MAX_STP_OFFSET = (252 * 4);
if (Offset <= MAX_STP_OFFSET && Header->Op == OP_LOADNAMEDVECTORCONSTANT) {
auto Const = Header->C<IR::IROp_LoadNamedVectorConstant>();
if (Const->Constant == IR::NamedVectorConstant::NAMED_VECTOR_ZERO) {
Ref Zero = _Constant(0);
Ref STP = _StoreContextPair(IR::OpSize::i64Bit, GPRClass, Zero, Zero, Offset);
// XXX: This works around InlineConstant not having an associated
// register class, else we'd just do InlineConstant above.
Ref InlineZero = _InlineConstant(0);
ReplaceNodeArgument(STP, 0, InlineZero);
ReplaceNodeArgument(STP, 1, InlineZero);
return;
}
}
}
_StoreContext(Size, Class, Value, Offset);
}
void FlushRegisterCache(bool SRAOnly = false, bool MMXOnly = false) {
// At block boundaries, fix up the carry flag.
if (!SRAOnly) {
RectifyCarryInvert(CFInvertedABI);
}
CalculateDeferredFlags();
if (!MMXOnly) {
CalculateDeferredFlags();
}
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
const auto VectorSize = GetGuestVectorLength();
// Write backwards. This is a heuristic to improve coalescing, since we
@@ -1204,6 +1209,11 @@ public:
Bits &= Mask;
}
if (MMXOnly) {
Mask &= ((1ull << (MM7Index - MM0Index + 1)) - 1) << MM0Index;
Bits &= Mask;
}
while (Bits != 0) {
uint32_t Index = 63 - std::countl_zero(Bits);
Ref Value = RegCache.Value[Index];
@@ -1241,10 +1251,10 @@ public:
_StoreContextPair(Size, Class, ValueNext, Value, Offset - SizeInt);
Bits &= ~NextBit;
} else {
_StoreContext(Size, Class, Value, Offset);
StoreContextHelper(Size, Class, Value, Offset);
// If Partial and MMX register, then we need to store all 1s in bits 64-80
if (Partial && Index >= MM0Index && Index <= MM7Index) {
_StoreContext(OpSize::i16Bit, IR::GPRClass, _Constant(0xFFFF), Offset + 8);
_StoreContext(OpSize::i16Bit, IR::GPRClass, Constant(0xFFFF), Offset + 8);
}
}
}
@@ -1257,6 +1267,10 @@ public:
RegCache.Partial &= ~Mask;
}
IR::OpSize GetGPROpSize() const {
return Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
protected:
void RecordX87Use() override {
CurrentHeader->HasX87 = true;
@@ -1310,6 +1324,7 @@ private:
struct JumpTargetInfo {
Ref BlockEntry;
bool HaveEmitted;
bool IsEntryPoint;
};
FEXCore::Context::ContextImpl* CTX {};
@@ -1358,11 +1373,6 @@ private:
Ref ADDSUBPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2);
void AVXVectorALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorUnaryOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
void AVXVariableShiftImpl(OpcodeArgs, IROps IROp);
Ref AESKeyGenAssistImpl(OpcodeArgs);
@@ -1495,7 +1505,7 @@ private:
#undef OpcodeArgs
Ref AppendSegmentOffset(Ref Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
Ref GetSegment(uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
Ref GetSegment(uint32_t Flags, uint32_t DefaultPrefix = FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX, bool Override = false);
void UpdatePrefixFromSegment(Ref Segment, uint32_t SegmentReg);
@@ -1503,14 +1513,32 @@ private:
void StoreGPRRegister(uint32_t GPR, const Ref Src, IR::OpSize Size = OpSize::iInvalid, uint8_t Offset = 0);
void StoreXMMRegister(uint32_t XMM, const Ref Src);
Ref GetRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0);
Ref _GetRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset, bool Inline) {
const auto GPRSize = GetGPROpSize();
const auto Offs = Op->PC + Op->InstSize + Offset - Entry;
return Inline ? _InlineEntrypointOffset(GPRSize, Offs) : _EntrypointOffset(GPRSize, Offs);
}
bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
Ref GetRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0) {
return _GetRelocatedPC(Op, Offset, false);
}
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0) {
ExitFunction(_GetRelocatedPC(Op, Offset, true /* Inline */));
}
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
ExitFunction(_GetRelocatedPC(Op, Offset, true /* Inline */), Hint, CallReturnAddress, CallReturnBlock);
}
[[nodiscard]]
static bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
// Literals are immediates as sources but memory addresses as destinations.
return !(Load && Operand.IsLiteral()) && !Operand.IsGPR();
}
bool IsNonTSOReg(MemoryAccessType Access, uint8_t Reg) {
[[nodiscard]]
static bool IsNonTSOReg(MemoryAccessType Access, uint8_t Reg) {
return Access == MemoryAccessType::DEFAULT && Reg == X86State::REG_RSP;
}
@@ -1617,7 +1645,7 @@ private:
}
void ZeroNZCV() {
CachedNZCV = _Constant(0);
CachedNZCV = Constant(0);
NZCVDirty = true;
}
@@ -1630,9 +1658,9 @@ private:
// This is currently worse for 8/16-bit, but that should be optimized. TODO
if (SrcSize >= OpSize::i32Bit) {
if (SetPF) {
CalculatePF(_SubWithFlags(SrcSize, Res, _Constant(0)));
CalculatePF(SubWithFlags(SrcSize, Res, (uint64_t)0));
} else {
_SubNZCV(SrcSize, Res, _Constant(0));
_SubNZCV(SrcSize, Res, Constant(0));
}
CFInverted = true;
@@ -1703,7 +1731,7 @@ private:
} else {
// Invert as a GPR
unsigned Bit = IndexNZCV(FEXCore::X86State::RFLAG_CF_RAW_LOC);
SetNZCV(_Xor(OpSize::i32Bit, GetNZCV(), _Constant(1u << Bit)));
SetNZCV(_Xor(OpSize::i32Bit, GetNZCV(), Constant(1u << Bit)));
CalculateDeferredFlags();
}
@@ -1741,7 +1769,7 @@ private:
}
HandleNZCVWrite();
_SubNZCV(OpSize::i32Bit, _Constant(0), Value);
_SubNZCV(OpSize::i32Bit, Constant(0), Value);
CFInverted = true;
}
@@ -1766,25 +1794,25 @@ private:
StoreRegister(Core::CPUState::AF_AS_GREG, false, Value);
} else if (BitOffset == FEXCore::X86State::RFLAG_DF_RAW_LOC) {
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, _Constant(1), Value, ShiftType::LSL, 1));
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF = _Select(FEXCore::IR::COND_EQ, Value, _Constant(0), _And(OpSize::i32Bit, PackedTF, _Constant(~1)), _Constant(1));
auto NewPackedTF = _Select(FEXCore::IR::COND_EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContext(OpSize::i8Bit, GPRClass, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
} else {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
}
}
void SetAF(unsigned Constant) {
void SetAF(unsigned K) {
// AF is stored in bit 4 of the AF flag byte, with garbage in the other
// bits. This allows us to defer the extract in the usual case. When it is
// read, bit 4 is extracted. In order to write a constant value of AF, that
// means we need to left-shift here to compensate.
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Constant(Constant << 4));
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(K << 4));
}
void ZeroPF_AF();
@@ -1800,7 +1828,8 @@ private:
InvalidateReg(Core::CPUState::AF_AS_GREG);
}
CondClassType CondForNZCVBit(unsigned BitOffset, bool Invert) {
[[nodiscard]]
static CondClassType CondForNZCVBit(unsigned BitOffset, bool Invert) {
switch (BitOffset) {
case X86State::RFLAG_SF_RAW_LOC: return {Invert ? COND_PL : COND_MI};
case X86State::RFLAG_ZF_RAW_LOC: return {Invert ? COND_NEQ : COND_EQ};
@@ -1826,7 +1855,8 @@ private:
static const int AVXHigh0Index = 48;
static const int AVXHigh15Index = 63;
uint32_t CacheIndexToContextOffset(int Index) {
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
@@ -1835,7 +1865,8 @@ private:
}
}
RegisterClassType CacheIndexClass(int Index) {
[[nodiscard]]
static RegisterClassType CacheIndexClass(int Index) {
if ((Index >= MM0Index && Index <= MM7Index) || Index >= FPR0Index) {
return FPRClass;
} else {
@@ -1843,7 +1874,8 @@ private:
}
}
IR::OpSize CacheIndexToOpSize(int Index) {
[[nodiscard]]
static IR::OpSize CacheIndexToOpSize(int Index) {
// MMX registers are rounded up to 128-bit since they are shared with 80-bit
// x87 registers, even though MMX is logically only 64-bit.
if (Index >= AVXHigh0Index || ((Index >= MM0Index && Index <= MM7Index))) {
@@ -1951,7 +1983,7 @@ private:
}
Ref LoadGPR(uint8_t Reg) {
return LoadRegCache(Reg, GPR0Index + Reg, GPRClass, CTX->GetGPROpSize());
return LoadRegCache(Reg, GPR0Index + Reg, GPRClass, GetGPROpSize());
}
Ref LoadContext(IR::OpSize Size, uint8_t Index) {
@@ -2005,14 +2037,14 @@ private:
auto Value = _Bfe(OpSize::i32Bit, 1, IndexNZCV(BitOffset), GetNZCV());
if (Invert) {
return _Xor(OpSize::i32Bit, Value, _Constant(1));
return _Xor(OpSize::i32Bit, Value, Constant(1));
} else {
return Value;
}
} else {
// Because we explicitly inverted for CF above, we use the unsafe
// _NZCVSelect rather than the safe CF-aware version.
return _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(BitOffset, Invert), _Constant(1), _Constant(0));
return _NZCVSelect01(CondForNZCVBit(BitOffset, Invert));
}
} else if (BitOffset == FEXCore::X86State::RFLAG_PF_RAW_LOC) {
return LoadGPR(Core::CPUState::PF_AS_GREG);
@@ -2020,7 +2052,7 @@ private:
return LoadGPR(Core::CPUState::AF_AS_GREG);
} else if (BitOffset == FEXCore::X86State::RFLAG_DF_RAW_LOC) {
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), _Constant(63));
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(Core::CPUState, flags[BitOffset]));
}
@@ -2092,7 +2124,7 @@ private:
// Zero AF. Note that the comparison sets the raw PF to 0/1 above, so
// PF[4] is 0 so the XOR with PF will have no effect, so setting the AF
// byte to zero will indeed zero AF as intended.
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(0));
}
// Convert NZCV from the Arm representation to an eXternal representation
@@ -2109,7 +2141,7 @@ private:
void ConvertNZCVToX87() {
LOGMAN_THROW_A_FMT(NZCVDirty && CachedNZCV, "NZCV must be saved");
Ref V = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_OF_RAW_LOC, false), _Constant(1), _Constant(0));
Ref V = _NZCVSelect01(CondForNZCVBit(FEXCore::X86State::RFLAG_OF_RAW_LOC, false));
if (CTX->HostFeatures.SupportsFlagM2) {
// Convert to x86 flags, saves us from or'ing after.
@@ -2117,8 +2149,8 @@ private:
}
// CF is inverted after FCMP
Ref C = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_CF_RAW_LOC, true), _Constant(1), _Constant(0));
Ref Z = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_ZF_RAW_LOC, false), _Constant(1), _Constant(0));
Ref C = _NZCVSelect01(CondForNZCVBit(FEXCore::X86State::RFLAG_CF_RAW_LOC, true));
Ref Z = _NZCVSelect01(CondForNZCVBit(FEXCore::X86State::RFLAG_ZF_RAW_LOC, false));
if (!CTX->HostFeatures.SupportsFlagM2) {
C = _Or(OpSize::i32Bit, C, V);
@@ -2126,7 +2158,7 @@ private:
}
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(C);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(V);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(Z);
}
@@ -2176,9 +2208,9 @@ private:
return CachedNamedVectorConstants[NamedConstant][log2_size_bytes];
}
auto Constant = _LoadNamedVectorConstant(Size, NamedConstant);
CachedNamedVectorConstants[NamedConstant][log2_size_bytes] = Constant;
return Constant;
auto K = _LoadNamedVectorConstant(Size, NamedConstant);
CachedNamedVectorConstants[NamedConstant][log2_size_bytes] = K;
return K;
}
Ref LoadAndCacheIndexedNamedVectorConstant(IR::OpSize Size, FEXCore::IR::IndexNamedVectorConstant NamedIndexedConstant, uint32_t Index) {
IndexNamedVectorMapKey Key {
@@ -2192,9 +2224,9 @@ private:
return it->second;
}
auto Constant = _LoadNamedVectorIndexedConstant(Size, NamedIndexedConstant, Index);
CachedIndexedNamedVectorConstants.insert_or_assign(Key, Constant);
return Constant;
auto K = _LoadNamedVectorIndexedConstant(Size, NamedIndexedConstant, Index);
CachedIndexedNamedVectorConstants.insert_or_assign(Key, K);
return K;
}
Ref LoadUncachedZeroVector(IR::OpSize Size) {
@@ -2212,9 +2244,9 @@ private:
CachedIndexedNamedVectorConstants.clear();
}
std::pair<bool, CondClassType> DecodeNZCVCondition(uint8_t OP);
std::optional<CondClassType> DecodeNZCVCondition(uint8_t OP);
Ref SelectBit(Ref Cmp, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
Ref SelectCC(uint8_t OP, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
Ref SelectCC0All1(uint8_t OP);
/**
* @brief Flushes NZCV. Mostly vestigial.
@@ -2249,7 +2281,7 @@ private:
}
// Otherwise, prepare to branch.
auto Zero = _Constant(0);
auto Zero = Constant(0);
// If the shift is zero, do not touch the flags.
auto SetBlock = CreateNewCodeBlockAfter(GetCurrentBlock());
@@ -2289,7 +2321,6 @@ private:
* @name These functions are used by the deferred flag handling while it is calculating and storing flags in to RFLAGs.
* @{ */
Ref LoadPFRaw(bool Mask, bool Invert);
Ref SelectPF(bool Invert, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
Ref LoadAF();
void FixupAF();
void SetAFAndFixup(Ref AF);
@@ -2325,17 +2356,19 @@ private:
void ChgStateX87_MMX() override {
LOGMAN_THROW_A_FMT(MMXState == MMXState_X87, "Expected state to be x87");
_StackForceSlow();
SetX87Top(_Constant(0)); // top reset to zero
StoreContext(AbridgedFTWIndex, _Constant(0xFFFFUL)); // all valid
SetX87Top(Constant(0)); // top reset to zero
StoreContext(AbridgedFTWIndex, Constant(0xFFFFUL)); // all valid
MMXState = MMXState_MMX;
}
void ChgStateMMX_X87() override {
LOGMAN_THROW_A_FMT(MMXState == MMXState_MMX, "Expected state to be MMX");
// The opcode dispatcher register cache is used for MMX, but the x87 pass register cache is used for x87, spill to
// context to ensure coherence.
FlushRegisterCache(false, true);
// We explicitly initialize to x87 state in StartNewBlock.
// So if we ever change this to do something else, we need to
// make sure that we consider if we need to explicitly set it there.
FlushRegisterCache();
MMXState = MMXState_X87;
}
@@ -2351,10 +2384,17 @@ private:
bool BlockSetRIP {false};
bool Multiblock {};
bool Is64BitMode {};
uint64_t Entry {};
// Set if mono hacks are enabled and the current block is the mono callsite backpatcher, in which case the
// XCHG ops that would patch code are replaced with a hook that performs the write and manually invalidates
// the target address.
bool IsMonoBackpatcherBlock {false};
IROp_IRHeader* CurrentHeader {};
bool IsTSOEnabled(FEXCore::IR::RegisterClassType Class) {
[[nodiscard]]
bool IsTSOEnabled(FEXCore::IR::RegisterClassType Class) const {
if (ForceTSO == ForceTSOMode::ForceEnabled) {
return true;
} else if (ForceTSO == ForceTSOMode::ForceDisabled) {
@@ -2384,7 +2424,7 @@ private:
Ref _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
if (AtomicTSO) {
return _LoadMemTSO(Class, Size, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
@@ -2404,7 +2444,7 @@ private:
A.Offset = 0;
}
Out.Base = LoadEffectiveAddress(this, A, CTX->GetGPROpSize(), true, false);
Out.Base = LoadEffectiveAddress(this, A, GetGPROpSize(), true, false);
return Out;
}
@@ -2435,7 +2475,7 @@ private:
Ref _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, Ref Value, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
if (AtomicTSO) {
return _StoreMemTSO(Class, Size, Value, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
@@ -2486,7 +2526,7 @@ private:
void Push(IR::OpSize Size, Ref Value) {
auto OldSP = LoadGPRRegister(X86State::REG_RSP);
auto NewSP = _Push(CTX->GetGPROpSize(), Size, Value, OldSP);
auto NewSP = _Push(GetGPROpSize(), Size, Value, OldSP);
StoreGPRRegister(X86State::REG_RSP, NewSP);
FlushRegisterCache();
}
@@ -2516,11 +2556,11 @@ private:
}
ArithRef And(uint64_t K) {
return IsConstant ? ArithRef(E, C & K) : ArithRef(E, E->_And(OpSize::i64Bit, R, E->_Constant(K)));
return IsConstant ? ArithRef(E, C & K) : ArithRef(E, E->_And(OpSize::i64Bit, R, E->Constant(K)));
}
ArithRef Presub(uint64_t K) {
return IsConstant ? ArithRef(E, K - C) : ArithRef(E, E->_Sub(OpSize::i64Bit, E->_Constant(K), R));
return IsConstant ? ArithRef(E, K - C) : ArithRef(E, E->Sub(OpSize::i64Bit, E->Constant(K), R));
}
ArithRef Lshl(uint64_t Shift) {
@@ -2529,7 +2569,7 @@ private:
} else if (IsConstant) {
return ArithRef(E, C << Shift);
} else {
return ArithRef(E, E->_Lshl(OpSize::i64Bit, R, E->_Constant(Shift)));
return ArithRef(E, E->_Lshl(OpSize::i64Bit, R, E->Constant(Shift)));
}
}
@@ -2569,7 +2609,7 @@ private:
}
if (IsConstant) {
return E->_Bfi(OpSize::i64Bit, Size, Start, Bitfield, E->_Constant(C));
return E->_Bfi(OpSize::i64Bit, Size, Start, Bitfield, E->Constant(C));
} else {
return E->_Bfi(OpSize::i64Bit, Size, Start, Bitfield, R);
}
@@ -2585,15 +2625,15 @@ private:
return ArithRef(E, Result);
} else {
return ArithRef(E, E->_Lshl(Size, E->_Constant(1), R));
return ArithRef(E, E->_Lshl(Size, E->Constant(1), R));
}
}
Ref Ref() {
return IsConstant ? E->_Constant(C) : R;
return IsConstant ? E->Constant(C) : R;
}
bool IsDefinitelyZero() {
bool IsDefinitelyZero() const {
return IsConstant && C == 0;
}
};
@@ -2612,8 +2652,6 @@ private:
return ArithRef(this, K);
}
void InstallHostSpecificOpcodeHandlers();
///< Segment telemetry tracking
uint32_t SegmentsNeedReadCheck {~0U};
void CheckLegacySegmentWrite(Ref NewNode, uint32_t SegmentReg);
@@ -2622,21 +2660,19 @@ private:
constexpr inline void InstallToTable(auto& FinalTable, const auto& LocalTable) {
for (const auto& Op : LocalTable) {
auto OpNum = std::get<0>(Op);
auto Dispatcher = std::get<2>(Op);
for (uint8_t i = 0; i < std::get<1>(Op); ++i) {
auto OpNum = Op.Op;
auto Dispatcher = Op.Ptr;
for (uint8_t i = 0; i < Op.Count; ++i) {
auto& TableOp = FinalTable[OpNum + i];
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
if (TableOp.OpcodeDispatcher) {
if (TableOp.OpcodeDispatcher.OpDispatch) {
ERROR_AND_DIE_FMT("Duplicate Entry {}", TableOp.Name);
}
#endif
TableOp.OpcodeDispatcher = Dispatcher;
TableOp.OpcodeDispatcher.OpDispatch = Dispatcher;
}
}
}
void InstallOpcodeHandlers(Context::OperatingMode Mode);
} // namespace FEXCore::IR
File diff suppressed because it is too large. Load diff
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable[] = {
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable[] = {
// Instructions
{0x00, 6, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ALUOp, FEXCore::IR::IROps::OP_ADD, FEXCore::IR::IROps::OP_ATOMICFETCHADD, 0>},
@@ -57,6 +57,7 @@ constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispat
{0xC2, 2, &OpDispatchBuilder::RETOp},
{0xC8, 1, &OpDispatchBuilder::EnterOp},
{0xC9, 1, &OpDispatchBuilder::LEAVEOp},
{0xCA, 2, &OpDispatchBuilder::RETFARIndirectOp},
{0xCC, 2, &OpDispatchBuilder::INTOp},
{0xCF, 1, &OpDispatchBuilder::IRETOp},
{0xD7, 2, &OpDispatchBuilder::XLATOp},
@@ -75,33 +76,4 @@ constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispat
{0xFA, 2, &OpDispatchBuilder::PermissionRestrictedOp},
{0xFC, 2, &OpDispatchBuilder::FLAGControlOp},
};
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable_64[] = {
{0x63, 1, &OpDispatchBuilder::MOVSXDOp},
{0xA0, 4, &OpDispatchBuilder::MOVOffsetOp},
};
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable_32[] = {
{0x06, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x07, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x0E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX>},
{0x16, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX>},
{0x17, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX>},
{0x1E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX>},
{0x1F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX>},
{0x27, 1, &OpDispatchBuilder::DAAOp},
{0x2F, 1, &OpDispatchBuilder::DASOp},
{0x37, 1, &OpDispatchBuilder::AAAOp},
{0x3F, 1, &OpDispatchBuilder::AASOp},
{0x40, 8, &OpDispatchBuilder::INCOp},
{0x48, 8, &OpDispatchBuilder::DECOp},
{0x60, 1, &OpDispatchBuilder::PUSHAOp},
{0x61, 1, &OpDispatchBuilder::POPAOp},
{0xA0, 4, &OpDispatchBuilder::MOVOffsetOp},
{0xCE, 1, &OpDispatchBuilder::INTOp},
{0xD4, 1, &OpDispatchBuilder::AAMOp},
{0xD5, 1, &OpDispatchBuilder::AADOp},
{0xD6, 1, &OpDispatchBuilder::SALCOp},
};
} // namespace FEXCore::IR
@@ -11,10 +11,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
#include <array>
#include <cstdint>
#include <tuple>
#include <utility>
namespace FEXCore::IR {
class OrderedNode;
@@ -22,24 +19,20 @@ class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref RotatedNode {};
if (CTX->HostFeatures.SupportsSHA) {
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
} else {
// SHA1 extension missing, manually rotate.
// Emulate rotate.
auto ShiftLeft = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Dest, 30);
RotatedNode = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeft, Dest, 2);
}
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
@@ -47,6 +40,10 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
@@ -59,334 +56,145 @@ void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
auto Src2 = SHADataShuffle(Src);
// The result is swizzled differently than expected
Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
} else {
// Shift the incoming source left by a 32-bit element, inserting Zeros.
// This could be slightly improved to use a VInsGPR with the zero register.
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
auto Src2Shift = _VExtr(OpSize::i128Bit, OpSize::i8Bit, Src, ZeroRegister, 12);
auto Xor1 = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, Src2Shift);
// Emulate rotate.
auto ShiftLeftXor1 = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Xor1, 1);
auto RotatedXor1 = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXor1, Xor1, 31);
// Element0 didn't get XOR'd with anything, so do it now.
auto ExtractUpper = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, RotatedXor1, 3);
auto XorLower = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, ExtractUpper);
// Emulate rotate.
auto ShiftLeftXorLower = _VShlI(OpSize::i128Bit, OpSize::i32Bit, XorLower, 1);
auto RotatedXorLower = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXorLower, XorLower, 31);
Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 0, 0, RotatedXor1, RotatedXorLower);
}
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
auto Src2 = SHADataShuffle(Src);
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
using FnType = Ref (*)(OpDispatchBuilder&, Ref, Ref, Ref);
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1c?
return Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._Andn(OpSize::i32Bit, D, B));
};
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p with different key
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1m
return Self.BitwiseAtLeastTwo(B, C, D);
};
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
constexpr std::array<uint32_t, 4> k_array {
0x5A827999U,
0x6ED9EBA1U,
0x8F1BBCDCU,
0xCA62C1D6U,
};
constexpr std::array<FnType, 4> fn_array {
f0,
f1,
f2,
f3,
};
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Ref ConstantVector {};
switch (Imm8) {
case 0:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0);
break;
case 1:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1);
break;
case 2:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2);
break;
case 3:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3);
break;
}
Ref ConstantVector {};
switch (Imm8) {
case 0:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0);
break;
case 1:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1);
break;
case 2:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2);
break;
case 3:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3);
break;
}
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
Src2 = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src2, ConstantVector);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
Src2 = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src2, ConstantVector);
switch (Imm8) {
case 0: Result = SHADataShuffle(_VSha1C(Src1, ZeroRegister, Src2)); break;
case 2: Result = SHADataShuffle(_VSha1M(Src1, ZeroRegister, Src2)); break;
case 1:
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
} else {
const FnType Fn = fn_array[Imm8];
auto K = _Constant(OpSize::i32Bit, k_array[Imm8]);
auto W0E = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto B = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto C = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto D = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
auto A1 =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto D1 = C;
auto E1 = D;
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
auto ANext =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto DNext = C;
auto ENext = D;
return {ANext, BNext, CNext, DNext, ENext};
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Dest3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Dest2, std::get<2>(Final));
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Dest1, std::get<3>(Final));
switch (Imm8) {
case 0: Result = SHADataShuffle(_VSha1C(Src1, ZeroRegister, Src2)); break;
case 2: Result = SHADataShuffle(_VSha1M(Src1, ZeroRegister, Src2)); break;
case 1:
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Result = _VSha256U0(Dest, Src);
} else {
const auto Sigma0 = [this](Ref W) -> Ref {
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 7)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 18))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 3)));
};
auto W4 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto W3 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto W2 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto W1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto W0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
auto Sig3 = _Add(OpSize::i32Bit, W3, Sigma0(W4));
auto Sig2 = _Add(OpSize::i32Bit, W2, Sigma0(W3));
auto Sig1 = _Add(OpSize::i32Bit, W1, Sigma0(W2));
auto Sig0 = _Add(OpSize::i32Bit, W0, Sigma0(W1));
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, Sig3);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, Sig2);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, Sig1);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, Sig0);
}
auto Result = _VSha256U0(Dest, Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](Ref W) -> Ref {
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 17)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 19))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 10)));
};
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Src2 = _VZip2(OpSize::i128Bit, OpSize::i64Bit, DupDst, Src);
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Src2 = _VZip2(OpSize::i128Bit, OpSize::i64Bit, DupDst, Src);
Result = _VSha256U1(Src1, Src2);
} else {
auto W14 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto W15 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto W16 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0), Sigma1(W14));
auto W17 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1), Sigma1(W15));
auto W18 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2), Sigma1(W16));
auto W19 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3), Sigma1(W17));
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, W19);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, W18);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, W17);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, W16);
}
auto Result = _VSha256U1(Src1, Src2);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
Ref OpDispatchBuilder::BitwiseAtLeastTwo(Ref A, Ref B, Ref C) {
// Returns whether at least 2/3 of A/B/C is true.
// Expressed as (A & (B | C)) | (B & C)
//
// Equivalent to expression in SHA calculations: (A & B) ^ (A & C) ^ (B & C)
auto And = _And(OpSize::i32Bit, B, C);
auto Or = _Or(OpSize::i32Bit, B, C);
return _Or(OpSize::i32Bit, _And(OpSize::i32Bit, A, Or), And);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
auto shuffle_abcd = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `abcd` configuration from x86 format.
auto Tmp = _VZip2(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_abcd = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `abcd` configuration from x86 format.
auto Tmp = _VZip2(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_efgh = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `efgh` configuration from x86 format.
auto Tmp = _VZip(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_efgh = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `efgh` configuration from x86 format.
auto Tmp = _VZip(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto ABCD = shuffle_abcd(Dest, Src);
auto EFGH = shuffle_efgh(Dest, Src);
auto ABCD = shuffle_abcd(Dest, Src);
auto EFGH = shuffle_efgh(Dest, Src);
// x86 uses only the bottom 64-bits of the key, so duplicate to match ARM64 semantics.
auto Key = _VDupElement(OpSize::i128Bit, OpSize::i64Bit, XMM0, 0);
// x86 uses only the bottom 64-bits of the key, so duplicate to match ARM64 semantics.
auto Key = _VDupElement(OpSize::i128Bit, OpSize::i64Bit, XMM0, 0);
auto A = _VSha256H(ABCD, EFGH, Key);
auto B = _VSha256H2(EFGH, ABCD, Key);
Result = shuffle_abcd(A, B);
} else {
const auto Ch = [this](Ref E, Ref F, Ref G) -> Ref {
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
};
const auto Sigma0 = [this](Ref A) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 2)), A, ShiftType::ROR, 13),
A, ShiftType::ROR, 22);
};
const auto Sigma1 = [this](Ref E) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(OpSize::i32Bit, 6)), E, ShiftType::ROR, 11),
E, ShiftType::ROR, 25);
};
auto E0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 1);
auto F0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Ref Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
auto WK0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 0);
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
auto H0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
Q0 = _Add(OpSize::i32Bit, Q0, H0);
auto A0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto B0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto A1 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q0, BitwiseAtLeastTwo(A0, B0, C0)), Sigma0(A0));
auto D0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
Ref Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
auto WK1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 1);
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
// Rematerialize G0. Costs a move but saves spilling, coming out ahead.
G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Q1 = _Add(OpSize::i32Bit, Q1, G0);
auto A2 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q1, BitwiseAtLeastTwo(A1, A0, B0)), Sigma0(A1));
// Rematerialize C0. As with G0.
C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto E2 = _Add(OpSize::i32Bit, Q1, C0);
auto Res3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, A2);
auto Res2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Res3, A1);
auto Res1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Res2, E2);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Res1, E1);
}
auto A = _VSha256H(ABCD, EFGH, Key);
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
@@ -395,7 +203,7 @@ void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
@@ -408,6 +216,10 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
@@ -416,7 +228,7 @@ void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
@@ -429,6 +241,10 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
@@ -437,7 +253,7 @@ void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
@@ -450,6 +266,10 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
@@ -458,7 +278,7 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
@@ -479,11 +299,20 @@ Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsPMULL_128Bit) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
@@ -493,6 +322,10 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsPMULL_128Bit) {
UnimplementedOp(Op);
return;
}
const auto DstSize = OpSizeFromDst(Op);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_DDDTable[] = {
constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
@@ -28,7 +28,7 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
void OpDispatchBuilder::ZeroPF_AF() {
// PF is stored inverted, so invert it when we zero.
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(_Constant(1));
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(Constant(1));
SetAF(0);
}
@@ -201,7 +201,7 @@ void OpDispatchBuilder::FixupAF() {
auto PFRaw = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
auto AFRaw = GetRFLAG(FEXCore::X86State::RFLAG_AF_RAW_LOC);
// Again 64-bit as masking is more expensive given our ConstProp design.
// Again 64-bit as masking is more expensive.
Ref XorRes = _Xor(OpSize::i64Bit, AFRaw, PFRaw);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -247,7 +247,7 @@ void OpDispatchBuilder::CalculateAF(Ref Src1, Ref Src2) {
// We store the XOR of the arguments. At read time, we XOR with the
// appropriate bit of the result (available as the PF flag) and extract the
// appropriate bit. Again 64-bit to avoid masking.
Ref XorRes = Src1 == Src2 ? _Constant(0) : _Xor(OpSize::i64Bit, Src1, Src2);
Ref XorRes = Src1 == Src2 ? Constant(0) : _Xor(OpSize::i64Bit, Src1, Src2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -267,8 +267,6 @@ Ref OpDispatchBuilder::IncrementByCarry(OpSize OpSize, Ref Src) {
}
Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
Ref Res;
@@ -288,11 +286,11 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Src2PlusCF = IncrementByCarry(OpSize, Src2);
// Need to zero-extend for the comparison.
Res = _Add(OpSize, Src1, Src2PlusCF);
Res = Add(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
// TODO: We can fold that second Bfe in (cmp uxth).
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Res, Src2PlusCF, One, Zero);
auto SelectCFInv = Select01(OpSize, CondClassType {COND_UGE}, Res, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -304,8 +302,6 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
}
Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(Src1, Src2);
@@ -325,10 +321,10 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2
auto Src2PlusCF = IncrementByCarry(OpSize, Src2);
Res = _Sub(OpSize, Src1, Src2PlusCF);
Res = Sub(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Src1, Src2PlusCF, One, Zero);
auto SelectCFInv = Select01(OpSize, CondClassType {COND_UGE}, Src1, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -349,10 +345,10 @@ Ref OpDispatchBuilder::CalculateFlags_SUB(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Res;
if (SrcSize >= OpSize::i32Bit) {
Res = _SubWithFlags(SrcSize, Src1, Src2);
Res = SubWithFlags(SrcSize, Src1, Src2);
} else {
_SubNZCV(SrcSize, Src1, Src2);
Res = _Sub(OpSize::i32Bit, Src1, Src2);
Res = Sub(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
@@ -379,10 +375,10 @@ Ref OpDispatchBuilder::CalculateFlags_ADD(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Res;
if (SrcSize >= OpSize::i32Bit) {
Res = _AddWithFlags(SrcSize, Src1, Src2);
Res = AddWithFlags(SrcSize, Src1, Src2);
} else {
_AddNZCV(SrcSize, Src1, Src2);
Res = _Add(OpSize::i32Bit, Src1, Src2);
Res = Add(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
@@ -6,9 +6,10 @@ namespace FEXCore::IR {
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F38Table[] = {
constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_66, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_NONE, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i16Bit>},
@@ -71,9 +72,28 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(PF_38_66, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(PF_38_NONE, 0xC8), 1, &OpDispatchBuilder::SHA1NEXTEOp},
{OPD(PF_38_NONE, 0xC9), 1, &OpDispatchBuilder::SHA1MSG1Op},
{OPD(PF_38_NONE, 0xCA), 1, &OpDispatchBuilder::SHA1MSG2Op},
{OPD(PF_38_NONE, 0xCB), 1, &OpDispatchBuilder::SHA256RNDS2Op},
{OPD(PF_38_NONE, 0xCC), 1, &OpDispatchBuilder::SHA256MSG1Op},
{OPD(PF_38_NONE, 0xCD), 1, &OpDispatchBuilder::SHA256MSG2Op},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_66, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66, 0xF6), 1, &OpDispatchBuilder::ADXOp},
{OPD(PF_38_F3, 0xF6), 1, &OpDispatchBuilder::ADXOp},
};
@@ -8,7 +8,7 @@ namespace FEXCore::IR {
#define PF_3A_66 1
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> Table[] = {
constexpr DispatchTableEntry Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
@@ -29,6 +29,7 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
@@ -36,14 +37,16 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
return std::to_array(Table);
};
auto REX0 = OpDispatchTableGenH0F3AREX.template operator()<0>();
auto REX1 = OpDispatchTableGenH0F3AREX.template operator()<1>();
auto concat = []<typename T, size_t N1, size_t N2>(std::array<T, N1> const& lhs,
std::array<T, N2> const& rhs) consteval -> std::array<T, N1 + N2> {
auto concat = []<typename T, size_t N1, size_t N2>(const std::array<T, N1>& lhs,
const std::array<T, N2>& rhs) consteval -> std::array<T, N1 + N2> {
std::array<T, N1 + N2> Table {};
for (size_t i = 0; i < N1; ++i) {
Table[i] = lhs[i];
@@ -60,16 +63,11 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATableNeedsREX0[] = {
constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i64Bit>},
{OPD(1, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i64Bit>},
};
#undef PF_3A_NONE
#undef PF_3A_66
@@ -5,7 +5,7 @@
namespace FEXCore::IR {
using X86Tables::OpToIndex;
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_PrimaryGroupTables[] = {
constexpr DispatchTableEntry OpDispatch_PrimaryGroupTables[] = {
// GROUP 1
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -117,7 +117,9 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, &OpDispatchBuilder::INCOp}, // INC
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, &OpDispatchBuilder::DECOp}, // DEC
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, &OpDispatchBuilder::CALLAbsoluteOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, &OpDispatchBuilder::CALLFARIndirectOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, &OpDispatchBuilder::JUMPAbsoluteOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 5), 1, &OpDispatchBuilder::JUMPFARIndirectOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 6), 1, &OpDispatchBuilder::PUSHOp},
// GROUP 11
@@ -8,7 +8,7 @@ constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryGroupTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 6
{OPD(FEXCore::X86Tables::TYPE_GROUP_6, PF_NONE, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_6, PF_F3, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
@@ -69,10 +69,16 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 7), 1, &OpDispatchBuilder::RDPIDOp},
// GROUP 12
@@ -113,11 +119,13 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_NONE, 7), 1, &OpDispatchBuilder::StoreFenceOrCLFlush}, // SFENCE (or CLFLUSH)
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 5), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 6), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 6), 1, &OpDispatchBuilder::UMonitorOrCLRSSBSY},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 6), 1, &OpDispatchBuilder::CLWB},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 6), 1, &OpDispatchBuilder::CLWBOrTPause},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 7), 1, &OpDispatchBuilder::CLFLUSHOPT},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F2, 6), 1, &OpDispatchBuilder::UMWaitOp},
// GROUP 16
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_NONE, 0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, true, 1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_NONE, 1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 1>},
@@ -154,18 +162,6 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_F2, 0), 8, &OpDispatchBuilder::NOPOp},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryGroupTables_64[] = {
// GROUP 15
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 0), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::ReadSegmentReg, OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 1), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::ReadSegmentReg, OpDispatchBuilder::Segment::GS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 2), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::WriteSegmentReg, OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 3), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::WriteSegmentReg, OpDispatchBuilder::Segment::GS>},
};
#undef OPD
} // namespace FEXCore::IR
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryModRMTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryModRMTables[] = {
// REG /1
{((0 << 3) | 0), 1, &OpDispatchBuilder::UnimplementedOp},
{((0 << 3) | 1), 1, &OpDispatchBuilder::UnimplementedOp},
@@ -17,6 +17,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
// REG /7
{((3 << 3) | 0), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
} // namespace FEXCore::IR
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable[] = {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
@@ -150,7 +150,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
#endif
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryRepModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
@@ -181,7 +181,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryRepNEModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
@@ -207,7 +207,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryOpSizeModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
@@ -313,20 +313,4 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i16Bit>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable_64[] = {
{0x05, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SyscallOp, true>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
{0xA9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable_32[] = {
{0x05, 1, &OpDispatchBuilder::NOPOp},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
{0xA9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
};
} // namespace FEXCore::IR
@@ -4,7 +4,7 @@
namespace FEXCore::IR {
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_VEXTable[] = {
constexpr DispatchTableEntry OpDispatch_VEXTable[] = {
{OPD(2, 0b00, 0xF2), 1, &OpDispatchBuilder::ANDNBMIOp}, {OPD(2, 0b00, 0xF5), 1, &OpDispatchBuilder::BZHI},
{OPD(2, 0b10, 0xF5), 1, &OpDispatchBuilder::PEXT}, {OPD(2, 0b11, 0xF5), 1, &OpDispatchBuilder::PDEP},
{OPD(2, 0b11, 0xF6), 1, &OpDispatchBuilder::MULX}, {OPD(2, 0b00, 0xF7), 1, &OpDispatchBuilder::BEXTRBMIOp},
@@ -16,7 +16,7 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
#undef OPD
#define OPD(group, pp, opcode) (((group - X86Tables::InstType::TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_VEXGroupTable[] = {
constexpr DispatchTableEntry OpDispatch_VEXGroupTable[] = {
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b001), 1, &OpDispatchBuilder::BLSRBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b010), 1, &OpDispatchBuilder::BLSMSKBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b011), 1, &OpDispatchBuilder::BLSIBMIOp},
@@ -434,7 +434,7 @@ Ref OpDispatchBuilder::InsertCVTGPR_To_FPRImpl(OpcodeArgs, IR::OpSize DstSize, I
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
return _VSToFGPRInsert(DstSize, DstElementSize, SrcSize, Src1, Src2, ZeroUpperBits);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
@@ -477,7 +477,7 @@ Ref OpDispatchBuilder::InsertScalar_CVT_Float_To_FloatImpl(OpcodeArgs, IR::OpSiz
// We load the full vector width when dealing with a source vector,
// so that we don't do any unnecessary zero extension to the scalar
// element that we're going to operate on.
const auto SrcSize = OpSizeFromSrc(Op);
const auto SrcSize = Src2Op.IsGPR() ? OpSize::i128Bit : SrcElementSize;
Ref Src1 = LoadSource_WithOpSize(FPRClass, Op, Src1Op, DstSize, Op->Flags);
Ref Src2 = LoadSource_WithOpSize(FPRClass, Op, Src2Op, SrcSize, Op->Flags, {.AllowUpperGarbage = true});
@@ -740,8 +740,8 @@ void OpDispatchBuilder::MOVMSKOp(OpcodeArgs, IR::OpSize ElementSize) {
// Inserting the full lower 32-bits offset 31 so the sign bit ends up at offset 63.
GPR = _Bfi(OpSize::i64Bit, 32, 31, GPR, GPR);
// Shift right to only get the two sign bits we care about.
GPR = _Lshr(OpSize::i64Bit, GPR, _Constant(62));
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, CTX->GetGPROpSize(), OpSize::iInvalid);
GPR = _Lshr(OpSize::i64Bit, GPR, Constant(62));
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
} else if (Size == OpSize::i128Bit && ElementSize == OpSize::i32Bit) {
// Shift all the sign bits to the bottom of their respective elements.
Src = _VUShrI(Size, OpSize::i32Bit, Src, 31);
@@ -753,9 +753,9 @@ void OpDispatchBuilder::MOVMSKOp(OpcodeArgs, IR::OpSize ElementSize) {
Src = _VAddV(Size, OpSize::i32Bit, Src);
// Extract to a GPR.
Ref GPR = _VExtractToGPR(Size, OpSize::i32Bit, Src, 0);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, CTX->GetGPROpSize(), OpSize::iInvalid);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
} else {
Ref CurrentVal = _Constant(0);
Ref CurrentVal = Constant(0);
for (unsigned i = 0; i < NumElements; ++i) {
// Extract the top bit of the element
@@ -1504,7 +1504,7 @@ void OpDispatchBuilder::VBROADCASTOp(OpcodeArgs, IR::OpSize ElementSize) {
Result = _VDupElement(DstSize, ElementSize, Src, 0);
} else {
// Get the address to broadcast from into a GPR.
Ref Address = MakeSegmentAddress(Op, Op->Src[0], CTX->GetGPROpSize());
Ref Address = MakeSegmentAddress(Op, Op->Src[0], GetGPROpSize());
Result = _VBroadcastFromMem(DstSize, ElementSize, Address);
}
@@ -1523,7 +1523,7 @@ Ref OpDispatchBuilder::PINSROpImpl(OpcodeArgs, IR::OpSize ElementSize, const X86
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
return _VInsGPR(Size, ElementSize, Index, Src1, Src2);
}
@@ -1644,7 +1644,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs, IR::OpSize ElementSize) {
Index &= NumElements - 1;
if (Op->Dest.IsGPR()) {
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
// Extract already zero extends the result.
Ref Result = _VExtractToGPR(OpSize::i128Bit, OverridenElementSize, Src, Index);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, GPRSize, OpSize::iInvalid);
@@ -2065,7 +2065,7 @@ Ref OpDispatchBuilder::CVTGPR_To_FPRImpl(OpcodeArgs, IR::OpSize DstElementSize,
Ref Converted {};
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
Converted = _Float_FromGPR_S(DstElementSize, SrcSize, Src2);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
@@ -2121,7 +2121,7 @@ Ref OpDispatchBuilder::CVTFPR_To_GPRImpl(OpcodeArgs, Ref Src, IR::OpSize SrcElem
Ref Converted = _Float_ToGPR_ZS(GPRSize, SrcElementSize, Src);
bool Dst32 = GPRSize == OpSize::i32Bit;
Ref MaxI = Dst32 ? _Constant(0x80000000) : _Constant(0x8000000000000000);
Ref MaxI = Dst32 ? Constant(0x80000000) : Constant(0x8000000000000000);
Ref MaxF = LoadAndCacheNamedVectorConstant(SrcElementSize, (SrcElementSize == OpSize::i32Bit) ?
(Dst32 ? NAMED_VECTOR_CVTMAX_F32_I32 : NAMED_VECTOR_CVTMAX_F32_I64) :
(Dst32 ? NAMED_VECTOR_CVTMAX_F64_I32 : NAMED_VECTOR_CVTMAX_F64_I64));
@@ -2134,7 +2134,7 @@ void OpDispatchBuilder::CVTFPR_To_GPR(OpcodeArgs) {
// If loading a vector, use the full size, so we don't
// unnecessarily zero extend the vector. Otherwise, if
// memory, then we want to load the element size exactly.
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : SrcElementSize;
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
Ref Result = CVTFPR_To_GPRImpl(Op, Src, SrcElementSize, HostRoundingMode);
StoreResult(GPRClass, Op, Result, OpSize::iInvalid);
@@ -2355,7 +2355,7 @@ void OpDispatchBuilder::VMASKMOVOpImpl(OpcodeArgs, IR::OpSize ElementSize, IR::O
const X86Tables::DecodedOperand& MaskOp, const X86Tables::DecodedOperand& DataOp) {
const auto MakeAddress = [this, Op](const X86Tables::DecodedOperand& Data) {
return MakeSegmentAddress(Op, Data, CTX->GetGPROpSize());
return MakeSegmentAddress(Op, Data, GetGPROpSize());
};
Ref Mask = LoadSource_WithOpSize(FPRClass, Op, MaskOp, DataSize, Op->Flags);
@@ -2398,7 +2398,7 @@ void OpDispatchBuilder::MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType) {
Ref Result {};
if (Op->Src[0].IsGPR()) {
// Loading from GPR and moving to Vector.
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], CTX->GetGPROpSize(), Op->Flags);
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], GetGPROpSize(), Op->Flags);
// zext to 128bit
Result = _VCastFromGPR(OpSize::i128Bit, OpSizeFromSrc(Op), Src);
} else {
@@ -2504,7 +2504,7 @@ Ref OpDispatchBuilder::XSaveBase(X86Tables::DecodedOp Op) {
void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// NOTE: Mask should be EAX and EDX concatenated, but we only need to test
// for features that are in the lower 32 bits, so EAX only is sufficient.
const auto OpSize = CTX->GetGPROpSize();
const auto OpSize = GetGPROpSize();
const auto StoreIfFlagSet = [this, OpSize](uint32_t BitIndex, auto fn, uint32_t FieldSize = 1) {
Ref Mask = LoadGPRRegister(X86State::REG_RAX);
@@ -2539,8 +2539,7 @@ void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// We need to save MXCSR and MXCSR_MASK if either SSE or AVX are requested to be saved
{
StoreIfFlagSet(
1, [this, Op] { SaveMXCSRState(XSaveBase(Op)); }, 2);
StoreIfFlagSet(1, [this, Op] { SaveMXCSRState(XSaveBase(Op)); }, 2);
}
// Update XSTATE_BV region of the XSAVE header
@@ -2553,7 +2552,7 @@ void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// XSTATE_BV section of the header is 8 bytes in size, but we only really
// care about setting at most 3 bits in the first byte. We zero out the rest.
_StoreMem(GPRClass, OpSize::i64Bit, RequestedFeatures, Base, _Constant(512), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, OpSize::i64Bit, RequestedFeatures, Base, Constant(512), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
}
}
@@ -2577,11 +2576,11 @@ void OpDispatchBuilder::SaveX87State(OpcodeArgs, Ref MemBase) {
_StoreMem(GPRClass, OpSize::i16Bit, MemBase, FCW, OpSize::i16Bit);
}
{ _StoreMem(GPRClass, OpSize::i16Bit, ReconstructFSW_Helper(), MemBase, _Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, OpSize::i16Bit, ReconstructFSW_Helper(), MemBase, Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1); }
{
// Abridged FTW
_StoreMem(GPRClass, OpSize::i8Bit, LoadContext(AbridgedFTWIndex), MemBase, _Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, OpSize::i8Bit, LoadContext(AbridgedFTWIndex), MemBase, Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
}
// BYTE | 0 1 | 2 3 | 4 | 5 | 6 7 | 8 9 | a b | c d | e f |
@@ -2635,7 +2634,7 @@ void OpDispatchBuilder::SaveX87State(OpcodeArgs, Ref MemBase) {
}
void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
_StoreMemPair(FPRClass, OpSize::i128Bit, LoadXMMRegister(i), LoadXMMRegister(i + 1), MemBase, i * 16 + 160);
@@ -2644,11 +2643,11 @@ void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
void OpDispatchBuilder::SaveMXCSRState(Ref MemBase) {
// Store MXCSR and the mask for all bits.
_StoreMemPair(GPRClass, OpSize::i32Bit, GetMXCSR(), _Constant(0xFFFF), MemBase, 24);
_StoreMemPair(GPRClass, OpSize::i32Bit, GetMXCSR(), Constant(0xFFFF), MemBase, 24);
}
void OpDispatchBuilder::SaveAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
Ref Upper0 = _VDupElement(OpSize::i256Bit, OpSize::i128Bit, LoadXMMRegister(i + 0), 1);
@@ -2662,7 +2661,7 @@ Ref OpDispatchBuilder::GetMXCSR() {
Ref MXCSR = _LoadContext(OpSize::i32Bit, GPRClass, offsetof(FEXCore::Core::CPUState, mxcsr));
// Mask out unsupported bits
// Keeps FZ, RC, exception masks, and DAZ
MXCSR = _And(OpSize::i32Bit, MXCSR, _Constant(0xFFC0));
MXCSR = _And(OpSize::i32Bit, MXCSR, Constant(0xFFC0));
return MXCSR;
}
@@ -2672,12 +2671,12 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
RestoreX87State(Mem);
RestoreSSEState(Mem);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Mem, _Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Mem, Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
}
void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
const auto OpSize = CTX->GetGPROpSize();
const auto OpSize = GetGPROpSize();
// If a bit in our XSTATE_BV is set, then we restore from that region of the XSAVE area,
// otherwise, if not set, then we need to set the relevant data the bit corresponds to
@@ -2689,7 +2688,7 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
// Note: we rematerialize Base/Mask in each block to avoid crossblock
// liveness.
Ref Base = XSaveBase(Op);
Ref Mask = _LoadMem(GPRClass, OpSize::i64Bit, Base, _Constant(512), OpSize::i64Bit, MEM_OFFSET_SXTX, 1);
Ref Mask = _LoadMem(GPRClass, OpSize::i64Bit, Base, Constant(512), OpSize::i64Bit, MEM_OFFSET_SXTX, 1);
Ref BitFlag = _Bfe(OpSize, FieldSize, BitIndex, Mask);
auto CondJump_ = CondJump(BitFlag, {COND_NEQ});
@@ -2715,13 +2714,11 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
// x87
{
RestoreIfFlagSetOrDefault(
0, [this, Op] { RestoreX87State(XSaveBase(Op)); }, [this, Op] { DefaultX87State(Op); });
RestoreIfFlagSetOrDefault(0, [this, Op] { RestoreX87State(XSaveBase(Op)); }, [this, Op] { DefaultX87State(Op); });
}
// SSE
{
RestoreIfFlagSetOrDefault(
1, [this, Op] { RestoreSSEState(XSaveBase(Op)); }, [this] { DefaultSSEState(); });
RestoreIfFlagSetOrDefault(1, [this, Op] { RestoreSSEState(XSaveBase(Op)); }, [this] { DefaultSSEState(); });
}
// AVX
if (CTX->HostFeatures.SupportsAVX) {
@@ -2734,9 +2731,9 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
RestoreIfFlagSetOrDefault(
1,
[this, Op] {
Ref Base = XSaveBase(Op);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Base, _Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
Ref Base = XSaveBase(Op);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Base, Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
},
[] { /* Intentionally do nothing*/ }, 2);
}
@@ -2747,13 +2744,13 @@ void OpDispatchBuilder::RestoreX87State(Ref MemBase) {
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
{
auto NewFSW = _LoadMem(GPRClass, OpSize::i16Bit, MemBase, _Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, OpSize::i16Bit, MemBase, Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
}
{
// Abridged FTW
StoreContext(AbridgedFTWIndex, _LoadMem(GPRClass, OpSize::i8Bit, MemBase, _Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1));
StoreContext(AbridgedFTWIndex, _LoadMem(GPRClass, OpSize::i8Bit, MemBase, Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1));
}
for (uint32_t i = 0; i < Core::CPUState::NUM_MMS; i += 2) {
@@ -2765,7 +2762,7 @@ void OpDispatchBuilder::RestoreX87State(Ref MemBase) {
}
void OpDispatchBuilder::RestoreSSEState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
auto XMMRegs = LoadMemPair(FPRClass, OpSize::i128Bit, MemBase, i * 16 + 160);
@@ -2777,7 +2774,7 @@ void OpDispatchBuilder::RestoreSSEState(Ref MemBase) {
void OpDispatchBuilder::RestoreMXCSRState(Ref MXCSR) {
// Mask out unsupported bits
MXCSR = _And(OpSize::i32Bit, MXCSR, _Constant(0xFFC0));
MXCSR = _And(OpSize::i32Bit, MXCSR, Constant(0xFFC0));
_StoreContext(OpSize::i32Bit, GPRClass, MXCSR, offsetof(FEXCore::Core::CPUState, mxcsr));
// We only support the rounding mode and FTZ bit being set
@@ -2786,7 +2783,7 @@ void OpDispatchBuilder::RestoreMXCSRState(Ref MXCSR) {
}
void OpDispatchBuilder::RestoreAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
Ref XMMReg0 = LoadXMMRegister(i + 0);
@@ -2811,7 +2808,7 @@ void OpDispatchBuilder::DefaultX87State(OpcodeArgs) {
}
void OpDispatchBuilder::DefaultSSEState() {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
Ref ZeroVector = LoadZeroVector(OpSize::i128Bit);
for (uint32_t i = 0; i < NumRegs; ++i) {
@@ -2820,7 +2817,7 @@ void OpDispatchBuilder::DefaultSSEState() {
}
void OpDispatchBuilder::DefaultAVXState() {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i++) {
Ref Reg = LoadXMMRegister(i);
@@ -2877,7 +2874,7 @@ void OpDispatchBuilder::VPALIGNROp(OpcodeArgs) {
template<IR::OpSize ElementSize>
void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
Ref Src1 = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetGuestVectorLength(), Op->Flags);
Ref Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
@@ -2888,12 +2885,12 @@ template void OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>(OpcodeArgs);
template void OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>(OpcodeArgs);
void OpDispatchBuilder::LDMXCSR(OpcodeArgs) {
Ref Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags);
Ref Dest = LoadSource_WithOpSize(GPRClass, Op, Op->Dest, OpSize::i32Bit, Op->Flags);
RestoreMXCSRState(Dest);
}
void OpDispatchBuilder::STMXCSR(OpcodeArgs) {
StoreResult(GPRClass, Op, GetMXCSR(), OpSize::iInvalid);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GetMXCSR(), OpSize::i32Bit, OpSize::iInvalid);
}
template<IR::OpSize ElementSize>
@@ -3953,10 +3950,7 @@ void OpDispatchBuilder::PTestOpImpl(OpSize Size, Ref Dest, Ref Src) {
Test1 = _VExtractToGPR(Size, OpSize::i16Bit, Test1, 0);
Test2 = _VExtractToGPR(Size, OpSize::i16Bit, Test2, 0);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
Test2 = _Select(FEXCore::IR::COND_NEQ, Test2, ZeroConst, OneConst, ZeroConst);
Test2 = To01(OpSize::i64Bit, Test2);
// Careful, these flags are different between {V,}PTEST and VTESTP{S,D}
// Set ZF according to Test1. SF will be zeroed since we do a 32-bit test on
@@ -3979,7 +3973,7 @@ void OpDispatchBuilder::VTESTOpImpl(OpSize SrcSize, IR::OpSize ElementSize, Ref
const auto ElementSizeInBits = IR::OpSizeAsBits(ElementSize);
const auto MaskConstant = uint64_t {1} << (ElementSizeInBits - 1);
Ref Mask = _VDupFromGPR(SrcSize, ElementSize, _Constant(MaskConstant));
Ref Mask = _VDupFromGPR(SrcSize, ElementSize, Constant(MaskConstant));
Ref AndTest = _VAnd(SrcSize, OpSize::i8Bit, Src2, Src1);
Ref AndNotTest = _VAndn(SrcSize, OpSize::i8Bit, Src2, Src1);
@@ -3993,10 +3987,7 @@ void OpDispatchBuilder::VTESTOpImpl(OpSize SrcSize, IR::OpSize ElementSize, Ref
Ref AndGPR = _VExtractToGPR(SrcSize, OpSize::i16Bit, MaxAnd, 0);
Ref AndNotGPR = _VExtractToGPR(SrcSize, OpSize::i16Bit, MaxAndNot, 0);
Ref ZeroConst = _Constant(0);
Ref OneConst = _Constant(1);
Ref CFInv = _Select(IR::COND_NEQ, AndNotGPR, ZeroConst, OneConst, ZeroConst);
Ref CFInv = To01(OpSize::i64Bit, AndNotGPR);
// As in PTest, this sets Z appropriately while zeroing the rest of NZCV.
SetNZ_ZeroCV(OpSize::i32Bit, AndGPR);
@@ -4583,7 +4574,7 @@ void OpDispatchBuilder::VPERMDOp(OpcodeArgs) {
// Get rid of any junk unrelated to the relevant selector index bits (bits [2:0])
Ref IndexMask = _VectorImm(DstSize, OpSize::i32Bit, 0b111);
Ref AddConst = _Constant(0x03020100);
Ref AddConst = Constant(0x03020100);
Ref Repeating3210 = _VDupFromGPR(DstSize, OpSize::i32Bit, AddConst);
Ref FinalIndices = VPERMDIndices(OpSizeFromDst(Op), Indices, IndexMask, Repeating3210);
@@ -4732,7 +4723,7 @@ void OpDispatchBuilder::VPBLENDWOp(OpcodeArgs) {
void OpDispatchBuilder::VZEROOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto IsVZEROALL = DstSize == OpSize::i256Bit;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
if (IsVZEROALL) {
// NOTE: Despite the name being VZEROALL, this will still only ever
@@ -4818,7 +4809,7 @@ Ref OpDispatchBuilder::VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize,
Ref ShiftedIndices = _VShlI(DstSize, OpSize::i8Bit, IndexTrn3, IndexShift);
uint64_t VConstant = IsPD ? 0x0706050403020100 : 0x03020100;
Ref VectorConst = _VDupFromGPR(DstSize, ElementSize, _Constant(VConstant));
Ref VectorConst = _VDupFromGPR(DstSize, ElementSize, Constant(VConstant));
Ref FinalIndices {};
if (Is256Bit) {
@@ -4877,7 +4868,7 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask
IntermediateResult = _VPCMPISTRX(Src1, Src2, Control);
}
Ref ZeroConst = _Constant(0);
Ref ZeroConst = Constant(0);
if (IsMask) {
// For the masked variant of the instructions, if control[6] is set, then we
@@ -4914,7 +4905,7 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask
Ref ResultNoFlags = _Bfe(OpSize::i32Bit, 16, 0, IntermediateResult);
Ref IfZero = _Constant(16 >> (Control & 1));
Ref IfZero = Constant(16 >> (Control & 1));
Ref IfNotZero = UseMSBIndex ? _FindMSB(IR::OpSize::i32Bit, ResultNoFlags) : _FindLSB(IR::OpSize::i32Bit, ResultNoFlags);
Ref Result = _Select(IR::COND_EQ, ResultNoFlags, ZeroConst, IfZero, IfNotZero);
@@ -4947,9 +4938,14 @@ void OpDispatchBuilder::VFMAImpl(OpcodeArgs, IROps IROp, bool Scalar, uint8_t Sr
const OpSize ElementSize = Op->Flags & X86Tables::DecodeFlags::FLAG_OPTION_AVX_W ? OpSize::i64Bit : OpSize::i32Bit;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, Size, Op->Flags);
Ref Src1 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Size, Op->Flags);
Ref Src2 {};
if (Op->Src[1].IsGPR()) {
Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], Size, Op->Flags);
} else {
Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
}
Ref Sources[3] = {
Dest,
@@ -5005,7 +5001,7 @@ void OpDispatchBuilder::VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx,
}
OpDispatchBuilder::RefVSIB OpDispatchBuilder::LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags) {
[[maybe_unused]] const bool IsVSIB = (Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0;
const bool IsVSIB = (Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0;
LOGMAN_THROW_A_FMT(Operand.IsSIB() && IsVSIB, "Trying to load VSIB for something that isn't the correct type!");
// VSIB is a very special case which has a ton of encoded data.
@@ -5082,9 +5078,9 @@ void OpDispatchBuilder::VPGATHER(OpcodeArgs) {
///< BaseAddr doesn't need to exist, calculate that here.
Ref BaseAddr = VSIB.BaseAddr;
if (BaseAddr && VSIB.Displacement) {
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, _Constant(VSIB.Displacement));
BaseAddr = Add(OpSize::i64Bit, BaseAddr, VSIB.Displacement);
} else if (VSIB.Displacement) {
BaseAddr = _Constant(VSIB.Displacement);
BaseAddr = Constant(VSIB.Displacement);
} else if (!BaseAddr) {
BaseAddr = Invalid();
}
@@ -32,21 +32,25 @@ Ref OpDispatchBuilder::GetX87Top() {
}
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref NewAbridgedFTW {};
// For the output, we want a 1-bit for each pair not equal to 11 (Empty).
static_assert(static_cast<uint8_t>(FPState::X87Tag::Empty) == 0b11);
for (int i = 0; i < 8; i++) {
Ref RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
Ref RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, _Constant(1), _Constant(0));
// Make even bits 1 if the pair is equal to 11, and 0 otherwise.
FTW = _AndShift(OpSize::i32Bit, FTW, FTW, ShiftType::LSR, 1);
if (i) {
NewAbridgedFTW = _Orlshl(OpSize::i32Bit, NewAbridgedFTW, RegValid, i);
} else {
NewAbridgedFTW = RegValid;
}
}
// Invert FTW and clear the odd bits. Even bits are 1 if the pair
// is not equal to 11, and odd bits are 0.
FTW = _Andn(OpSize::i32Bit, Constant(0x55555555), FTW);
StoreContext(AbridgedFTWIndex, NewAbridgedFTW);
// All that's left is to compact away the odd bits. That is a Morton
// deinterleave operation, which has a standard solution. See
// https://stackoverflow.com/questions/3137266/how-to-de-interleave-bits-unmortonizing
FTW = _And(OpSize::i32Bit, _Orlshr(OpSize::i32Bit, FTW, FTW, 1), Constant(0x33333333));
FTW = _And(OpSize::i32Bit, _Orlshr(OpSize::i32Bit, FTW, FTW, 2), Constant(0x0f0f0f0f));
FTW = _Orlshr(OpSize::i32Bit, FTW, FTW, 4);
// ...and that's it. StoreContext implicitly does the final masking.
StoreContext(AbridgedFTWIndex, FTW);
}
void OpDispatchBuilder::SetX87Top(Ref Value) {
@@ -84,9 +88,9 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
_PopStackDestroy();
}
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant Constant) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, Constant);
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit, true);
}
@@ -104,16 +108,16 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
SaveNZCV();
// Extract sign and make integer absolute
auto zero = _Constant(0);
auto zero = Constant(0);
_SubNZCV(OpSize::i64Bit, Data, zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, _Constant(0x8000), zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClassType {COND_MI});
// left justify the absolute integer
auto shift = _Sub(OpSize::i64Bit, _Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shift = Sub(OpSize::i64Bit, Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shifted = _Lshl(OpSize::i64Bit, absolute, shift);
auto adjusted_exponent = _Sub(OpSize::i64Bit, _Constant(0x3fff + 63), shift);
auto adjusted_exponent = Sub(OpSize::i64Bit, Constant(0x3fff + 63), shift);
auto zeroed_exponent = _Select(COND_EQ, absolute, zero, zero, adjusted_exponent);
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
@@ -125,7 +129,7 @@ void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
@@ -148,6 +152,30 @@ void OpDispatchBuilder::FSTToStack(OpcodeArgs) {
void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref Data = _ReadStackValue(0);
// For 16-bit integers, we need to manually check for overflow
// since _F80CVTInt doesn't handle 16-bit overflow detection properly
if (Size == OpSize::i16Bit) {
// Extract the 80-bit float value to check for special cases
// Get the upper 64 bits which contain sign and exponent and then the exponent from upper.
Ref Upper = _VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, Data, 1);
Ref Exponent = _And(OpSize::i64Bit, Upper, Constant(0x7fff));
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
Ref IsSpecial = _NZCVSelect01({COND_EQ});
// For overflow detection, check if exponent indicates a value >= 2^15
// Biased exponent for 2^15 is 0x3fff + 15 = 0x400e
SubWithFlags(OpSize::i64Bit, Exponent, 0x400e);
Ref IsOverflow = _NZCVSelect01({COND_UGE});
// Set Invalid Operation flag if overflow or special value
Ref InvalidFlag = _Or(OpSize::i64Bit, IsSpecial, IsOverflow);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(InvalidFlag);
}
Data = _F80CVTInt(Size, Data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, OpSize::i8Bit);
@@ -312,17 +340,17 @@ Ref OpDispatchBuilder::GetX87FTW_Helper() {
// https://graphics.stanford.edu/~seander/bithacks.html#InterleaveBMN
Ref X = LoadContext(AbridgedFTWIndex);
X = _Orlshl(OpSize::i32Bit, X, X, 4);
X = _And(OpSize::i32Bit, X, _Constant(0x0f0f0f0f));
X = _And(OpSize::i32Bit, X, Constant(0x0f0f0f0f));
X = _Orlshl(OpSize::i32Bit, X, X, 2);
X = _And(OpSize::i32Bit, X, _Constant(0x33333333));
X = _And(OpSize::i32Bit, X, Constant(0x33333333));
X = _Orlshl(OpSize::i32Bit, X, X, 1);
X = _And(OpSize::i32Bit, X, _Constant(0x55555555));
X = _And(OpSize::i32Bit, X, Constant(0x55555555));
X = _Orlshl(OpSize::i32Bit, X, X, 1);
// The above sequence sets valid to 11 and empty to 00, so invert to finalize.
static_assert(static_cast<uint8_t>(FPState::X87Tag::Valid) == 0b00);
static_assert(static_cast<uint8_t>(FPState::X87Tag::Empty) == 0b11);
return _Xor(OpSize::i32Bit, X, _Constant(0xffff));
return _Xor(OpSize::i32Bit, X, Constant(0xffff));
}
void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
@@ -359,33 +387,33 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
}
@@ -397,11 +425,13 @@ Ref OpDispatchBuilder::ReconstructX87StateFromFSW_Helper(Ref FSW) {
auto C1 = _Bfe(OpSize::i32Bit, 1, 9, FSW);
auto C2 = _Bfe(OpSize::i32Bit, 1, 10, FSW);
auto C3 = _Bfe(OpSize::i32Bit, 1, 14, FSW);
auto IE = _Bfe(OpSize::i32Bit, 1, 0, FSW);
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(C0);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(C1);
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(C2);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(C3);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(IE);
return Top;
}
@@ -415,13 +445,13 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 1));
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 1);
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 2));
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 2);
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
}
}
@@ -455,45 +485,44 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto SevenConst = Constant(7);
const auto LoadSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
// The final st(7) needs a bit of special handling here
@@ -504,9 +533,9 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(OpSize::i128Bit, OpSize::i16Bit, data, 4);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// reset to default
FNINIT(Op);
@@ -523,29 +552,27 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = NewFCW;
auto roundShift = _Constant(10);
auto roundMask = _Constant(3);
auto roundShift = Constant(10);
auto roundMask = Constant(3);
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
roundingMode = _And(OpSize::i32Bit, roundingMode, roundMask);
_SetRoundingMode(roundingMode, false, roundingMode);
}
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
auto SevenConst = Constant(7);
auto low = Constant(~0ULL);
auto high = Constant(0xFFFF);
Ref Mask = _VLoadTwoGPRs(low, high);
const auto StoreSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
@@ -554,16 +581,16 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
// The final st(7) needs a bit of special handling here
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref RegHigh =
_LoadMem(FPRClass, OpSize::i16Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_LoadMem(FPRClass, OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(OpSize::i128Bit, OpSize::i16Bit, 4, 0, Reg, RegHigh);
if (ReducedPrecisionMode) {
Reg = _F80CVT(OpSize::i64Bit, Reg); // Convert to double precision
@@ -592,13 +619,13 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
if (Offset != 0) {
_F80StackXchange(Offset);
}
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
}
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0x3FF0000000000000)) :
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
}
@@ -639,7 +666,7 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
if (WhichFlags == FCOMIFlags::FLAGS_X87) {
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
} else {
@@ -649,10 +676,13 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
// PF is stored inverted, so invert from the host flag.
// TODO: This could perhaps be optimized?
auto PF = _Xor(OpSize::i32Bit, HostFlag_Unordered, _Constant(1));
auto PF = _Xor(OpSize::i32Bit, HostFlag_Unordered, Constant(1));
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(PF);
}
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
if (PopTwice) {
_PopStackDestroy();
_PopStackDestroy();
@@ -671,15 +701,18 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
HostFlag_ZF = _Or(OpSize::i32Bit, HostFlag_ZF, HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
}
void OpDispatchBuilder::X87OpHelper(OpcodeArgs, FEXCore::IR::IROps IROp, bool ZeroC2) {
DeriveOp(Result, IROp, _F80SCALEStack());
if (ZeroC2) {
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(Constant(0));
}
}
@@ -705,7 +738,7 @@ void OpDispatchBuilder::X87ModifySTP(OpcodeArgs, bool Inc) {
Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
// Start with the top value
auto Top = T ? T : GetX87Top();
Ref FSW = _Lshl(OpSize::i64Bit, Top, _Constant(11));
Ref FSW = _Lshl(OpSize::i64Bit, Top, Constant(11));
// We must construct the FSW from our various bits
auto C0 = GetRFLAG(FEXCore::X86State::X87FLAG_C0_LOC);
@@ -720,6 +753,9 @@ Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
auto C3 = GetRFLAG(FEXCore::X86State::X87FLAG_C3_LOC);
FSW = _Orlshl(OpSize::i64Bit, FSW, C3, 14);
auto IE = GetRFLAG(FEXCore::X86State::X87FLAG_IE_LOC);
FSW = _Or(OpSize::i64Bit, FSW, IE);
return FSW;
}
@@ -732,15 +768,20 @@ void OpDispatchBuilder::X87FNSTSW(OpcodeArgs) {
StoreResult(GPRClass, Op, StatusWord, OpSize::iInvalid);
}
void OpDispatchBuilder::FNCLEX(OpcodeArgs) {
// Clear the exception flag bit
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(_Constant(0));
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
auto Zero = _Constant(0);
auto Zero = Constant(0);
if (ReducedPrecisionMode) {
_SetRoundingMode(Zero, false, Zero);
}
// Init FCW to 0x037F
auto NewFCW = _Constant(OpSize::i16Bit, 0x037F);
auto NewFCW = Constant(0x037F);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Set top to zero
@@ -755,6 +796,7 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(Zero);
}
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
@@ -800,11 +842,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
default: LOGMAN_MSG_A_FMT("Unhandled FCMOV op: 0x{:x}", Opcode); break;
}
auto ZeroConst = _Constant(0);
auto AllOneConst = _Constant(0xffff'ffff'ffff'ffffull);
Ref SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SrcCond);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SelectCC0All1(CC));
_F80VBSLStack(OpSize::i128Bit, VecCond, Op->OP & 7, 0);
}
@@ -820,11 +858,9 @@ void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
// Claim this is a normal number
// We don't support anything else
auto TopValid = _StackValidTag(0);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// In the case of top being invalid then C3:C2:C0 is 0b101
auto C3 = _Select(FEXCore::IR::COND_NEQ, TopValid, OneConst, OneConst, ZeroConst);
auto C3 = Select01(OpSize::i32Bit, CondClassType {COND_NEQ}, TopValid, Constant(1));
auto C2 = TopValid;
auto C0 = C3; // Mirror C3 until something other than zero is supported
@@ -36,12 +36,12 @@ void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
}
@@ -87,7 +87,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
}
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(Num));
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit, true);
}
@@ -377,21 +377,21 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Gpr = _VExtractToGPR(OpSize::i64Bit, OpSize::i64Bit, Node, 0);
// zero case
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0xfff0'0000'0000'0000UL));
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = _Sub(OpSize::i64Bit, ExpNZ, _Constant(1023));
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
Ref ExpNZV = _Float_FromGPR_S(OpSize::i64Bit, OpSize::i64Bit, ExpNZ);
Ref SigNZ = _And(OpSize::i64Bit, Gpr, _Constant(0x800f'ffff'ffff'ffffLL));
SigNZ = _Or(OpSize::i64Bit, SigNZ, _Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZ = _And(OpSize::i64Bit, Gpr, Constant(0x800f'ffff'ffff'ffffLL));
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
_TestNZ(OpSize::i64Bit, Gpr, _Constant(0x7fff'ffff'ffff'ffffUL));
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, ExpZV, ExpNZV);
@@ -33,7 +33,7 @@ X86GeneratedCode::X86GeneratedCode() {
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), &SignalReturnCode.at(0), SignalReturnCode.size());
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
@@ -1,29 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
meta: frontend|x86-tables ~ Metadata that drives the frontend x86/64 decoding
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
namespace FEXCore::X86Tables {
void InitializeBaseTables(Context::OperatingMode Mode);
void InitializeSecondaryTables(Context::OperatingMode Mode);
void InitializeSecondaryGroupTables(Context::OperatingMode Mode);
void InitializePrimaryGroupTables(Context::OperatingMode Mode);
void InitializeH0F3ATables(Context::OperatingMode Mode);
void InitializeInfoTables(Context::OperatingMode Mode) {
InitializeBaseTables(Mode);
InitializeSecondaryTables(Mode);
InitializeSecondaryGroupTables(Mode);
InitializePrimaryGroupTables(Mode);
InitializeH0F3ATables(Mode);
}
} // namespace FEXCore::X86Tables
@@ -15,300 +15,425 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
enum Primary_LUT {
ENTRY_06,
ENTRY_07,
ENTRY_0E,
ENTRY_16,
ENTRY_17,
ENTRY_1E,
ENTRY_1F,
ENTRY_27,
ENTRY_2F,
ENTRY_37,
ENTRY_3F,
ENTRY_40,
ENTRY_48,
ENTRY_60,
ENTRY_61,
ENTRY_63,
ENTRY_9A,
ENTRY_A0,
ENTRY_A1,
ENTRY_A2,
ENTRY_A3,
ENTRY_CE,
ENTRY_D4,
ENTRY_D5,
ENTRY_D6,
ENTRY_EA,
ENTRY_MAX,
};
constexpr std::array<X86InstInfo[2], ENTRY_MAX> Primary_ArchSelect_LUT = {{
// ENTRY_06
{
{"PUSH ES", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_07
{
{"POP ES", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_0E
{
{"PUSH CS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_16
{
{"PUSH SS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_17
{
{"POP SS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_1E
{
{"PUSH DS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_1F
{
{"POP DS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_27
{
{"DAA", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::DAAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_2F
{
{"DAS", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::DASOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_37
{
{"AAA", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::AAAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_3F
{
{"AAS", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::AASOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_40
{
{"INC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, { .OpDispatch = &IR::OpDispatchBuilder::INCOp } },
// REX
{"", TYPE_REX_PREFIX, FLAGS_NONE, 0},
},
// ENTRY_48
{
{"DEC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, { .OpDispatch = &IR::OpDispatchBuilder::DECOp } },
{"", TYPE_REX_PREFIX, FLAGS_NONE, 0},
},
// ENTRY_60
{
{"PUSHA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::PUSHAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_61
{
{"POPA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::POPAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_63
{
{"ARPL", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"MOVSXD", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM, 0, { .OpDispatch = &IR::OpDispatchBuilder::MOVSXDOp } },
},
// ENTRY_9A
{
{"CALLF", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_A0
{
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A1
{
{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A2
{
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A3
{
{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_CE
{
{"INTO", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = &IR::OpDispatchBuilder::INTOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D4
{
{"AAM", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, { .OpDispatch = &IR::OpDispatchBuilder::AAMOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D5
{
{"AAD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, { .OpDispatch = &IR::OpDispatchBuilder::AADOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D6
{
{"SALC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::SALCOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_EA
{
{"JMPF", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
}};
const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct BaseOpTable[] = {
// Prefixes
// Operand size overide
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0}},
// Address size override
{0x67, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x26, 1, X86InstInfo{"ES", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x2E, 1, X86InstInfo{"CS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x36, 1, X86InstInfo{"SS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x3E, 1, X86InstInfo{"DS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x67, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0}},
{0x26, 1, X86InstInfo{"ES", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x2E, 1, X86InstInfo{"CS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x36, 1, X86InstInfo{"SS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x3E, 1, X86InstInfo{"DS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
// These are still invalid on 64bit
{0x64, 1, X86InstInfo{"FS", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x65, 1, X86InstInfo{"GS", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF0, 1, X86InstInfo{"LOCK", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF2, 1, X86InstInfo{"REPNE", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x64, 1, X86InstInfo{"FS", TYPE_PREFIX, FLAGS_NONE, 0}},
{0x65, 1, X86InstInfo{"GS", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF0, 1, X86InstInfo{"LOCK", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF2, 1, X86InstInfo{"REPNE", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0}},
// Instructions
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x05, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x05, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x06, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_06] }}},
{0x07, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_07] }}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x15, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x0E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_0E] }}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x1D, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x15, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x16, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_16] }}},
{0x17, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_17] }}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x1D, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x1E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1E] }}},
{0x1F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1F] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x35, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x27, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_27] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x2F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_2F] }}},
{0x38, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x39, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x3A, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x3B, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x3C, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x3D, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x35, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x50, 8, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0x58, 8, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0x37, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_37] }}},
{0x38, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x39, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x3A, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x3B, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM, 0}},
{0x3C, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x3D, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x3F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_3F] }}},
{0x62, 1, X86InstInfo{"", TYPE_GROUP_EVEX, FLAGS_NONE, 0, nullptr}},
{0x40, 8, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_40] }}},
{0x48, 8, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_48] }}},
{0x68, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SRC_SEXT, 4, nullptr}},
{0x69, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x6A, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SRC_SEXT , 1, nullptr}},
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1, nullptr}},
{0x50, 8, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0x58, 8, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0x60, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_60] }}},
{0x61, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_61] }}},
{0x62, 1, X86InstInfo{"", TYPE_GROUP_EVEX, FLAGS_NONE, 0}},
{0x63, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_63] }}},
{0x68, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SRC_SEXT, 4}},
{0x69, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x6A, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SRC_SEXT , 1}},
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1}},
// This should just throw a GP
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x72, 1, X86InstInfo{"JB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x73, 1, X86InstInfo{"JNB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x74, 1, X86InstInfo{"JZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x75, 1, X86InstInfo{"JNZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x76, 1, X86InstInfo{"JBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x77, 1, X86InstInfo{"JNBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x78, 1, X86InstInfo{"JS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x79, 1, X86InstInfo{"JNS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7A, 1, X86InstInfo{"JP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7B, 1, X86InstInfo{"JNP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7C, 1, X86InstInfo{"JL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7D, 1, X86InstInfo{"JNL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7E, 1, X86InstInfo{"JLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7F, 1, X86InstInfo{"JNLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x72, 1, X86InstInfo{"JB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x73, 1, X86InstInfo{"JNB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x74, 1, X86InstInfo{"JZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x75, 1, X86InstInfo{"JNZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x76, 1, X86InstInfo{"JBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x77, 1, X86InstInfo{"JNBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x78, 1, X86InstInfo{"JS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x79, 1, X86InstInfo{"JNS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7A, 1, X86InstInfo{"JP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7B, 1, X86InstInfo{"JNP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7C, 1, X86InstInfo{"JL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7D, 1, X86InstInfo{"JNL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7E, 1, X86InstInfo{"JLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7F, 1, X86InstInfo{"JNLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x8A, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x8B, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x8C, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM, 0, nullptr}},
{0x8E, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{0x8F, 1, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_ZERO_REG | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x90, 8, X86InstInfo{"XCHG", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x98, 1, X86InstInfo{"CDQE", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x99, 1, X86InstInfo{"CQO", TYPE_INST, FLAGS_SF_DST_RDX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x8A, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x8B, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM, 0}},
{0x8C, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM, 0}},
{0x8E, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{0x8F, 1, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_ZERO_REG | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0x90, 8, X86InstInfo{"XCHG", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_SF_SRC_RAX, 0}},
{0x98, 1, X86InstInfo{"CDQE", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0}},
{0x99, 1, X86InstInfo{"CQO", TYPE_INST, FLAGS_SF_DST_RDX | FLAGS_SF_SRC_RAX, 0}},
// These three are all X87 instructions
{0x9B, 1, X86InstInfo{"FWAIT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9C, 1, X86InstInfo{"PUSHF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0, nullptr}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_BLOCK_END, 0, nullptr}},
{0x9B, 1, X86InstInfo{"FWAIT", TYPE_INST, FLAGS_NONE, 0}},
{0x9C, 1, X86InstInfo{"PUSHF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_BLOCK_END, 0}},
{0x9E, 1, X86InstInfo{"SAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9F, 1, X86InstInfo{"LAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9E, 1, X86InstInfo{"SAHF", TYPE_INST, FLAGS_NONE, 0}},
{0x9F, 1, X86InstInfo{"LAHF", TYPE_INST, FLAGS_NONE, 0}},
{0xA4, 1, X86InstInfo{"MOVSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA5, 1, X86InstInfo{"MOVS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA6, 1, X86InstInfo{"CMPSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA7, 1, X86InstInfo{"CMPS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA0, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A0] }}},
{0xA1, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A1] }}},
{0xA2, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A2] }}},
{0xA3, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A3] }}},
{0xA8, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0xA9, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0xAA, 1, X86InstInfo{"STOS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAB, 1, X86InstInfo{"STOS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAC, 1, X86InstInfo{"LODS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xAD, 1, X86InstInfo{"LODS", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xAE, 1, X86InstInfo{"SCAS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAF, 1, X86InstInfo{"SCAS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xA4, 1, X86InstInfo{"MOVSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA5, 1, X86InstInfo{"MOVS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA6, 1, X86InstInfo{"CMPSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA7, 1, X86InstInfo{"CMPS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xB0, 8, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_REX_IN_BYTE , 1, nullptr}},
{0xB8, 8, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_DISPLACE_SIZE_MUL_2, 4, nullptr}},
{0xA8, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0xA9, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0xAA, 1, X86InstInfo{"STOS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAB, 1, X86InstInfo{"STOS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAC, 1, X86InstInfo{"LODS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xAD, 1, X86InstInfo{"LODS", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xAE, 1, X86InstInfo{"SCAS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAF, 1, X86InstInfo{"SCAS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2, nullptr}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0xCA, 2, X86InstInfo{"RETF", TYPE_PRIV, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1, nullptr}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xB0, 8, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_REX_IN_BYTE , 1}},
{0xB8, 8, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_DISPLACE_SIZE_MUL_2, 4}},
{0xD7, 1, X86InstInfo{"XLAT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0xCA, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2}},
{0xCB, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1}},
{0xCE, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_CE] }}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
{0xE0, 1, X86InstInfo{"LOOPNE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE1, 1, X86InstInfo{"LOOPE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE2, 1, X86InstInfo{"LOOP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0xD4, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D4] }}},
{0xD5, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D5] }}},
{0xD6, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D6] }}},
{0xD7, 1, X86InstInfo{"XLAT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xE0, 1, X86InstInfo{"LOOPNE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE1, 1, X86InstInfo{"LOOPE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE2, 1, X86InstInfo{"LOOP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
// Should just throw GP
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 1}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 1}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1, nullptr}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END | FLAGS_CALL , 4}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1}},
// Should just throw GP
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0}},
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0}},
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0}},
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0}},
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0}},
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0}},
// Two Byte table
{0x0F, 1, X86InstInfo{"", TYPE_SECONDARY_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x0F, 1, X86InstInfo{"", TYPE_SECONDARY_TABLE_PREFIX, FLAGS_NONE, 0}},
// x87 table
{0xD8, 8, X86InstInfo{"", TYPE_X87_TABLE_PREFIX, FLAGS_MODRM, 0, nullptr}},
{0xD8, 8, X86InstInfo{"", TYPE_X87_TABLE_PREFIX, FLAGS_MODRM, 0}},
// ModRM table
// MoreBytes field repurposed for valid bits mask
{0x80, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 0, nullptr}},
{0x81, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 1, nullptr}},
{0x82, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 2, nullptr}},
{0x83, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 3, nullptr}},
{0xC0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 0, nullptr}},
{0xC1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 1, nullptr}},
{0xD0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 2, nullptr}},
{0xD1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 3, nullptr}},
{0xD2, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 4, nullptr}},
{0xD3, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 5, nullptr}},
{0xF6, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 0, nullptr}},
{0xF7, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 1, nullptr}},
{0xFE, 1, X86InstInfo{"", TYPE_GROUP_4, FLAGS_MODRM, 0, nullptr}},
{0xFF, 1, X86InstInfo{"", TYPE_GROUP_5, FLAGS_MODRM, 0, nullptr}},
{0x80, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 0}},
{0x81, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 1}},
{0x82, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 2}},
{0x83, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 3}},
{0xC0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 0}},
{0xC1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 1}},
{0xD0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 2}},
{0xD1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 3}},
{0xD2, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 4}},
{0xD3, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 5}},
{0xF6, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 0}},
{0xF7, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 1}},
{0xFE, 1, X86InstInfo{"", TYPE_GROUP_4, FLAGS_MODRM, 0}},
{0xFF, 1, X86InstInfo{"", TYPE_GROUP_5, FLAGS_MODRM, 0}},
// Group 11
{0xC6, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 0, nullptr}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 1, nullptr}},
{0xC6, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 0}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 1}},
// VEX table
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0}},
};
GenerateTable(&Table.at(0), BaseOpTable, std::size(BaseOpTable));
GenerateTable(Table.data(), BaseOpTable, std::size(BaseOpTable));
IR::InstallToTable(Table, IR::OpDispatch_BaseOpTable);
return Table;
}();
void InitializeBaseTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct BaseOpTable_64[] = {
{0x06, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x0E, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x16, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x1E, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x27, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2F, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x37, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x3F, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// REX
{0x40, 16, X86InstInfo{"", TYPE_REX_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x60, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x63, 1, X86InstInfo{"MOVSXD", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM, 0, nullptr}},
{0x9A, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xA0, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA2, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA1, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA3, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xCE, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xD4, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// `L1OM` Larrabee instructions used this as an escape byte.
// FEX will never support this.
{0xD6, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xEA, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
static constexpr U8U8InfoStruct BaseOpTable_32[] = {
{0x06, 1, X86InstInfo{"PUSH ES", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x07, 1, X86InstInfo{"POP ES", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x0E, 1, X86InstInfo{"PUSH CS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x16, 1, X86InstInfo{"PUSH SS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x17, 1, X86InstInfo{"POP SS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1E, 1, X86InstInfo{"PUSH DS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1F, 1, X86InstInfo{"POP DS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x27, 1, X86InstInfo{"DAA", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x2F, 1, X86InstInfo{"DAS", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x37, 1, X86InstInfo{"AAA", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x3F, 1, X86InstInfo{"AAS", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x40, 8, X86InstInfo{"INC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
{0x48, 8, X86InstInfo{"DEC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
{0x60, 1, X86InstInfo{"PUSHA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x61, 1, X86InstInfo{"POPA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x63, 1, X86InstInfo{"ARPL", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x9A, 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xA0, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA2, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA1, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA3, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xCE, 1, X86InstInfo{"INTO", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xD4, 1, X86InstInfo{"AAM", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xD5, 1, X86InstInfo{"AAD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xD6, 1, X86InstInfo{"SALC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
if (Mode == Context::MODE_64BIT) {
GenerateTable(&BaseOps.at(0), BaseOpTable_64, std::size(BaseOpTable_64));
IR::InstallToTable(BaseOps, IR::OpDispatch_BaseOpTable_64);
}
else {
GenerateTable(&BaseOps.at(0), BaseOpTable_32, std::size(BaseOpTable_32));
IR::InstallToTable(BaseOps, IR::OpDispatch_BaseOpTable_32);
}
}
}
@@ -13,48 +13,48 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps = []() consteval {
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct DDDNowOpTable[] = {
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x1D, 1, X86InstInfo{"PF2ID", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x1D, 1, X86InstInfo{"PF2ID", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
// Inverse 3DNow! These two instructions are Geode product line specific
// No CPUID for these, you're expected to read ID_CONFIG_MSR (1250h) bit 1
{0x86, 1, X86InstInfo{"PFRCPV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x87, 1, X86InstInfo{"PFRSQRTV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x86, 1, X86InstInfo{"PFRCPV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x87, 1, X86InstInfo{"PFRSQRTV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x8A, 1, X86InstInfo{"PFNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x8E, 1, X86InstInfo{"PFPNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x8A, 1, X86InstInfo{"PFNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x8E, 1, X86InstInfo{"PFPNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x90, 1, X86InstInfo{"PFCMPGE", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x94, 1, X86InstInfo{"PFMIN", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x96, 1, X86InstInfo{"PFRCP", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x97, 1, X86InstInfo{"PFRSQRT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x90, 1, X86InstInfo{"PFCMPGE", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x94, 1, X86InstInfo{"PFMIN", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x96, 1, X86InstInfo{"PFRCP", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x97, 1, X86InstInfo{"PFRSQRT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x9A, 1, X86InstInfo{"PFSUB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x9E, 1, X86InstInfo{"PFADD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x9A, 1, X86InstInfo{"PFSUB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0x9E, 1, X86InstInfo{"PFADD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xA0, 1, X86InstInfo{"PFCMPGT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA4, 1, X86InstInfo{"PFMAX", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA6, 1, X86InstInfo{"PFRCPIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA7, 1, X86InstInfo{"PFRSQIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA0, 1, X86InstInfo{"PFCMPGT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xA4, 1, X86InstInfo{"PFMAX", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xA6, 1, X86InstInfo{"PFRCPIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xA7, 1, X86InstInfo{"PFRSQIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xAA, 1, X86InstInfo{"PFSUBR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xAE, 1, X86InstInfo{"PFACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xAA, 1, X86InstInfo{"PFSUBR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xAE, 1, X86InstInfo{"PFACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xB0, 1, X86InstInfo{"PFCMPEQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB4, 1, X86InstInfo{"PFMUL", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB6, 1, X86InstInfo{"PFRCPIT2", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB0, 1, X86InstInfo{"PFCMPEQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xB4, 1, X86InstInfo{"PFMUL", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xB6, 1, X86InstInfo{"PFRCPIT2", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xBB, 1, X86InstInfo{"PSWAPD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xBF, 1, X86InstInfo{"PAVGUSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xBB, 1, X86InstInfo{"PSWAPD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{0xBF, 1, X86InstInfo{"PAVGUSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
};
GenerateTable(&Table.at(0), DDDNowOpTable, std::size(DDDNowOpTable));
GenerateTable(Table.data(), DDDNowOpTable, std::size(DDDNowOpTable));
IR::InstallToTable(Table, IR::OpDispatch_DDDTable);
return Table;
@@ -13,7 +13,7 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps = []() consteval {
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> Table{};
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
@@ -23,103 +23,103 @@ std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps = []() consteval {
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x02), 1, X86InstInfo{"PHADDD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x02), 1, X86InstInfo{"PHADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x03), 1, X86InstInfo{"PHADDSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x03), 1, X86InstInfo{"PHADDSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x04), 1, X86InstInfo{"PMADDUBSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x04), 1, X86InstInfo{"PMADDUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x05), 1, X86InstInfo{"PHSUBW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x05), 1, X86InstInfo{"PHSUBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x06), 1, X86InstInfo{"PHSUBD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x06), 1, X86InstInfo{"PHSUBD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x07), 1, X86InstInfo{"PHSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x07), 1, X86InstInfo{"PHSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x08), 1, X86InstInfo{"PSIGNB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x08), 1, X86InstInfo{"PSIGNB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x09), 1, X86InstInfo{"PSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x09), 1, X86InstInfo{"PSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x0A), 1, X86InstInfo{"PSIGND", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x0A), 1, X86InstInfo{"PSIGND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x02), 1, X86InstInfo{"PHADDD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x02), 1, X86InstInfo{"PHADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x03), 1, X86InstInfo{"PHADDSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x03), 1, X86InstInfo{"PHADDSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x04), 1, X86InstInfo{"PMADDUBSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x04), 1, X86InstInfo{"PMADDUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x05), 1, X86InstInfo{"PHSUBW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x05), 1, X86InstInfo{"PHSUBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x06), 1, X86InstInfo{"PHSUBD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x06), 1, X86InstInfo{"PHSUBD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x07), 1, X86InstInfo{"PHSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x07), 1, X86InstInfo{"PHSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x08), 1, X86InstInfo{"PSIGNB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x08), 1, X86InstInfo{"PSIGNB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x09), 1, X86InstInfo{"PSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x09), 1, X86InstInfo{"PSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x0A), 1, X86InstInfo{"PSIGND", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x0A), 1, X86InstInfo{"PSIGND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x0B), 1, X86InstInfo{"PMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x10), 1, X86InstInfo{"PBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x14), 1, X86InstInfo{"BLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x15), 1, X86InstInfo{"BLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x17), 1, X86InstInfo{"PTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x1D), 1, X86InstInfo{"PABSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x1D), 1, X86InstInfo{"PABSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x10), 1, X86InstInfo{"PBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x14), 1, X86InstInfo{"BLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x15), 1, X86InstInfo{"BLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x17), 1, X86InstInfo{"PTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x1C), 1, X86InstInfo{"PABSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x1D), 1, X86InstInfo{"PABSW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x1D), 1, X86InstInfo{"PABSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0}},
{OPD(PF_38_66, 0x1E), 1, X86InstInfo{"PABSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x20), 1, X86InstInfo{"PMOVSXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x21), 1, X86InstInfo{"PMOVSXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x22), 1, X86InstInfo{"PMOVSXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x23), 1, X86InstInfo{"PMOVSXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x24), 1, X86InstInfo{"PMOVSXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x25), 1, X86InstInfo{"PMOVSXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x28), 1, X86InstInfo{"PMULDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x29), 1, X86InstInfo{"PCMPEQQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x2B), 1, X86InstInfo{"PACKUSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x20), 1, X86InstInfo{"PMOVSXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x21), 1, X86InstInfo{"PMOVSXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x22), 1, X86InstInfo{"PMOVSXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x23), 1, X86InstInfo{"PMOVSXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x24), 1, X86InstInfo{"PMOVSXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x25), 1, X86InstInfo{"PMOVSXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x28), 1, X86InstInfo{"PMULDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x29), 1, X86InstInfo{"PCMPEQQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x2A), 1, X86InstInfo{"MOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x2B), 1, X86InstInfo{"PACKUSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x30), 1, X86InstInfo{"PMOVZXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x31), 1, X86InstInfo{"PMOVZXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x32), 1, X86InstInfo{"PMOVZXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x37), 1, X86InstInfo{"PCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3B), 1, X86InstInfo{"PMINUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3C), 1, X86InstInfo{"PMAXSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3D), 1, X86InstInfo{"PMAXSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3E), 1, X86InstInfo{"PMAXUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3F), 1, X86InstInfo{"PMAXUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x30), 1, X86InstInfo{"PMOVZXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x31), 1, X86InstInfo{"PMOVZXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x32), 1, X86InstInfo{"PMOVZXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x37), 1, X86InstInfo{"PCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3B), 1, X86InstInfo{"PMINUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3C), 1, X86InstInfo{"PMAXSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3D), 1, X86InstInfo{"PMAXSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3E), 1, X86InstInfo{"PMAXUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x3F), 1, X86InstInfo{"PMAXUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xC8), 1, X86InstInfo{"SHA1NEXTE", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC9), 1, X86InstInfo{"SHA1MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCA), 1, X86InstInfo{"SHA1MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC8), 1, X86InstInfo{"SHA1NEXTE", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xC9), 1, X86InstInfo{"SHA1MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xCA), 1, X86InstInfo{"SHA1MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xCB), 1, X86InstInfo{"SHA256RNDS2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCC), 1, X86InstInfo{"SHA256MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCD), 1, X86InstInfo{"SHA256MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCB), 1, X86InstInfo{"SHA256RNDS2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xCC), 1, X86InstInfo{"SHA256MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xCD), 1, X86InstInfo{"SHA256MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0xDB), 1, X86InstInfo{"AESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDC), 1, X86InstInfo{"AESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDD), 1, X86InstInfo{"AESENCLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDE), 1, X86InstInfo{"AESDEC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDF), 1, X86InstInfo{"AESDECLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDB), 1, X86InstInfo{"AESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0xDC), 1, X86InstInfo{"AESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0xDD), 1, X86InstInfo{"AESENCLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0xDE), 1, X86InstInfo{"AESDEC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_66, 0xDF), 1, X86InstInfo{"AESDECLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0}},
{OPD(PF_38_NONE, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_NONE, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_NONE, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(PF_38_NONE, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(PF_38_66, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_66, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_66, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(PF_38_66, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0}},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(PF_38_66, 0xF6), 1, X86InstInfo{"ADCX", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0, nullptr}},
{OPD(PF_38_F3, 0xF6), 1, X86InstInfo{"ADOX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66, 0xF6), 1, X86InstInfo{"ADCX", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{OPD(PF_38_F3, 0xF6), 1, X86InstInfo{"ADOX", TYPE_INST, FLAGS_MODRM, 0}},
};
#undef OPD
GenerateTable(&Table.at(0), H0F38Table, std::size(H0F38Table));
GenerateTable(Table.data(), H0F38Table, std::size(H0F38Table));
IR::InstallToTable(Table, IR::OpDispatch_H0F38Table);
return Table;
@@ -19,71 +19,82 @@ using namespace InstFlags;
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = 1;
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps = []() consteval {
enum H0F3A_LUT {
ENTRY_1_3A_66_16,
ENTRY_1_3A_66_22,
ENTRY_MAX,
};
constexpr std::array<X86InstInfo[2], ENTRY_MAX> H0F3A_ArchSelect_LUT = {{
// ENTRY_1_3A_66_16
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"PEXTRQ", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PExtrOp, IR::OpSize::i64Bit> }},
},
// ENTRY_1_3A_66_22
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::PINSROp<IR::OpSize::i64Bit> }},
},
}};
constexpr std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps = []() consteval {
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> Table{};
auto TableGen = []<uint16_t REX>() consteval {
constexpr U16U8InfoStruct Table[] = {
{OPD(REX, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(REX, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x20), 1, X86InstInfo{"PINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1}},
{OPD(REX, PF_3A_66, 0x21), 1, X86InstInfo{"INSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x40), 1, X86InstInfo{"DPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x41), 1, X86InstInfo{"DPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{OPD(REX, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(REX, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
};
return std::to_array(Table);
};
constexpr auto H0F3ATable_IgnoresREX0 = TableGen.template operator()<0>();
constexpr auto H0F3ATable_IgnoresREX1 = TableGen.template operator()<1>();
GenerateTable(&Table.at(0), &H0F3ATable_IgnoresREX0.at(0), H0F3ATable_IgnoresREX0.size());
GenerateTable(&Table.at(0), &H0F3ATable_IgnoresREX1.at(0), H0F3ATable_IgnoresREX1.size());
GenerateTable(Table.data(), H0F3ATable_IgnoresREX0.data(), H0F3ATable_IgnoresREX0.size());
GenerateTable(Table.data(), H0F3ATable_IgnoresREX1.data(), H0F3ATable_IgnoresREX1.size());
constexpr U16U8InfoStruct TableNeedsREX[] = {
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
constexpr U16U8InfoStruct TableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1}},
{OPD(0, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1}},
};
GenerateTable(&Table.at(0), TableNeedsREX, std::size(TableNeedsREX));
GenerateTable(Table.data(), TableNeedsREX0, std::size(TableNeedsREX0));
constexpr U16U8InfoStruct TableNeedsREX1[] = {
{OPD(1, PF_3A_66, 0x16), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = H0F3A_ArchSelect_LUT[ENTRY_1_3A_66_16] }}},
{OPD(1, PF_3A_66, 0x22), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = H0F3A_ArchSelect_LUT[ENTRY_1_3A_66_22] }}},
};
GenerateTable(Table.data(), TableNeedsREX1, std::size(TableNeedsREX1));
IR::InstallToTable(Table, IR::OpDispatch_H0F3ATableIgnoreREX);
IR::InstallToTable(Table, IR::OpDispatch_H0F3ATableNeedsREX0);
return Table;
}();
void InitializeH0F3ATables(Context::OperatingMode Mode) {
static constexpr U16U8InfoStruct H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRQ", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x22), 1, X86InstInfo{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
};
#undef OPD
if (Mode == Context::MODE_64BIT) {
GenerateTable(&H0F3ATableOps.at(0), H0F3ATable_64, std::size(H0F3ATable_64));
IR::InstallToTable(H0F3ATableOps, IR::OpDispatch_H0F3ATable_64);
}
}
}
@@ -14,168 +14,197 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps = []() consteval {
enum PrimaryGroup_LUT {
ENTRY_1_82_0,
ENTRY_1_82_1,
ENTRY_1_82_2,
ENTRY_1_82_3,
ENTRY_1_82_4,
ENTRY_1_82_5,
ENTRY_1_82_6,
ENTRY_1_82_7,
ENTRY_MAX,
};
constexpr std::array<X86InstInfo[2], ENTRY_MAX> PrimaryGroup_ArchSelect_LUT = {{
{
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::CMPOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
}};
constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> Table{};
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr U16U8InfoStruct PrimaryGroupOpTable[] = {
// GROUP_1 | 0x80 | reg
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
// Duplicates the 0x80 opcode group
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_0] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 1), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_1] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 2), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_2] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 3), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_3] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 4), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_4] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 5), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_5] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 6), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_6] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 7), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_7] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
// GROUP 2
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xC1), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD0), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD1), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 1), 1, X86InstInfo{"ROR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 2), 1, X86InstInfo{"RCL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 3), 1, X86InstInfo{"RCR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 4), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 5), 1, X86InstInfo{"SHR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 6), 1, X86InstInfo{"SHL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD2), 7), 1, X86InstInfo{"SAR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0, nullptr}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 0), 1, X86InstInfo{"ROL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 1), 1, X86InstInfo{"ROR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 2), 1, X86InstInfo{"RCL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 3), 1, X86InstInfo{"RCR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 4), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 5), 1, X86InstInfo{"SHR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 6), 1, X86InstInfo{"SHL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
{OPD(TYPE_GROUP_2, OpToIndex(0xD3), 7), 1, X86InstInfo{"SAR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX, 0}},
// GROUP 3
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 0), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 1), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 4), 1, X86InstInfo{"MUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 5), 1, X86InstInfo{"IMUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 1, X86InstInfo{"DIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 7), 1, X86InstInfo{"IDIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 0), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 1), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 4), 1, X86InstInfo{"MUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 5), 1, X86InstInfo{"IMUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 1, X86InstInfo{"DIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 7), 1, X86InstInfo{"IDIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 0), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 1), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 4), 1, X86InstInfo{"MUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 5), 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 1, X86InstInfo{"DIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 7), 1, X86InstInfo{"IDIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 0), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 1), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 4), 1, X86InstInfo{"MUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 5), 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 1, X86InstInfo{"DIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 7), 1, X86InstInfo{"IDIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// GROUP 4
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 2), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 2), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 5
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 5), 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 6), 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END | FLAGS_CALL , 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 5), 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 6), 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 11
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 0), 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT, 1, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 7), 1, X86InstInfo{"XABORT", TYPE_INST, FLAGS_MODRM, 1, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 0), 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 7), 1, X86InstInfo{"XBEGIN", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_SETS_RIP | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 0), 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT, 1}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 7), 1, X86InstInfo{"XABORT", TYPE_INST, FLAGS_MODRM, 1}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 0), 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 7), 1, X86InstInfo{"XBEGIN", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_SETS_RIP | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
};
GenerateTable(&Table.at(0), PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
GenerateTable(Table.data(), PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
IR::InstallToTable(Table, IR::OpDispatch_PrimaryGroupTables);
return Table;
}();
void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
const U16U8InfoStruct PrimaryGroupOpTable_64[] = {
// Invalid in 64bit mode
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
const U16U8InfoStruct PrimaryGroupOpTable_32[] = {
// Duplicates the 0x80 opcode group
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
};
#undef OPD
if (Mode == Context::MODE_64BIT) {
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
}
else {
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable_32, std::size(PrimaryGroupOpTable_32));
}
}
}
@@ -13,14 +13,45 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_6) << 5) | (prefix) << 3 | (Reg))
constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
enum SecondGroup_LUT {
ENTRY_15_F3_0,
ENTRY_15_F3_1,
ENTRY_15_F3_2,
ENTRY_15_F3_3,
ENTRY_MAX,
};
constexpr std::array<X86InstInfo[2], ENTRY_MAX> SecondGroup_ArchSelect_LUT = {{
// ENTRY_15_F3_0
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"RDFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ReadSegmentReg, IR::OpDispatchBuilder::Segment::FS> } },
},
// ENTRY_15_F3_1
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"RDGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ReadSegmentReg, IR::OpDispatchBuilder::Segment::GS> } },
},
// ENTRY_15_F3_2
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"WRFSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::WriteSegmentReg, IR::OpDispatchBuilder::Segment::FS> } },
},
// ENTRY_15_F3_3
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"WRGSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::WriteSegmentReg, IR::OpDispatchBuilder::Segment::GS> } },
},
}};
constexpr auto SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
// GROUP 2
@@ -30,115 +61,115 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// Pulls from other MODRM table
// GROUP 6
{OPD(TYPE_GROUP_6, PF_NONE, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_66, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_66, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_66, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_66, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_66, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_6, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 7
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_66, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_66, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"SIDT", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"INVLPG", TYPE_SECOND_GROUP_MODRM, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// GROUP 8
{OPD(TYPE_GROUP_8, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_8, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
// GROUP 9
@@ -147,302 +178,302 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// RDRAND/RDSEED only work with no prefix (Other than 66h)
// CMPXCHG8B/16B works with all prefixes
// Tooling fails to decode CMPXCHG with prefix
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 10
{OPD(TYPE_GROUP_10, PF_NONE, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_NONE, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_NONE, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F3, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F3, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_66, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_66, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_10, PF_F2, 0), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 1), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 2), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 3), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 4), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 5), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 6), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_10, PF_F2, 7), 1, X86InstInfo{"UD1", TYPE_INST, FLAGS_BLOCK_END, 0}},
// GROUP 12
{OPD(TYPE_GROUP_12, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_NONE, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_12, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_NONE, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_12, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_NONE, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_12, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_66, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_12, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_66, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_12, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_66, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_12, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_12, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 13
{OPD(TYPE_GROUP_13, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_NONE, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_13, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_NONE, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_13, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_NONE, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_13, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_66, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_13, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_66, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_13, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_66, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_13, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_13, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 14
{OPD(TYPE_GROUP_14, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_NONE, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_14, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_NONE, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1}},
{OPD(TYPE_GROUP_14, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 3), 1, X86InstInfo{"PSRLDQ",TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 7), 1, X86InstInfo{"PSLLDQ",TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_66, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_14, PF_66, 3), 1, X86InstInfo{"PSRLDQ",TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_14, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_66, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_14, PF_66, 7), 1, X86InstInfo{"PSLLDQ",TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 1}},
{OPD(TYPE_GROUP_14, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_14, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 15
{OPD(TYPE_GROUP_15, PF_NONE, 0), 1, X86InstInfo{"FXSAVE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}}, // MMX/x87
{OPD(TYPE_GROUP_15, PF_NONE, 1), 1, X86InstInfo{"FXRSTOR", TYPE_INST, FLAGS_MODRM, 0, nullptr}}, // MMX/x87
{OPD(TYPE_GROUP_15, PF_NONE, 2), 1, X86InstInfo{"LDMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 3), 1, X86InstInfo{"STMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 4), 1, X86InstInfo{"XSAVE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 5), 1, X86InstInfo{"LFENCE/XRSTOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 6), 1, X86InstInfo{"MFENCE/XSAVEOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 7), 1, X86InstInfo{"SFENCE/CLFLUSH", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 0), 1, X86InstInfo{"FXSAVE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}}, // MMX/x87
{OPD(TYPE_GROUP_15, PF_NONE, 1), 1, X86InstInfo{"FXRSTOR", TYPE_INST, FLAGS_MODRM, 0}}, // MMX/x87
{OPD(TYPE_GROUP_15, PF_NONE, 2), 1, X86InstInfo{"LDMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_15, PF_NONE, 3), 1, X86InstInfo{"STMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_15, PF_NONE, 4), 1, X86InstInfo{"XSAVE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_NONE, 5), 1, X86InstInfo{"LFENCE/XRSTOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_NONE, 6), 1, X86InstInfo{"MFENCE/XSAVEOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_NONE, 7), 1, X86InstInfo{"SFENCE/CLFLUSH", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_F3, 0), 1, X86InstInfo{"RDFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 1), 1, X86InstInfo{"RDGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"WRFSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 5), 1, X86InstInfo{"INCSSPQ", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"CLRSSBSY", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 0), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = SecondGroup_ArchSelect_LUT[ENTRY_15_F3_0] }}},
{OPD(TYPE_GROUP_15, PF_F3, 1), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = SecondGroup_ArchSelect_LUT[ENTRY_15_F3_1] }}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = SecondGroup_ArchSelect_LUT[ENTRY_15_F3_2] }}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = SecondGroup_ArchSelect_LUT[ENTRY_15_F3_3] }}},
{OPD(TYPE_GROUP_15, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F3, 5), 1, X86InstInfo{"INCSSPQ", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"UMONITOR/CLRSSBSY", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 6), 1, X86InstInfo{"CLWB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 7), 1, X86InstInfo{"CLFLUSHOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_66, 6), 1, X86InstInfo{"CLWB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_66, 7), 1, X86InstInfo{"CLFLUSHOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 6), 1, X86InstInfo{"UMWAIT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_15, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 16
// AMD documentation claims again that this entire group is n/a to prefix
// Tooling once again fails to disassemble oens with the prefix. Disable until proven otherwise
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
// GROUP 17
{OPD(TYPE_GROUP_17, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 0), 1, X86InstInfo{"EXTRQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 2, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_66, 0), 1, X86InstInfo{"EXTRQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 2}},
{OPD(TYPE_GROUP_17, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_17, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_17, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP P
// AMD documentation claims n/a for all instructions in Group P
@@ -450,54 +481,48 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// It claims that /3 is still Prefetch Mod
// Tooling fails to decode past the /2 encoding but runs fine in hardware
// Hardware also runs all the prefixes correctly
{OPD(TYPE_GROUP_P, PF_NONE, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_NONE, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_NONE, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F3, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F3, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_66, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_66, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_P, PF_F2, 0), 1, X86InstInfo{"PREFETCH Ex", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 1), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 2), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 3), 1, X86InstInfo{"PREFETCH Mod", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 4), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 5), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 6), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_P, PF_F2, 7), 1, X86InstInfo{"PREFETCH Res", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
};
#undef OPD
GenerateTable(&Table.at(0), SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
GenerateTable(Table.data(), SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
IR::InstallToTable(Table, IR::OpDispatch_SecondaryGroupTables);
return Table;
}();
void InitializeSecondaryGroupTables(Context::OperatingMode Mode) {
if (Mode == Context::MODE_64BIT) {
IR::InstallToTable(SecondInstGroupOps, IR::OpDispatch_SecondaryGroupTables_64);
}
}
}
@@ -12,51 +12,51 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps = []() consteval {
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
// REG /1
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_INST, FLAGS_NONE, 0}},
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_INST, FLAGS_NONE, 0}},
{((0 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((0 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((0 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((0 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((0 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((0 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// REG /2
{((1 << 3) | 0), 1, X86InstInfo{"XGETBV", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 1), 1, X86InstInfo{"XSETBV", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 5), 1, X86InstInfo{"XEND", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 6), 1, X86InstInfo{"XTEST", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 0), 1, X86InstInfo{"XGETBV", TYPE_INST, FLAGS_NONE, 0}},
{((1 << 3) | 1), 1, X86InstInfo{"XSETBV", TYPE_PRIV, FLAGS_NONE, 0}},
{((1 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((1 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((1 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((1 << 3) | 5), 1, X86InstInfo{"XEND", TYPE_INST, FLAGS_NONE, 0}},
{((1 << 3) | 6), 1, X86InstInfo{"XTEST", TYPE_INST, FLAGS_NONE, 0}},
{((1 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// REG /3
{((2 << 3) | 0), 1, X86InstInfo{"VMRUN", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 1), 1, X86InstInfo{"VMMCALL", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 2), 1, X86InstInfo{"VMLOAD", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 3), 1, X86InstInfo{"VMSAVE", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 4), 1, X86InstInfo{"STGI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 5), 1, X86InstInfo{"CLGI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 6), 1, X86InstInfo{"SKINIT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 7), 1, X86InstInfo{"INVLPGA", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((2 << 3) | 0), 1, X86InstInfo{"VMRUN", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 1), 1, X86InstInfo{"VMMCALL", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 2), 1, X86InstInfo{"VMLOAD", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 3), 1, X86InstInfo{"VMSAVE", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 4), 1, X86InstInfo{"STGI", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 5), 1, X86InstInfo{"CLGI", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 6), 1, X86InstInfo{"SKINIT", TYPE_PRIV, FLAGS_NONE, 0}},
{((2 << 3) | 7), 1, X86InstInfo{"INVLPGA", TYPE_INST, FLAGS_NONE, 0}},
// REG /7
{((3 << 3) | 0), 1, X86InstInfo{"SWAPGS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 1), 1, X86InstInfo{"RDTSCP", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 2), 1, X86InstInfo{"MONITORX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 3), 1, X86InstInfo{"MWAITX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 4), 1, X86InstInfo{"CLZERO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_SRC_RAX | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{((3 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 0), 1, X86InstInfo{"SWAPGS", TYPE_INST, FLAGS_NONE, 0}},
{((3 << 3) | 1), 1, X86InstInfo{"RDTSCP", TYPE_INST, FLAGS_NONE, 0}},
{((3 << 3) | 2), 1, X86InstInfo{"MONITORX", TYPE_PRIV, FLAGS_NONE, 0}},
{((3 << 3) | 3), 1, X86InstInfo{"MWAITX", TYPE_PRIV, FLAGS_NONE, 0}},
{((3 << 3) | 4), 1, X86InstInfo{"CLZERO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_SRC_RAX | FLAGS_DEBUG_MEM_ACCESS, 0}},
{((3 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((3 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
};
GenerateTable(&Table.at(0), SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
GenerateTable(Table.data(), SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
IR::InstallToTable(Table, IR::OpDispatch_SecondaryModRMTables);
return Table;
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -6,18 +6,20 @@ $end_info$
*/
#pragma once
#include <FEXCore/Core/Context.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <cstddef>
#include <cstdint>
#include <cstring>
#include <type_traits>
namespace FEXCore::IR {
///< Forward declaration of OpDispatchBuilder
class OpDispatchBuilder;
}
namespace FEXCore::X86Tables {
///< Forward declaration of X86InstInfo
struct X86InstInfo;
namespace DecodeFlags {
@@ -32,13 +34,15 @@ constexpr uint32_t FLAG_REX_WIDENING = (1 << 7);
constexpr uint32_t FLAG_REX_XGPR_B = (1 << 8);
constexpr uint32_t FLAG_REX_XGPR_X = (1 << 9);
constexpr uint32_t FLAG_REX_XGPR_R = (1 << 10);
constexpr uint32_t FLAG_ES_PREFIX = (1 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (1 << 12);
constexpr uint32_t FLAG_SS_PREFIX = (1 << 13);
constexpr uint32_t FLAG_DS_PREFIX = (1 << 14);
constexpr uint32_t FLAG_FS_PREFIX = (1 << 15);
constexpr uint32_t FLAG_GS_PREFIX = (1 << 16);
constexpr uint32_t FLAG_SEGMENTS = (0b11'1111 << 11);
constexpr uint32_t FLAG_NO_PREFIX = (0b000 << 11);
constexpr uint32_t FLAG_ES_PREFIX = (0b001 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (0b010 << 11);
constexpr uint32_t FLAG_SS_PREFIX = (0b011 << 11);
constexpr uint32_t FLAG_DS_PREFIX = (0b100 << 11);
constexpr uint32_t FLAG_FS_PREFIX = (0b101 << 11);
constexpr uint32_t FLAG_GS_PREFIX = (0b110 << 11);
constexpr uint32_t FLAG_SEGMENTS = (0b111 << 11);
// Bits 14, 15, 16 - Unused
constexpr uint32_t FLAG_REP_PREFIX = (1 << 17);
constexpr uint32_t FLAG_REPNE_PREFIX = (1 << 18);
@@ -189,6 +193,7 @@ struct DecodedInst {
X86InstInfo const* TableInfo;
uint32_t Flags;
uint8_t OPRaw;
uint16_t OP;
uint8_t ModRM;
@@ -197,6 +202,7 @@ struct DecodedInst {
uint8_t LastEscapePrefix;
bool DecodedModRM;
bool DecodedSIB;
bool ForceTSO;
};
union ModRMDecoded {
@@ -229,6 +235,9 @@ enum InstType {
TYPE_X87 = TYPE_INST,
TYPE_INVALID,
TYPE_COPY_OTHER,
// Changes `X86InstInfo::OpcodeDispatcher` member to use the `Indirect` version.
// Points to a 2 member array of X86InstInfo to choose instruction description based on executing bitness.
TYPE_ARCH_DISPATCHER,
// Must be in order
// Groups 1, 1a, 2, 3, 4, 5, 11 are for the primary op table
@@ -353,6 +362,14 @@ constexpr InstFlagType FLAGS_VEX_1ST_SRC = (0b10ULL << 22);
constexpr InstFlagType FLAGS_VEX_2ND_SRC = (0b11ULL << 22);
// Whether or not the instruction has a VSIB byte
constexpr InstFlagType FLAGS_VEX_VSIB = (1ULL << 24);
constexpr InstFlagType FLAGS_VEX_L_IGNORE = (1ULL << 25);
constexpr InstFlagType FLAGS_VEX_L_0 = (1ULL << 26);
constexpr InstFlagType FLAGS_VEX_L_1 = (1ULL << 27);
constexpr InstFlagType FLAGS_REX_W_0 = (1ULL << 28);
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
@@ -420,12 +437,17 @@ constexpr uint8_t OpToIndex(uint8_t Op) {
using DecodedOp = DecodedInst const*;
using OpDispatchPtr = void (IR::OpDispatchBuilder::*)(DecodedOp);
union OpDispatchPtrWrapper {
OpDispatchPtr OpDispatch;
const struct X86InstInfo *Indirect;
};
struct X86InstInfo {
char const *Name;
InstType Type;
InstFlags::InstFlagType Flags; ///< Must be larger than InstFlags enum
uint8_t MoreBytes;
OpDispatchPtr OpcodeDispatcher;
OpDispatchPtrWrapper OpcodeDispatcher;
bool operator==(const X86InstInfo &b) const {
if (strcmp(Name, b.Name) != 0 ||
@@ -467,23 +489,27 @@ constexpr size_t MAX_VEX_TABLE_SIZE = (1 << 13);
// group select (3 bits for now) | ModRM opcode (3 bits)
constexpr size_t MAX_VEX_GROUP_TABLE_SIZE = (1 << 7);
extern std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps;
extern std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps;
extern std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps;
extern std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps;
extern std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps;
extern const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps;
extern const std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps;
extern const std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps;
extern const std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps;
extern const std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps;
extern std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps;
extern std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps;
extern std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps;
extern std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops;
extern std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps;
extern std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps;
extern std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps;
extern const std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps;
extern const std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps;
extern const std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps;
extern const std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87F80Ops;
extern const std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87F64Ops;
extern const std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps;
extern const std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps;
extern const std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps;
// VEX
extern std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps;
extern std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps;
extern const std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps;
extern const std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps;
extern const std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps_AVX128;
extern const std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps_AVX128;
template <typename OpcodeType>
struct X86TablesInfoStruct {
@@ -504,7 +530,7 @@ constexpr static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInf
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
LOGMAN_MSG_A_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
}
if (FinalTable[OpNum + i].OpcodeDispatcher) {
if (FinalTable[OpNum + i].OpcodeDispatcher.OpDispatch) {
LOGMAN_MSG_A_FMT("Already installed an OpcodeDispatcher for 0x{:x}", OpNum + i);
}
FinalTable[OpNum + i] = Info;
@@ -513,7 +539,7 @@ constexpr static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInf
};
template<typename OpcodeType>
constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, X86InstInfo *OtherLocal) {
constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, const X86InstInfo *OtherLocal) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
@@ -577,8 +603,6 @@ constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86Tables
}
};
void InitializeInfoTables(Context::OperatingMode Mode);
}
template <>
@@ -6,234 +6,760 @@ $end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <iterator>
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops = []() consteval {
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> Table{};
using namespace IR;
// Top bit indicating if it needs to be repeated with {0x40, 0x80} or'd in
// All OPDReg versions need it
#define OPDReg(op, reg) ((1 << 15) | ((op - 0xD8) << 8) | (reg << 3))
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i32Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i32Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i32Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i32Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i32Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i32Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xD0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xD8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD9, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64, OpSize::i32Bit>},
// 1 = Invalid
{OPDReg(0xD9, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i32Bit>},
{OPDReg(0xD9, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i32Bit>},
{OPDReg(0xD9, 4) | 0x00, 8, &OpDispatchBuilder::X87LDENVF64},
{OPDReg(0xD9, 5) | 0x00, 8, &OpDispatchBuilder::X87FLDCWF64},
{OPDReg(0xD9, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSTENV},
{OPDReg(0xD9, 7) | 0x00, 8, &OpDispatchBuilder::X87FSTCW},
{OPD(0xD9, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDFromStack>},
{OPD(0xD9, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xD9, 0xD0), 1, &OpDispatchBuilder::NOPOp}, // FNOP
// D1 = Invalid
// D8 = Invalid
{OPD(0xD9, 0xE0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80STACKCHANGESIGN, false>},
{OPD(0xD9, 0xE1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80STACKABS, false>},
// E2 = Invalid
{OPD(0xD9, 0xE4), 1, &OpDispatchBuilder::FTSTF64},
{OPD(0xD9, 0xE5), 1, &OpDispatchBuilder::X87FXAM},
// E6 = Invalid
{OPD(0xD9, 0xE8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x3FF0000000000000>}, // 1.0
{OPD(0xD9, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x400A934F0979A372>}, // log2l(10)
{OPD(0xD9, 0xEA), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x3FF71547652B82FE>}, // log2l(e)
{OPD(0xD9, 0xEB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x400921FB54442D18>}, // pi
{OPD(0xD9, 0xEC), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x3FD34413509F79FF>}, // log10l(2)
{OPD(0xD9, 0xED), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0x3FE62E42FEFA39EF>}, // log(2)
{OPD(0xD9, 0xEE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64_Const, 0>}, // 0.0
// EF = Invalid
{OPD(0xD9, 0xF0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80F2XM1STACK, false>},
{OPD(0xD9, 0xF1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87FYL2X, false>},
{OPD(0xD9, 0xF2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80PTANSTACK, true>},
{OPD(0xD9, 0xF3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80ATANSTACK, false>},
{OPD(0xD9, 0xF4), 1, &OpDispatchBuilder::X87FXTRACTF64},
{OPD(0xD9, 0xF5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80FPREM1STACK, true>},
{OPD(0xD9, 0xF6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, false>},
{OPD(0xD9, 0xF7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xD9, 0xF8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80FPREMSTACK, true>},
{OPD(0xD9, 0xF9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87FYL2X, true>},
{OPD(0xD9, 0xFA), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SQRTSTACK, false>},
{OPD(0xD9, 0xFB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SINCOSSTACK, true>},
{OPD(0xD9, 0xFC), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80ROUNDSTACK, false>},
{OPD(0xD9, 0xFD), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SCALESTACK, false>},
{OPD(0xD9, 0xFE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SINSTACK, true>},
{OPD(0xD9, 0xFF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80COSSTACK, true>},
{OPDReg(0xDA, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::i32Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::i32Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i32Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i32Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i32Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i32Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i32Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i32Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDA, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
// E8 = Invalid
{OPD(0xDA, 0xE9), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
// EA = Invalid
// F0 = Invalid
// F8 = Invalid
{OPDReg(0xDB, 0) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDB, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, true>},
{OPDReg(0xDB, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, false>},
{OPDReg(0xDB, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, false>},
// 4 = Invalid
{OPDReg(0xDB, 5) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64, OpSize::f80Bit>},
// 6 = Invalid
{OPDReg(0xDB, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::f80Bit>},
{OPD(0xDB, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
{OPD(0xDB, 0xE2), 1, &OpDispatchBuilder::FNCLEX},
{OPD(0xDB, 0xE3), 1, &OpDispatchBuilder::FNINIT},
// E4 = Invalid
{OPD(0xDB, 0xE8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDB, 0xF0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
// F8 = Invalid
{OPDReg(0xDC, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::i64Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::i64Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i64Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i64Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i64Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i64Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i64Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i64Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDC, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDD, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDF64, OpSize::i64Bit>},
{OPDReg(0xDD, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, true>},
{OPDReg(0xDD, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i64Bit>},
{OPDReg(0xDD, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i64Bit>},
{OPDReg(0xDD, 4) | 0x00, 8, &OpDispatchBuilder::X87FRSTOR},
// 5 = Invalid
{OPDReg(0xDD, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSAVE},
{OPDReg(0xDD, 7) | 0x00, 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDD, 0xC0), 8, &OpDispatchBuilder::X87FFREE},
{OPD(0xDD, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>}, // register-register from regular X87
{OPD(0xDD, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>}, //^
{OPD(0xDD, 0xE0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDD, 0xE8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::i16Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::i16Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i16Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::i16Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i16Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::i16Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i16Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::i16Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDE, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xD9), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
{OPD(0xDE, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDF, 0) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDF, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, true>},
{OPDReg(0xDF, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, false>},
{OPDReg(0xDF, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, false>},
{OPDReg(0xDF, 4) | 0x00, 8, &OpDispatchBuilder::FBLDF64},
{OPDReg(0xDF, 5) | 0x00, 8, &OpDispatchBuilder::FILDF64},
{OPDReg(0xDF, 6) | 0x00, 8, &OpDispatchBuilder::FBSTPF64},
{OPDReg(0xDF, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FISTF64, false>},
// XXX: This should also set the x87 tag bits to empty
// We don't support this currently, so just pop the stack
{OPD(0xDF, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xDF, 0xE0), 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDF, 0xE8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDF, 0xF0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
}};
constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i32Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i32Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xD8, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i32Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i32Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i32Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i32Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xD8, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xD8, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD9, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD, OpSize::i32Bit>},
// 1 = Invalid
{OPDReg(0xD9, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i32Bit>},
{OPDReg(0xD9, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i32Bit>},
{OPDReg(0xD9, 4) | 0x00, 8, &OpDispatchBuilder::X87LDENV},
{OPDReg(0xD9, 5) | 0x00, 8, &OpDispatchBuilder::X87FLDCW}, // XXX: stubbed FLDCW
{OPDReg(0xD9, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSTENV},
{OPDReg(0xD9, 7) | 0x00, 8, &OpDispatchBuilder::X87FSTCW},
{OPD(0xD9, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLDFromStack>},
{OPD(0xD9, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xD9, 0xD0), 1, &OpDispatchBuilder::NOPOp}, // FNOP
// D1 = Invalid
// D8 = Invalid
{OPD(0xD9, 0xE0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80STACKCHANGESIGN, false>},
{OPD(0xD9, 0xE1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80STACKABS, false>},
// E2 = Invalid
{OPD(0xD9, 0xE4), 1, &OpDispatchBuilder::FTST},
{OPD(0xD9, 0xE5), 1, &OpDispatchBuilder::X87FXAM},
// E6 = Invalid
{OPD(0xD9, 0xE8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_ONE>}, // 1.0
{OPD(0xD9, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_LOG2_10>}, // log2l(10)
{OPD(0xD9, 0xEA), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_LOG2_E>}, // log2l(e)
{OPD(0xD9, 0xEB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_PI>}, // pi
{OPD(0xD9, 0xEC), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_LOG10_2>}, // log10l(2)
{OPD(0xD9, 0xED), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_X87_LOG_2>}, // log(2)
{OPD(0xD9, 0xEE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD_Const, NamedVectorConstant::NAMED_VECTOR_ZERO>}, // 0.0
// EF = Invalid
{OPD(0xD9, 0xF0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80F2XM1STACK, false>},
{OPD(0xD9, 0xF1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87FYL2X, false>},
{OPD(0xD9, 0xF2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80PTANSTACK, true>},
{OPD(0xD9, 0xF3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80ATANSTACK, false>},
{OPD(0xD9, 0xF4), 1, &OpDispatchBuilder::X87FXTRACT},
{OPD(0xD9, 0xF5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80FPREM1STACK, true>},
{OPD(0xD9, 0xF6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, false>},
{OPD(0xD9, 0xF7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xD9, 0xF8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80FPREMSTACK, true>},
{OPD(0xD9, 0xF9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87FYL2X, true>},
{OPD(0xD9, 0xFA), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SQRTSTACK, false>},
{OPD(0xD9, 0xFB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SINCOSSTACK, true>},
{OPD(0xD9, 0xFC), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80ROUNDSTACK, false>},
{OPD(0xD9, 0xFD), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SCALESTACK, false>},
{OPD(0xD9, 0xFE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80SINSTACK, true>},
{OPD(0xD9, 0xFF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87OpHelper, OP_F80COSSTACK, true>},
{OPDReg(0xDA, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::i32Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::i32Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i32Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i32Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDA, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i32Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 5) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i32Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i32Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDA, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i32Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDA, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDA, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
// E8 = Invalid
{OPD(0xDA, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
// EA = Invalid
// F0 = Invalid
// F8 = Invalid
{OPDReg(0xDB, 0) | 0x00, 8, &OpDispatchBuilder::FILD},
{OPDReg(0xDB, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, true>},
{OPDReg(0xDB, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, false>},
{OPDReg(0xDB, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, false>},
// 4 = Invalid
{OPDReg(0xDB, 5) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD, OpSize::f80Bit>},
// 6 = Invalid
{OPDReg(0xDB, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::f80Bit>},
{OPD(0xDB, 0xC0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xC8), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD0), 8, &OpDispatchBuilder::X87FCMOV},
{OPD(0xDB, 0xD8), 8, &OpDispatchBuilder::X87FCMOV},
// E0 = Invalid
{OPD(0xDB, 0xE2), 1, &OpDispatchBuilder::FNCLEX},
{OPD(0xDB, 0xE3), 1, &OpDispatchBuilder::FNINIT},
// E4 = Invalid
{OPD(0xDB, 0xE8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDB, 0xF0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
// F8 = Invalid
{OPDReg(0xDC, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::i64Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::i64Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i64Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i64Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDC, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i64Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 5) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i64Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i64Bit, false, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDC, 7) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i64Bit, false, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDC, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDD, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FLD, OpSize::i64Bit>},
{OPDReg(0xDD, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, true>},
{OPDReg(0xDD, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i64Bit>},
{OPDReg(0xDD, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FST, OpSize::i64Bit>},
{OPDReg(0xDD, 4) | 0x00, 8, &OpDispatchBuilder::X87FRSTOR},
// 5 = Invalid
{OPDReg(0xDD, 6) | 0x00, 8, &OpDispatchBuilder::X87FNSAVE},
{OPDReg(0xDD, 7) | 0x00, 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDD, 0xC0), 8, &OpDispatchBuilder::X87FFREE},
{OPD(0xDD, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDD, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDD, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDD, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::i16Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::i16Bit, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 2) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i16Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 3) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::i16Bit, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPDReg(0xDE, 4) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i16Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 5) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::i16Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 6) | 0x00, 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i16Bit, true, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xDE, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::i16Bit, true, true, OpDispatchBuilder::OpResult::RES_ST0>},
{OPD(0xDE, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xD9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
{OPD(0xDE, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xF8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPDReg(0xDF, 0) | 0x00, 8, &OpDispatchBuilder::FILD},
{OPDReg(0xDF, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, true>},
{OPDReg(0xDF, 2) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, false>},
{OPDReg(0xDF, 3) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, false>},
{OPDReg(0xDF, 4) | 0x00, 8, &OpDispatchBuilder::FBLD},
{OPDReg(0xDF, 5) | 0x00, 8, &OpDispatchBuilder::FILD},
{OPDReg(0xDF, 6) | 0x00, 8, &OpDispatchBuilder::FBSTP},
{OPDReg(0xDF, 7) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FIST, false>},
// XXX: This should also set the x87 tag bits to empty
// We don't support this currently, so just pop the stack
{OPD(0xDF, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xDF, 0xE0), 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDF, 0xE8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
{OPD(0xDF, 0xF0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
}};
#undef OPD
#undef OPDReg
auto GenerateX87TableLambda = [](const auto DispatchTable) consteval {
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
#define OPDReg(op, reg) (((op - 0xD8) << 8) | (reg << 3))
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct X87OpTable[] = {
// 0xD8
{OPDReg(0xD8, 0), 1, X86InstInfo{"FADD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 1), 1, X86InstInfo{"FMUL", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 2), 1, X86InstInfo{"FCOM", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 3), 1, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_MODRM | FLAGS_POP, 0, nullptr}},
{OPDReg(0xD8, 4), 1, X86InstInfo{"FSUB", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 5), 1, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 6), 1, X86InstInfo{"FDIV", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 7), 1, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 0), 1, X86InstInfo{"FADD", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 1), 1, X86InstInfo{"FMUL", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 2), 1, X86InstInfo{"FCOM", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 3), 1, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_MODRM | FLAGS_POP, 0}},
{OPDReg(0xD8, 4), 1, X86InstInfo{"FSUB", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 5), 1, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 6), 1, X86InstInfo{"FDIV", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD8, 7), 1, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_MODRM, 0}},
// / 0
{OPD(0xD8, 0xC0), 8, X86InstInfo{"FADD", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xC0), 8, X86InstInfo{"FADD", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xD8, 0xC8), 8, X86InstInfo{"FMUL", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xC8), 8, X86InstInfo{"FMUL", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xD8, 0xD0), 8, X86InstInfo{"FCOM", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xD0), 8, X86InstInfo{"FCOM", TYPE_X87, FLAGS_NONE, 0}},
// / 3
{OPD(0xD8, 0xD8), 8, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xD8, 0xD8), 8, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_POP, 0}},
// / 4
{OPD(0xD8, 0xE0), 8, X86InstInfo{"FSUB", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xE0), 8, X86InstInfo{"FSUB", TYPE_X87, FLAGS_NONE, 0}},
// / 5
{OPD(0xD8, 0xE8), 8, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xE8), 8, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_NONE, 0}},
// / 6
{OPD(0xD8, 0xF0), 8, X86InstInfo{"FDIV", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xF0), 8, X86InstInfo{"FDIV", TYPE_X87, FLAGS_NONE, 0}},
// / 7
{OPD(0xD8, 0xF8), 8, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD8, 0xF8), 8, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_NONE, 0}},
// 0xD9
{OPDReg(0xD9, 0), 1, X86InstInfo{"FLD", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD9, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPDReg(0xD9, 2), 1, X86InstInfo{"FST", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xD9, 3), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xD9, 4), 1, X86InstInfo{"FLDENV", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD9, 5), 1, X86InstInfo{"FLDCW", TYPE_X87, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD9, 6), 1, X86InstInfo{"FNSTENV", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xD9, 7), 1, X86InstInfo{"FNSTCW", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xD9, 0), 1, X86InstInfo{"FLD", TYPE_INST, FLAGS_MODRM, 0}},
{OPDReg(0xD9, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPDReg(0xD9, 2), 1, X86InstInfo{"FST", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xD9, 3), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xD9, 4), 1, X86InstInfo{"FLDENV", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xD9, 5), 1, X86InstInfo{"FLDCW", TYPE_X87, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xD9, 6), 1, X86InstInfo{"FNSTENV", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xD9, 7), 1, X86InstInfo{"FNSTCW", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// / 0
{OPD(0xD9, 0xC0), 8, X86InstInfo{"FLD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xC0), 8, X86InstInfo{"FLD", TYPE_INST, FLAGS_NONE, 0}},
// / 1
{OPD(0xD9, 0xC8), 8, X86InstInfo{"FXCH", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xC8), 8, X86InstInfo{"FXCH", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xD9, 0xD0), 1, X86InstInfo{"FNOP", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xD1), 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xD0), 1, X86InstInfo{"FNOP", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xD1), 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 3
{OPD(0xD9, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 4
{OPD(0xD9, 0xE0), 1, X86InstInfo{"FCHS", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE1), 1, X86InstInfo{"FABS", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE2), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE4), 1, X86InstInfo{"FTST", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE5), 1, X86InstInfo{"FXAM", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE6), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE0), 1, X86InstInfo{"FCHS", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE1), 1, X86InstInfo{"FABS", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE2), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE4), 1, X86InstInfo{"FTST", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE5), 1, X86InstInfo{"FXAM", TYPE_INST, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE6), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 5
{OPD(0xD9, 0xE8), 1, X86InstInfo{"FLD1", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE9), 1, X86InstInfo{"FLDL2T", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xEA), 1, X86InstInfo{"FLDL2E", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xEB), 1, X86InstInfo{"FLDPI", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xEC), 1, X86InstInfo{"FLDLG2", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xED), 1, X86InstInfo{"FLDLN2", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xEE), 1, X86InstInfo{"FLDZ", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xEF), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xE8), 1, X86InstInfo{"FLD1", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xE9), 1, X86InstInfo{"FLDL2T", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xEA), 1, X86InstInfo{"FLDL2E", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xEB), 1, X86InstInfo{"FLDPI", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xEC), 1, X86InstInfo{"FLDLG2", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xED), 1, X86InstInfo{"FLDLN2", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xEE), 1, X86InstInfo{"FLDZ", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xEF), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 6
{OPD(0xD9, 0xF0), 1, X86InstInfo{"F2XM1", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF1), 1, X86InstInfo{"FYL2X", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF2), 1, X86InstInfo{"FPTAN", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF3), 1, X86InstInfo{"FPATAN", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF4), 1, X86InstInfo{"FXTRACT", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF5), 1, X86InstInfo{"FPREM1", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF6), 1, X86InstInfo{"FDECSTP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xD9, 0xF7), 1, X86InstInfo{"FINCSTP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xD9, 0xF0), 1, X86InstInfo{"F2XM1", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF1), 1, X86InstInfo{"FYL2X", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF2), 1, X86InstInfo{"FPTAN", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF3), 1, X86InstInfo{"FPATAN", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF4), 1, X86InstInfo{"FXTRACT", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF5), 1, X86InstInfo{"FPREM1", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF6), 1, X86InstInfo{"FDECSTP", TYPE_X87, FLAGS_POP, 0}},
{OPD(0xD9, 0xF7), 1, X86InstInfo{"FINCSTP", TYPE_X87, FLAGS_POP, 0}},
// / 7
{OPD(0xD9, 0xF8), 1, X86InstInfo{"FPREM", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF9), 1, X86InstInfo{"FYL2XP1", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFA), 1, X86InstInfo{"FSQRT", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFB), 1, X86InstInfo{"FSINCOS", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFC), 1, X86InstInfo{"FRNDINT", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFD), 1, X86InstInfo{"FSCALE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFE), 1, X86InstInfo{"FSIN", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xFF), 1, X86InstInfo{"FCOS", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xD9, 0xF8), 1, X86InstInfo{"FPREM", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xF9), 1, X86InstInfo{"FYL2XP1", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFA), 1, X86InstInfo{"FSQRT", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFB), 1, X86InstInfo{"FSINCOS", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFC), 1, X86InstInfo{"FRNDINT", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFD), 1, X86InstInfo{"FSCALE", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFE), 1, X86InstInfo{"FSIN", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xD9, 0xFF), 1, X86InstInfo{"FCOS", TYPE_X87, FLAGS_NONE, 0}},
// 0xDA
{OPDReg(0xDA, 0), 1, X86InstInfo{"FIADD", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 1), 1, X86InstInfo{"FIMUL", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 2), 1, X86InstInfo{"FICOM", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 3), 1, X86InstInfo{"FICOMP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDA, 4), 1, X86InstInfo{"FISUB", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 5), 1, X86InstInfo{"FISUBR", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 6), 1, X86InstInfo{"FIDIV", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 7), 1, X86InstInfo{"FIDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDA, 0), 1, X86InstInfo{"FIADD", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 1), 1, X86InstInfo{"FIMUL", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 2), 1, X86InstInfo{"FICOM", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 3), 1, X86InstInfo{"FICOMP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_POP, 0}},
{OPDReg(0xDA, 4), 1, X86InstInfo{"FISUB", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 5), 1, X86InstInfo{"FISUBR", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 6), 1, X86InstInfo{"FIDIV", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDA, 7), 1, X86InstInfo{"FIDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
// / 0
{OPD(0xDA, 0xC0), 8, X86InstInfo{"FCMOVB", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xC0), 8, X86InstInfo{"FCMOVB", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xDA, 0xC8), 8, X86InstInfo{"FCMOVE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xC8), 8, X86InstInfo{"FCMOVE", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDA, 0xD0), 8, X86InstInfo{"FCMOVBE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xD0), 8, X86InstInfo{"FCMOVBE", TYPE_X87, FLAGS_NONE, 0}},
// / 3
{OPD(0xDA, 0xD8), 8, X86InstInfo{"FCMOVU", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xD8), 8, X86InstInfo{"FCMOVU", TYPE_X87, FLAGS_NONE, 0}},
// / 4
{OPD(0xDA, 0xE0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xE0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 5
{OPD(0xDA, 0xE8), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xE9), 1, X86InstInfo{"FUCOMPP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDA, 0xEA), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xE8), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDA, 0xE9), 1, X86InstInfo{"FUCOMPP", TYPE_X87, FLAGS_POP, 0}},
{OPD(0xDA, 0xEA), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 6
{OPD(0xDA, 0xF0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xF0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 7
{OPD(0xDA, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDA, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// 0xDB
{OPDReg(0xDB, 0), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDB, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDB, 2), 1, X86InstInfo{"FIST", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xDB, 3), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDB, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPDReg(0xDB, 5), 1, X86InstInfo{"FLD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDB, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPDReg(0xDB, 7), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDB, 0), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDB, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDB, 2), 1, X86InstInfo{"FIST", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xDB, 3), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDB, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPDReg(0xDB, 5), 1, X86InstInfo{"FLD", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xDB, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPDReg(0xDB, 7), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
// / 0
{OPD(0xDB, 0xC0), 8, X86InstInfo{"FCMOVNB", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xC0), 8, X86InstInfo{"FCMOVNB", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xDB, 0xC8), 8, X86InstInfo{"FCMOVNE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xC8), 8, X86InstInfo{"FCMOVNE", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDB, 0xD0), 8, X86InstInfo{"FCMOVNBE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xD0), 8, X86InstInfo{"FCMOVNBE", TYPE_X87, FLAGS_NONE, 0}},
// / 3
{OPD(0xDB, 0xD8), 8, X86InstInfo{"FCMOVNU", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xD8), 8, X86InstInfo{"FCMOVNU", TYPE_X87, FLAGS_NONE, 0}},
// / 4
{OPD(0xDB, 0xE0), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xE2), 1, X86InstInfo{"FNCLEX", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xE3), 1, X86InstInfo{"FNINIT", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xE4), 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xE0), 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDB, 0xE2), 1, X86InstInfo{"FNCLEX", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xDB, 0xE3), 1, X86InstInfo{"FNINIT", TYPE_X87, FLAGS_NONE, 0}},
{OPD(0xDB, 0xE4), 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 5
{OPD(0xDB, 0xE8), 8, X86InstInfo{"FUCOMI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xE8), 8, X86InstInfo{"FUCOMI", TYPE_INST, FLAGS_NONE, 0}},
// / 6
{OPD(0xDB, 0xF0), 8, X86InstInfo{"FCOMI", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xF0), 8, X86InstInfo{"FCOMI", TYPE_X87, FLAGS_NONE, 0}},
// / 7
{OPD(0xDB, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDB, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// 0xDC
{OPDReg(0xDC, 0), 1, X86InstInfo{"FADD", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 1), 1, X86InstInfo{"FMUL", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 2), 1, X86InstInfo{"FCOM", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 3), 1, X86InstInfo{"FCOMP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDC, 4), 1, X86InstInfo{"FSUB", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 5), 1, X86InstInfo{"FSUBR", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 6), 1, X86InstInfo{"FDIV", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 7), 1, X86InstInfo{"FDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDC, 0), 1, X86InstInfo{"FADD", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 1), 1, X86InstInfo{"FMUL", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 2), 1, X86InstInfo{"FCOM", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 3), 1, X86InstInfo{"FCOMP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_POP, 0}},
{OPDReg(0xDC, 4), 1, X86InstInfo{"FSUB", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 5), 1, X86InstInfo{"FSUBR", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 6), 1, X86InstInfo{"FDIV", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDC, 7), 1, X86InstInfo{"FDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
// / 0
{OPD(0xDC, 0xC0), 8, X86InstInfo{"FADD", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xC0), 8, X86InstInfo{"FADD", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xDC, 0xC8), 8, X86InstInfo{"FMUL", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xC8), 8, X86InstInfo{"FMUL", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDC, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 3
{OPD(0xDC, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 4
{OPD(0xDC, 0xE0), 8, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xE0), 8, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_NONE, 0}},
// / 5
{OPD(0xDC, 0xE8), 8, X86InstInfo{"FSUB", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xE8), 8, X86InstInfo{"FSUB", TYPE_X87, FLAGS_NONE, 0}},
// / 6
{OPD(0xDC, 0xF0), 8, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xF0), 8, X86InstInfo{"FDIVR", TYPE_X87, FLAGS_NONE, 0}},
// / 7
{OPD(0xDC, 0xF8), 8, X86InstInfo{"FDIV", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDC, 0xF8), 8, X86InstInfo{"FDIV", TYPE_X87, FLAGS_NONE, 0}},
// 0xDD
{OPDReg(0xDD, 0), 1, X86InstInfo{"FLD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDD, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDD, 2), 1, X86InstInfo{"FST", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xDD, 3), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDD, 4), 1, X86InstInfo{"FRSTOR", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDD, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPDReg(0xDD, 6), 1, X86InstInfo{"FNSAVE", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xDD, 7), 1, X86InstInfo{"FNSTSW", TYPE_X87, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xDD, 0), 1, X86InstInfo{"FLD", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xDD, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDD, 2), 1, X86InstInfo{"FST", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xDD, 3), 1, X86InstInfo{"FSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDD, 4), 1, X86InstInfo{"FRSTOR", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xDD, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPDReg(0xDD, 6), 1, X86InstInfo{"FNSAVE", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xDD, 7), 1, X86InstInfo{"FNSTSW", TYPE_X87, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// / 0
{OPD(0xDD, 0xC0), 8, X86InstInfo{"FFREE", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDD, 0xC0), 8, X86InstInfo{"FFREE", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xDD, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDD, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 2
{OPD(0xDD, 0xD0), 8, X86InstInfo{"FST", TYPE_INST, FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(0xDD, 0xD0), 8, X86InstInfo{"FST", TYPE_INST, FLAGS_SF_MOD_DST, 0}},
// / 3
{OPD(0xDD, 0xD8), 8, X86InstInfo{"FSTP", TYPE_X87, FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPD(0xDD, 0xD8), 8, X86InstInfo{"FSTP", TYPE_X87, FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
// / 4
{OPD(0xDD, 0xE0), 8, X86InstInfo{"FUCOM", TYPE_X87, FLAGS_NONE, 0, nullptr}},
{OPD(0xDD, 0xE0), 8, X86InstInfo{"FUCOM", TYPE_X87, FLAGS_NONE, 0}},
// / 5
{OPD(0xDD, 0xE8), 8, X86InstInfo{"FUCOMP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDD, 0xE8), 8, X86InstInfo{"FUCOMP", TYPE_X87, FLAGS_POP, 0}},
// / 6
{OPD(0xDD, 0xF0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDD, 0xF0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 7
{OPD(0xDD, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDD, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// 0xDE
{OPDReg(0xDE, 0), 1, X86InstInfo{"FIADD", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 1), 1, X86InstInfo{"FIMUL", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 2), 1, X86InstInfo{"FICOM", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 3), 1, X86InstInfo{"FICOMP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDE, 4), 1, X86InstInfo{"FISUB", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 5), 1, X86InstInfo{"FISUBR", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 6), 1, X86InstInfo{"FIDIV", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 7), 1, X86InstInfo{"FIDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDE, 0), 1, X86InstInfo{"FIADD", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 1), 1, X86InstInfo{"FIMUL", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 2), 1, X86InstInfo{"FICOM", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 3), 1, X86InstInfo{"FICOMP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_POP, 0}},
{OPDReg(0xDE, 4), 1, X86InstInfo{"FISUB", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 5), 1, X86InstInfo{"FISUBR", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 6), 1, X86InstInfo{"FIDIV", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDE, 7), 1, X86InstInfo{"FIDIVR", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
// / 0
{OPD(0xDE, 0xC0), 8, X86InstInfo{"FADDP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xC0), 8, X86InstInfo{"FADDP", TYPE_X87, FLAGS_POP, 0}},
// / 1
{OPD(0xDE, 0xC8), 8, X86InstInfo{"FMULP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xC8), 8, X86InstInfo{"FMULP", TYPE_X87, FLAGS_POP, 0}},
// / 2
{OPD(0xDE, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDE, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 3
{OPD(0xDE, 0xD8), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDE, 0xD9), 1, X86InstInfo{"FCOMPP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xDA), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDE, 0xD8), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDE, 0xD9), 1, X86InstInfo{"FCOMPP", TYPE_X87, FLAGS_POP, 0}},
{OPD(0xDE, 0xDA), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 4
{OPD(0xDE, 0xE0), 8, X86InstInfo{"FSUBRP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xE0), 8, X86InstInfo{"FSUBRP", TYPE_X87, FLAGS_POP, 0}},
// / 5
{OPD(0xDE, 0xE8), 8, X86InstInfo{"FSUBP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xE8), 8, X86InstInfo{"FSUBP", TYPE_X87, FLAGS_POP, 0}},
// / 6
{OPD(0xDE, 0xF0), 8, X86InstInfo{"FDIVRP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xF0), 8, X86InstInfo{"FDIVRP", TYPE_X87, FLAGS_POP, 0}},
// / 7
{OPD(0xDE, 0xF8), 8, X86InstInfo{"FDIVP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDE, 0xF8), 8, X86InstInfo{"FDIVP", TYPE_X87, FLAGS_POP, 0}},
// 0xDF
{OPDReg(0xDF, 0), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDF, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDF, 2), 1, X86InstInfo{"FIST", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPDReg(0xDF, 3), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDF, 4), 1, X86InstInfo{"FBLD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xDF, 5), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0, nullptr}},
{OPDReg(0xDF, 6), 1, X86InstInfo{"FBSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDF, 7), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_SF_MOD_DST | FLAGS_POP, 0, nullptr}},
{OPDReg(0xDF, 0), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{OPDReg(0xDF, 1), 1, X86InstInfo{"FISTTP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDF, 2), 1, X86InstInfo{"FIST", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPDReg(0xDF, 3), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDF, 4), 1, X86InstInfo{"FBLD", TYPE_X87, FLAGS_MODRM, 0}},
{OPDReg(0xDF, 5), 1, X86InstInfo{"FILD", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS, 0}},
{OPDReg(0xDF, 6), 1, X86InstInfo{"FBSTP", TYPE_X87, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
{OPDReg(0xDF, 7), 1, X86InstInfo{"FISTP", TYPE_X87, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_X87_FLAGS | FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
// / 0
// This instruction is a bit special. This is an undocumented(Almost) x87 instruction.
// https://en.wikipedia.org/wiki/X86_instruction_listings#Undocumented_x87_instructions
@@ -243,28 +769,52 @@ std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops = []() consteval {
// At some point the Nvidia OpenGL binary driver uses this instruction.
// GCC may also end up emitting this instruction in some rare edge case!
// Almost all x86 CPUs implement this, and it is expected to be around
{OPD(0xDF, 0xC0), 8, X86InstInfo{"FFREEP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDF, 0xC0), 8, X86InstInfo{"FFREEP", TYPE_X87, FLAGS_POP, 0}},
// / 1
{OPD(0xDF, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDF, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 2
{OPD(0xDF, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDF, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 3
{OPD(0xDF, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDF, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 4
{OPD(0xDF, 0xE0), 1, X86InstInfo{"FNSTSW", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{OPD(0xDF, 0xE1), 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDF, 0xE0), 1, X86InstInfo{"FNSTSW", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0}},
{OPD(0xDF, 0xE1), 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// / 5
{OPD(0xDF, 0xE8), 8, X86InstInfo{"FUCOMIP", TYPE_INST, FLAGS_POP, 0, nullptr}},
{OPD(0xDF, 0xE8), 8, X86InstInfo{"FUCOMIP", TYPE_INST, FLAGS_POP, 0}},
// / 6
{OPD(0xDF, 0xF0), 8, X86InstInfo{"FCOMIP", TYPE_X87, FLAGS_POP, 0, nullptr}},
{OPD(0xDF, 0xF0), 8, X86InstInfo{"FCOMIP", TYPE_X87, FLAGS_POP, 0}},
// / 7
{OPD(0xDF, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(0xDF, 0xF8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
};
#undef OPD
#undef OPDReg
GenerateX87Table(&Table.at(0), X87OpTable, std::size(X87OpTable));
auto InstallToX87Table = [](auto& FinalTable, auto& LocalTable) {
for (auto Op : LocalTable) {
auto OpNum = Op.Op;
bool Repeat = (OpNum & 0x8000) != 0;
OpNum = OpNum & 0x7FF;
auto Dispatcher = Op.Ptr;
for (uint8_t i = 0; i < Op.Count; ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].OpcodeDispatcher.OpDispatch == nullptr, "Duplicate Entry");
FinalTable[OpNum + i].OpcodeDispatcher.OpDispatch = Dispatcher;
// Flag to indicate if we need to repeat this op in {0x40, 0x80} ranges
if (Repeat) {
FinalTable[(OpNum | 0x40) + i].OpcodeDispatcher.OpDispatch = Dispatcher;
FinalTable[(OpNum | 0x80) + i].OpcodeDispatcher.OpDispatch = Dispatcher;
}
}
}
};
GenerateX87Table(Table.data(), X87OpTable, std::size(X87OpTable));
InstallToX87Table(Table, DispatchTable);
return Table;
}();
};
constexpr std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87F80Ops = GenerateX87TableLambda(X87F80OpTable);
constexpr std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87F64Ops = GenerateX87TableLambda(X87F64OpTable);
}
+2 -3
View File
@@ -90,7 +90,7 @@ void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry* Entry, uintptr_t VAFileStart,
mem += info->nlines * sizeof(gdb_line_mapping);
if (info->nlines) {
memcpy(lines, &Lines.at(0), info->nlines * sizeof(gdb_line_mapping));
memcpy(lines, Lines.data(), info->nlines * sizeof(gdb_line_mapping));
}
auto entry = new jit_code_entry {0, 0, 0, 0};
@@ -113,8 +113,7 @@ void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry* Entry, uintptr_t VAFileStart,
} // namespace FEXCore
#else
namespace FEXCore {
void GDBJITRegister([[maybe_unused]] FEXCore::IR::AOTIRCacheEntry* Entry, [[maybe_unused]] uintptr_t VAFileStart, [[maybe_unused]] uint64_t GuestRIP,
[[maybe_unused]] uintptr_t HostEntry, [[maybe_unused]] FEXCore::Core::DebugData* DebugData) {
void GDBJITRegister(FEXCore::IR::AOTIRCacheEntry*, uintptr_t, uint64_t, uintptr_t, FEXCore::Core::DebugData*) {
ERROR_AND_DIE_FMT("GDBSymbols support not compiled in");
}
} // namespace FEXCore
+2 -348
View File
@@ -2,311 +2,23 @@
#include "FEXHeaderUtils/Filesystem.h"
#include "Interface/Context/Context.h"
#include "Interface/IR/AOTIR.h"
#include "Interface/IR/IntrusiveIRList.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/string.h>
#include <Interface/Core/LookupCache.h>
#include <Interface/GDBJIT/GDBJIT.h>
#include <cstddef>
#include <cstdint>
#include <mutex>
#include <sys/stat.h>
#include <unistd.h>
#include <xxhash.h>
namespace FEXCore::IR {
AOTIRInlineEntry* AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
return (AOTIRInlineEntry*)(This + DataBase + DataOffset);
}
AOTIRInlineEntry* AOTIRInlineIndex::Find(uint64_t GuestStart) {
ssize_t l = 0;
ssize_t r = Count - 1;
while (l <= r) {
size_t m = l + (r - l) / 2;
if (Entries[m].GuestStart == GuestStart) {
return GetInlineEntry(Entries[m].DataOffset);
} else if (Entries[m].GuestStart < GuestStart) {
l = m + 1;
} else {
r = m - 1;
}
}
return nullptr;
}
IR::IRListView* AOTIRInlineEntry::GetIRData() {
return (IR::IRListView*)InlineData;
}
void AOTIRCaptureCacheEntry::AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash,
const FEXCore::IR::IRListView& IRList) {
auto Inserted = Index.emplace(GuestRIP, Stream->Offset());
if (Inserted.second) {
AOTIRInlineEntry entry {
.GuestHash = Hash,
.GuestLength = Length,
};
Stream->Write((const char*)&entry, sizeof(entry));
// IRData (inline)
IRList.Serialize(*Stream);
}
}
static bool readAll(int fd, void* data, size_t size) {
int rv = read(fd, data, size);
if (rv != size) {
return false;
} else {
return true;
}
}
static bool LoadAOTIRCache(AOTIRCacheEntry* Entry, int streamfd) {
#ifndef _WIN32
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != FEXCore::IR::AOTIR_COOKIE) {
return false;
}
fextl::string Module;
uint64_t ModSize;
uint64_t IndexSize;
lseek(streamfd, -sizeof(ModSize), SEEK_END);
if (!readAll(streamfd, (char*)&ModSize, sizeof(ModSize))) {
return false;
}
Module.resize(ModSize);
lseek(streamfd, -sizeof(ModSize) - ModSize, SEEK_END);
if (!readAll(streamfd, (char*)&Module[0], Module.size())) {
return false;
}
if (Entry->FileId != Module) {
return false;
}
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize))) {
return false;
}
struct stat fileinfo;
if (fstat(streamfd, &fileinfo) < 0) {
return false;
}
size_t Size = (fileinfo.st_size + 4095) & ~4095;
size_t IndexOffset = fileinfo.st_size - IndexSize - sizeof(ModSize) - ModSize - sizeof(IndexSize);
void* FilePtr = FEXCore::Allocator::mmap(nullptr, Size, PROT_READ, MAP_SHARED, streamfd, 0);
if (FilePtr == MAP_FAILED) {
return false;
}
auto Array = (AOTIRInlineIndex*)((char*)FilePtr + IndexOffset);
LOGMAN_THROW_A_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
Entry->Array = Array;
Entry->FilePtr = FilePtr;
Entry->Size = Size;
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
#else
return false;
#endif
}
void AOTIRCaptureCache::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
std::unique_lock lk(AOTIRCacheLock);
for (auto& [String, Entry] : AOTIRCaptureCacheMap) {
if (!Entry.Stream) {
continue;
}
const auto ModSize = String.size();
auto& stream = Entry.Stream;
// pad to 32 bytes
constexpr char Zero = 0;
while (stream->Offset() & 31) {
stream->Write(&Zero, 1);
}
AOTIRInlineIndex index {
.Count = Entry.Index.size(),
.DataBase = -stream->Offset(),
};
stream->Write((const char*)&index, sizeof(index));
for (const auto& [GuestStart, DataOffset] : Entry.Index) {
AOTIRInlineIndexEntry entry {
.GuestStart = GuestStart,
.DataOffset = DataOffset,
};
stream->Write((const char*)&entry, sizeof(entry));
}
// End of file header
const auto IndexSize = sizeof(AOTIRInlineIndex) + index.Count * sizeof(FEXCore::IR::AOTIRInlineIndexEntry);
stream->Write((const char*)&IndexSize, sizeof(IndexSize));
stream->Write(String.c_str(), ModSize);
stream->Write((const char*)&ModSize, sizeof(ModSize));
// Close the stream
stream->Close();
// Rename the file to atomically update the cache with the temporary file
AOTIRRenamer(String);
}
}
void AOTIRCaptureCache::AOTIRCaptureCacheWriteoutQueue_Flush() {
{
std::shared_lock lk {AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
for (;;) {
// This code is tricky to refactor so it doesn't allocate memory through glibc.
// The moved std::function object deallocates memory at the end of scope.
FEXCore::Allocator::YesIKnowImNotSupposedToUseTheGlibcAllocator glibc;
AOTIRCaptureCacheWriteoutLock.lock();
WriteOutFn fn = std::move(AOTIRCaptureCacheWriteoutQueue.front());
bool MaybeEmpty = false;
AOTIRCaptureCacheWriteoutQueue.pop();
MaybeEmpty = AOTIRCaptureCacheWriteoutQueue.size() == 0;
AOTIRCaptureCacheWriteoutLock.unlock();
fn();
if (MaybeEmpty) {
std::shared_lock lk {AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
}
LOGMAN_MSG_A_FMT("Must never get here");
}
void AOTIRCaptureCache::AOTIRCaptureCacheWriteoutQueue_Append(const WriteOutFn& fn) {
bool Flush = false;
{
std::unique_lock lk {AOTIRCaptureCacheWriteoutLock};
AOTIRCaptureCacheWriteoutQueue.push(fn);
if (AOTIRCaptureCacheWriteoutQueue.size() > 10000) {
Flush = true;
}
}
bool test_val = false;
if (Flush && AOTIRCaptureCacheWriteoutFlusing.compare_exchange_strong(test_val, true)) {
AOTIRCaptureCacheWriteoutQueue_Flush();
}
}
void AOTIRCaptureCache::WriteFilesWithCode(const Context::AOTIRCodeFileWriterFn& Writer) {
std::shared_lock lk(AOTIRCacheLock);
for (const auto& Entry : AOTIRCache) {
if (Entry.second.ContainsCode) {
Writer(Entry.second.FileId, Entry.second.Filename);
}
}
}
// IRStorageBase with memory owned by IR cache
class IRInlineStorage : public IRStorageBase {
AOTIRInlineEntry& entry;
public:
IRInlineStorage(AOTIRInlineEntry& entry)
: entry(entry) {}
IRListView GetIRView() override {
return entry.GetIRData();
}
};
std::optional<AOTIRCaptureCache::PreGenerateIRFetchResult>
AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
PreGenerateIRFetchResult Result {};
if (AOTIRCacheEntry.Entry) {
AOTIRCacheEntry.Entry->ContainsCode = true;
if (CTX->Config.AOTIRLoad()) {
auto Mod = AOTIRCacheEntry.Entry->Array;
if (Mod != nullptr) {
auto AOTEntry = Mod->Find(GuestRIP - AOTIRCacheEntry.VAFileStart);
if (AOTEntry) {
// verify hash
auto MappedStart = GuestRIP;
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
if (hash == AOTEntry->GuestHash) {
Result.IR = fextl::make_unique<IRInlineStorage>(*AOTEntry);
// LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
Result.DebugData = new FEXCore::Core::DebugData();
Result.StartAddr = MappedStart;
Result.Length = AOTEntry->GuestLength;
return Result;
} else {
LogMan::Msg::IFmt("AOTIR: hash check failed {:x}\n", MappedStart);
}
} else {
// LogMan::Msg::IFmt("AOTIR: Failed to find {:x}, {:x}, {}\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid);
}
}
}
}
return std::nullopt;
}
bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thread, void* CodePtr, uint64_t GuestRIP, uint64_t StartAddr,
uint64_t Length, fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR,
FEXCore::Core::DebugData* DebugData, bool GeneratedIR) {
uint64_t Length, FEXCore::Core::DebugData* DebugData) {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
if (CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
@@ -318,44 +30,7 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
if (CTX->Config.GDBSymbols()) {
GDBJITRegister(AOTIRCacheEntry.Entry, AOTIRCacheEntry.VAFileStart, GuestRIP, (uintptr_t)CodePtr, DebugData);
}
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && (CTX->Config.AOTIRCapture() || CTX->Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - AOTIRCacheEntry.VAFileStart;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.VAFileStart;
const auto& FileId = AOTIRCacheEntry.Entry->FileId;
// The lambda is converted to std::function. This is tricky to refactor so it doesn't allocate memory through glibc.
// NOTE: unique_ptr must be passed as a raw pointer since std::function requires lambda captures to be copyable
FEXCore::Allocator::YesIKnowImNotSupposedToUseTheGlibcAllocator glibc;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRRaw = IR.release(), FileId]() {
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR(IRRaw);
// It is guaranteed via AOTIRCaptureCacheWriteoutLock and AOTIRCaptureCacheWriteoutFlusing that this will not run concurrently
// Memory coherency is guaranteed via AOTIRCaptureCacheWriteoutLock
auto* AotFile = &AOTIRCaptureCacheMap[FileId];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(FileId);
uint64_t tag = FEXCore::IR::AOTIR_COOKIE;
AotFile->Stream->Write(&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IR->GetIRView());
});
if (CTX->Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
Thread->CPUBackend->ClearCache();
return true;
}
}
}
// Insert to caches if we generated IR
}
return false;
@@ -376,31 +51,10 @@ AOTIRCacheEntry* AOTIRCaptureCache::LoadAOTIRCacheEntry(const fextl::string& fil
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry {.FileId = fileid, .Filename = filename}});
auto Entry = &(Inserted.first->second);
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
if (CTX->Config.AOTIRLoad && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
FEXCore::IR::LoadAOTIRCache(Entry, streamfd);
close(streamfd);
}
}
return Entry;
}
return nullptr;
}
void AOTIRCaptureCache::UnloadAOTIRCacheEntry(AOTIRCacheEntry* Entry) {
#ifndef _WIN32
LOGMAN_THROW_A_FMT(Entry != nullptr, "Removing not existing entry");
if (Entry->Array) {
FEXCore::Allocator::munmap(Entry->FilePtr, Entry->Size);
Entry->Array = nullptr;
Entry->FilePtr = nullptr;
Entry->Size = 0;
}
#endif
}
} // namespace FEXCore::IR
+11 -111
View File
@@ -1,27 +1,21 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/IR/RegisterAllocationData.h"
#include "Interface/IR/IntrusiveIRList.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/queue.h>
#include <FEXCore/fextl/unordered_map.h>
#include <atomic>
#include <cstdint>
#include <functional>
#include <memory>
#include <shared_mutex>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <shared_mutex>
namespace FEXCore::CPU {
union Relocation;
} // namespace FEXCore::CPU
namespace FEXCore::Core {
struct InternalThreadState;
struct DebugDataSubblock {
uint32_t HostCodeOffset;
uint32_t HostCodeSize;
@@ -50,123 +44,29 @@ class ContextImpl;
}
namespace FEXCore::IR {
class IRListView;
constexpr auto COOKIE_VERSION = [](const char CookieText[4], uint32_t Version) {
uint64_t Cookie = Version;
Cookie <<= 32;
// Make the cookie text be the lower bits
Cookie |= CookieText[3];
Cookie <<= 8;
Cookie |= CookieText[2];
Cookie <<= 8;
Cookie |= CookieText[1];
Cookie <<= 8;
Cookie |= CookieText[0];
return Cookie;
};
constexpr static uint32_t AOTIR_VERSION = 0x0000'00004;
constexpr static uint64_t AOTIR_COOKIE = COOKIE_VERSION("FEXI", AOTIR_VERSION);
struct AOTIRInlineEntry {
uint64_t GuestHash;
uint64_t GuestLength;
/* IRData */
uint8_t InlineData[0];
IR::IRListView* GetIRData();
};
struct AOTIRInlineIndexEntry {
uint64_t GuestStart;
uint64_t DataOffset;
};
struct AOTIRInlineIndex {
uint64_t Count;
uint64_t DataBase;
AOTIRInlineIndexEntry Entries[0];
AOTIRInlineEntry* Find(uint64_t GuestStart);
AOTIRInlineEntry* GetInlineEntry(uint64_t DataOffset);
};
struct AOTIRCaptureCacheEntry {
fextl::unique_ptr<FEXCore::Context::AOTIRWriter> Stream;
fextl::map<uint64_t, uint64_t> Index;
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, const FEXCore::IR::IRListView& IRList);
};
struct AOTIRCacheEntry {
AOTIRInlineIndex* Array;
void* FilePtr;
size_t Size;
fextl::unique_ptr<FEXCore::HLE::SourcecodeMap> SourcecodeMap;
fextl::string FileId;
fextl::string Filename;
bool ContainsCode;
};
using AOTCacheType = fextl::unordered_map<fextl::string, FEXCore::IR::AOTIRCacheEntry>;
class AOTIRCaptureCache final {
public:
using WriteOutFn = std::function<void()>;
AOTIRCaptureCache(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx} {}
void FinalizeAOTIRCache();
void AOTIRCaptureCacheWriteoutQueue_Flush();
void AOTIRCaptureCacheWriteoutQueue_Append(const WriteOutFn& fn);
void WriteFilesWithCode(const Context::AOTIRCodeFileWriterFn& Writer);
struct PreGenerateIRFetchResult {
fextl::unique_ptr<IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
};
[[nodiscard]]
std::optional<PreGenerateIRFetchResult> PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
bool PostCompileCode(FEXCore::Core::InternalThreadState* Thread, void* CodePtr, uint64_t GuestRIP, uint64_t StartAddr, uint64_t Length,
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR, FEXCore::Core::DebugData* DebugData, bool GeneratedIR);
FEXCore::Core::DebugData* DebugData);
AOTIRCacheEntry* LoadAOTIRCacheEntry(const fextl::string& filename);
void UnloadAOTIRCacheEntry(AOTIRCacheEntry* Entry);
// Callbacks
void SetAOTIRLoader(Context::AOTIRLoaderCBFn CacheReader) {
AOTIRLoader = std::move(CacheReader);
}
void SetAOTIRWriter(Context::AOTIRWriterCBFn CacheWriter) {
AOTIRWriter = std::move(CacheWriter);
}
void SetAOTIRRenamer(Context::AOTIRRenamerCBFn CacheRenamer) {
AOTIRRenamer = std::move(CacheRenamer);
}
private:
FEXCore::Context::ContextImpl* CTX;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
std::atomic<bool> AOTIRCaptureCacheWriteoutFlusing;
fextl::queue<WriteOutFn> AOTIRCaptureCacheWriteoutQueue;
FEXCore::IR::AOTCacheType AOTIRCache;
Context::AOTIRLoaderCBFn AOTIRLoader;
Context::AOTIRWriterCBFn AOTIRWriter;
Context::AOTIRRenamerCBFn AOTIRRenamer;
fextl::unordered_map<fextl::string, FEXCore::IR::AOTIRCaptureCacheEntry> AOTIRCaptureCacheMap;
fextl::unordered_map<fextl::string, FEXCore::IR::AOTIRCacheEntry> AOTIRCache;
};
} // namespace FEXCore::IR
+39 -33
View File
@@ -1,12 +1,15 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/sstream.h>
#include <array>
namespace FEXCore::IR {
class OrderedNode;
@@ -54,23 +57,19 @@ struct NodeID final {
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool
operator<(NodeID lhs, NodeID rhs) noexcept {
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool
operator>(NodeID lhs, NodeID rhs) noexcept {
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool
operator<=(NodeID lhs, NodeID rhs) noexcept {
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool
operator>=(NodeID lhs, NodeID rhs) noexcept {
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
@@ -144,9 +143,22 @@ struct FEX_PACKED NodeWrapperBase final {
return NodeOffset & (1u << 31);
}
[[nodiscard]]
bool HasKill() const {
return NodeOffset & (1u << 30);
}
void ClearKill() {
NodeOffset &= ~(1u << 30);
}
void SetKill() {
NodeOffset |= (1u << 30);
}
[[nodiscard]]
bool IsPointer() const {
return !IsImmediate();
return !IsImmediate() && !HasKill();
}
[[nodiscard]]
@@ -183,8 +195,7 @@ struct FEX_PACKED NodeWrapperBase final {
}
[[nodiscard]]
friend constexpr bool
operator==(const NodeWrapperBase<Type>&, const NodeWrapperBase<Type>&) = default;
friend constexpr bool operator==(const NodeWrapperBase<Type>&, const NodeWrapperBase<Type>&) = default;
[[nodiscard]]
static NodeWrapperBase<Type> FromImmediate(uint32_t Immediate) {
@@ -423,8 +434,7 @@ struct FEX_PACKED RegisterClassType final {
return Val;
}
[[nodiscard]]
friend constexpr bool
operator==(const RegisterClassType&, const RegisterClassType&) = default;
friend constexpr bool operator==(const RegisterClassType&, const RegisterClassType&) = default;
};
struct FEX_PACKED CondClassType final {
@@ -433,8 +443,7 @@ struct FEX_PACKED CondClassType final {
return Val;
}
[[nodiscard]]
friend constexpr bool
operator==(const CondClassType&, const CondClassType&) = default;
friend constexpr bool operator==(const CondClassType&, const CondClassType&) = default;
};
struct FEX_PACKED MemOffsetType final {
@@ -443,8 +452,7 @@ struct FEX_PACKED MemOffsetType final {
return Val;
}
[[nodiscard]]
friend constexpr bool
operator==(const MemOffsetType&, const MemOffsetType&) = default;
friend constexpr bool operator==(const MemOffsetType&, const MemOffsetType&) = default;
};
struct FEX_PACKED TypeDefinition final {
@@ -479,8 +487,7 @@ struct FEX_PACKED TypeDefinition final {
}
[[nodiscard]]
friend constexpr bool
operator==(const TypeDefinition&, const TypeDefinition&) = default;
friend constexpr bool operator==(const TypeDefinition&, const TypeDefinition&) = default;
};
static_assert(std::is_trivially_copyable_v<TypeDefinition>);
@@ -493,8 +500,7 @@ struct FEX_PACKED FenceType final {
return Val;
}
[[nodiscard]]
friend constexpr bool
operator==(const FenceType&, const FenceType&) = default;
friend constexpr bool operator==(const FenceType&, const FenceType&) = default;
};
struct FEX_PACKED RoundType final {
@@ -503,8 +509,7 @@ struct FEX_PACKED RoundType final {
return Val;
}
[[nodiscard]]
friend constexpr bool
operator==(const RoundType&, const RoundType&) = default;
friend constexpr bool operator==(const RoundType&, const RoundType&) = default;
};
class NodeIterator;
@@ -515,7 +520,10 @@ class NodeIterator;
*/
class NodeIterator {
public:
using value_type = std::tuple<OrderedNode*, IROp_Header*>;
struct value_type final {
OrderedNode *Node;
IROp_Header *Header;
};
using size_type = std::size_t;
using difference_type = std::ptrdiff_t;
using reference = value_type&;
@@ -537,14 +545,12 @@ public:
, Node {Ptr} {}
[[nodiscard]]
bool
operator==(const NodeIterator& rhs) const {
bool operator==(const NodeIterator& rhs) const {
return Node.NodeOffset == rhs.Node.NodeOffset;
}
[[nodiscard]]
bool
operator!=(const NodeIterator& rhs) const {
bool operator!=(const NodeIterator& rhs) const {
return !operator==(rhs);
}
@@ -561,15 +567,13 @@ public:
}
[[nodiscard]]
value_type
operator*() {
value_type operator*() {
OrderedNode* RealNode = Node.GetNode(BaseList);
return {RealNode, RealNode->Op(IRList)};
}
[[nodiscard]]
value_type
operator()() {
value_type operator()() {
OrderedNode* RealNode = Node.GetNode(BaseList);
return {RealNode, RealNode->Op(IRList)};
}
@@ -620,6 +624,8 @@ enum class ShiftType : uint8_t {
ROR,
};
enum class BranchHint : uint8_t { None = 0, Call, Return, CheckTF };
// Converts a size stored as an integer in to an OpSize enum.
// This is a nop operation and will be eliminated by the compiler.
@@ -767,7 +773,7 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
return NodeID(NodeOffset / sizeof(IR::OrderedNode));
}
bool IsFragmentExit(FEXCore::IR::IROps Op);
[[nodiscard]]
bool IsBlockExit(FEXCore::IR::IROps Op);
void Dump(fextl::stringstream* out, const IRListView* IR);
Loaded 100 of 564 files, more files were not shown because too many files have changed in this diff. Show more