Compare commits

..
541 Commits
Author SHA1 Message Date
Ryan Houdek 59659a7184 Docs: Update for release FEX-2508 2025-08-01 18:24:33 -07:00
Ryan Houdek e6de17e72e Merge pull request #4752 from alyssarosenzweig/opt/defer-next-use-analysis-2
RegisterAllocationPass: defer next-use analysis
2025-08-01 16:38:31 -07:00
Ryan Houdek 0457bdc7ce Merge pull request #4751 from lioncash/perm
ASIMDOps: Remove unused permute overloads
2025-08-01 12:45:50 -07:00
Ryan Houdek 7e54c2735c Merge pull request #4750 from lioncash/op6
ASIMDOps: Move remaining base opcodes into implementing function
2025-08-01 12:45:24 -07:00
Alyssa Rosenzweig 7bcc58687f IR: remove a bunch of unused atomic ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:44 -04:00
Alyssa Rosenzweig 6a03df7b8d RedundantFlagCalculationElimination: drop dead syscall/atomic opts
I don't think these are worth it, and also currently they don't trigger ever.

n=100:
Difference at 95.0% confidence
	-0.00245961 +/- 0.00139573
	-0.524468% +/- 0.297615%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 14:01:37 -04:00
Alyssa Rosenzweig 7ee5a8065c RegisterAllocationPass: defer next-use analysis
This is expensive and only needed for spilling, so only do it for spilling. This
complicates the RA a bit but speeds us up on average since most blocks
don't spill. Total results of this change (including the prep commits that
slowed things down temporarily):

Difference at 95.0% confidence
	-1.71952% +/- 0.455996%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 2f2353765f RegisterAllocationPass: use kill bits
this is a lot lighter weight than next uses for the same purpose.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig 8b4c2f5093 RegisterAllocationPass: consider AnySpilled at start of iteration
if we spill for SRA, we don't need/want to execute this code path. this will be
load bearing by the end of this series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 13:24:52 -04:00
Alyssa Rosenzweig d80662a9ba RegisterAllocationPass: ignore kill bit in SRA
needed for the backwards pass internally due to ordering. a little awkward but
shrug.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:11:53 -04:00
Alyssa Rosenzweig 68c5c72dbc IR: model kill bits
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 12:04:50 -04:00
Lioncache b460cbc91a ASIMDOps: Remove unused permute overloads
These aren't used at all and don't really provide anything that the existing
non-templated overloads can't.
2025-08-01 11:51:55 -04:00
Lioncache f7939078c2 ASIMDOps: Remove unnecessary ins() overload
There's no difference between this and the non-templated version, in fact,
this variant wasn't even used at all, so we can just remove it.
2025-08-01 11:24:44 -04:00
Lioncache e103af3b93 ASIMDOps: Move base opcode into ASIMDScalarCopy()
Now we have no more duplicated opcodes in ASIMDOps.
2025-08-01 11:13:58 -04:00
Lioncache 45bec06d5b ASIMDOps: Move base opcode into ASIMDFloatConvBetweenInt()
Deduplicates the second last remaining instruction category in ASIMDOps
2025-08-01 10:49:52 -04:00
Alyssa Rosenzweig c82efe7621 RegisterAllocationPass: set AnySpilled less
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-01 10:39:11 -04:00
Ryan Houdek ba85dbf526 Merge pull request #4749 from lioncash/op5
ASIMDOps: Move most remaining base opcodes into their implementing functions
2025-07-31 12:21:31 -07:00
Lioncache e624e87b29 ASIMDOps: Move base opcode into ASIMDExtract()
Minor deduplication of the base opcode.
2025-07-31 10:44:06 -04:00
Lioncache 6ff8d0953f ASIMDOps: Move base opcode into ASIMDPermute()
Gets rid of the need to respecify the base opcode in every instruction implementation
2025-07-31 10:43:17 -04:00
Lioncache 9f2fcdf0b3 ASIMDOps: Move base opcode into Crypto2RegSHA()/Crypto3RegSHA()
Removes the need to specify the base opcode multiple times.
2025-07-31 10:23:38 -04:00
Lioncache cecb9bfe32 ASIMDOps: Move base opcode into CryptoAES()
Minor deduplication.
2025-07-31 10:15:15 -04:00
Lioncache 6a989d844a ASIMDOps: Move base opcode into ASIMDTable()
Gets rid of duplication of the opcode in all instruction implementations.
2025-07-31 10:10:35 -04:00
Ryan Houdek 318b3115a8 Merge pull request #4748 from lioncash/op4
ASIMDOps: Constrain more instructions with IsQOrDRegister
2025-07-30 16:22:14 -07:00
Lioncache 017952e89e ASIMDOps: Constrain more instructions with IsQOrDRegister
We had quite a few instructions that we weren't constraining with this,
now the instruction itself will show up in the failure output more immediately
instead of failing on the internal implementation function if anything is
incorrectly passed through.
2025-07-30 19:03:33 -04:00
LC b0c74e1b66 Merge pull request #4747 from Sonicadvance1/cpuid
CPUID: Update documentation comments
2025-07-30 18:31:54 -04:00
Ryan Houdek a190086504 Merge pull request #4746 from lioncash/op3
ASIMDOps: Move base opcode into ASIMDModifiedImm()/ASIMDShiftByImm()
2025-07-30 15:30:37 -07:00
Ryan Houdek 84704d1cc2 CPUID: Update documentation comments
Additional reserved bits have set uses now.
Additionally set the cpuid bit for bus-lock-detect, because FEX
definitely detects bus-locks.
2025-07-30 15:13:08 -07:00
Lioncache b1ef8bbc5a ASIMDOps: Move base opcode into ASIMDModifiedImm() 2025-07-30 18:03:43 -04:00
Lioncache f19a6343dd ASIMDOps: Move base opcode into ASIMDShiftByImm()
Avoids needing to duplicate the opcode all over the place
2025-07-30 18:03:36 -04:00
Ryan Houdek 23d0d7d4a4 Merge pull request #4745 from lioncash/op2
ASIMDOps: Move base opcode into ASIMD2RegMisc()
2025-07-30 13:50:04 -07:00
Lioncache ed7bc28c21 ASIMDOps: Simplify size conditionals in several 2-reg misc instructions
In many of these, the long ternaries are equivalent to a subtraction by 1,
which is much more efficient
2025-07-30 16:31:33 -04:00
Lioncache 05dcb160bd ASIMDOps: Move base opcode into ASIMD2RegMisc()
Avoids the need to specify the opcode repeatedly, making implementations less noisy.
2025-07-30 16:25:24 -04:00
Ryan Houdek e49155ff90 Merge pull request #4744 from lioncash/op
ASIMDOps: Move base opcode into implementation function for some categories
2025-07-30 12:46:07 -07:00
Ryan Houdek 265525e472 Merge pull request #4740 from Sonicadvance1/multiple_segments
FEXCore/Frontend: Ensure multiple prefix bytes work
2025-07-30 12:43:51 -07:00
Ryan Houdek d4dcbfa90b FEXCore/Frontend: Ensure multiple prefix bytes work
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.

Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
2025-07-30 12:30:13 -07:00
Ryan Houdek 6724673482 Merge pull request #4739 from Sonicadvance1/remove_check
Frontend: Remove arbitrary check
2025-07-30 12:26:13 -07:00
Ryan Houdek d04f75df29 Frontend: Remove arbitrary check
REX prefix isn't even encoded in to the instruction tables if a 32-bit
process is running. Just remove this.
2025-07-30 11:53:15 -07:00
Ryan Houdek a5260f4233 Merge pull request #4738 from bylaws/peggle
Implement inline SMC handling for linux FEX
2025-07-30 11:48:28 -07:00
Ryan Houdek 369ca5cb72 Merge pull request #4719 from Sonicadvance1/runtime_mode_switch_take2
Runtime mode switch take 2
2025-07-30 11:47:54 -07:00
Lioncache a3e02cbafa ASIMDOps: Simplify conditionals in saddlv/uaddlv
Really all these size conversions are emulating is a subtraction by 1.
Also we can drop in an assert that was missed in uaddlv
2025-07-30 11:21:06 -04:00
Lioncache 4cca2f6915 ASIMDOps: Move base opcode into ASIMDAcrossLanes() 2025-07-30 11:07:08 -04:00
Lioncache 64ca48c4b1 ASIMDOps: Move base opcode into ASIMD3Different()
Deduplicates the open-coded base opcode in the instruction implementations.
2025-07-30 10:59:11 -04:00
Lioncache 74a897271f ASIMDOps: Move base opcode into ASIMD3Same()
Moves the base opcode into the actual implementation, so that we
aren't open-coding it into every relevant instruction function.
2025-07-30 10:40:17 -04:00
Tony Wasserka c6733a6eec Merge pull request #4731 from Sonicadvance1/armtifacts
github: Upload armtifacts
2025-07-30 10:22:49 +02:00
Tony Wasserka f7e99678b4 Merge pull request #4730 from Sonicadvance1/fix_thunk_functional_path
unittests: Fixes thunk unittest path
2025-07-30 10:20:46 +02:00
Tony Wasserka 976b68ffec Merge pull request #4713 from Sonicadvance1/free_the_stats
Profiler: Decouple profile stats from the profiler option
2025-07-30 10:19:07 +02:00
Billy Laws 7b656fb009 SyscallsSMCTracking: Support inline SMC 2025-07-30 00:13:27 +01:00
Billy Laws 20331d52c3 LinuxSyscalls: Always reconstruct RIP and EFLAGS when spilling from the JIT 2025-07-30 00:13:27 +01:00
Billy Laws 2556acb82d TestHarnessRunner: Avoid frontend SMC handling 2025-07-30 00:13:27 +01:00
Ryan Houdek 248f0948b8 github: Upload armtifacts 2025-07-29 12:02:57 -07:00
Ryan Houdek 6c12db500e InstCountCI: Update for segment changes 2025-07-29 12:02:38 -07:00
Ryan Houdek 2a8f4dbbb6 Arm64EC: Update for GDT 2025-07-29 12:02:38 -07:00
Ryan Houdek 91598d5178 WOW64: Update for GDT 2025-07-29 12:02:37 -07:00
Ryan Houdek a6bb9739d4 OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2025-07-29 12:02:37 -07:00
Ryan Houdek aa871c797b FEXCore: Accurately store segment descriptors
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.

In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.

Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
2025-07-29 12:02:37 -07:00
Ryan Houdek 153d20ca59 Rename SHMStats 2025-07-29 12:02:19 -07:00
Ryan Houdek 03b04771be TestHarnessRunner: Become a real thread
Stop being so special as a host runner.
2025-07-29 12:02:19 -07:00
Ryan Houdek 392fa62dae Profiler: Decouple profile stats from the profiler option
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
2025-07-29 12:02:19 -07:00
Ryan Houdek af3918491c unittests: Fixes thunk unittest path
This hasn't been run in CI for awhile so this was missed when I was
testing it. I need to see what I can do to get this going in CI again.
2025-07-29 11:46:30 -07:00
LC c0762bbd82 Merge pull request #4737 from Sonicadvance1/i_dislike_tuple_13
FileManagement: Remove pair usage from GetEmulatedFDPath
2025-07-29 10:48:38 -04:00
LC b68272b413 Merge pull request #4733 from Sonicadvance1/i_dislike_tuple_9
OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
2025-07-29 10:46:08 -04:00
LC 90b9414b5f Merge pull request #4732 from Sonicadvance1/i_dislike_tuple_8
FEXCore: Remove unused refcount_shared_mutex
2025-07-29 10:43:25 -04:00
LC 5db494d4f7 Merge pull request #4736 from Sonicadvance1/i_dislike_tuple_12
x87StackOptimizationPass: Removes pair usage
2025-07-29 10:42:55 -04:00
Tony Wasserka d073066293 Merge pull request #4722 from Sonicadvance1/fix_4124
CMake: Work around QCom disabling SVE in their chips
2025-07-29 10:54:18 +02:00
LC e80270ae43 Merge pull request #4735 from Sonicadvance1/i_dislike_tuple_11
64BitAllocator: Removes pair usage in allocator
2025-07-28 21:41:36 -04:00
LC cd56b85e88 Merge pull request #4734 from Sonicadvance1/i_dislike_tuple_10
Arm64: Remove pair usage in 128-bit loader
2025-07-28 21:40:57 -04:00
Ryan Houdek 64fbf55bdd FileManagement: Remove pair usage from GetEmulatedFDPath
NFC
2025-07-28 16:21:49 -07:00
Ryan Houdek 14c1ee10b6 x87StackOptimizationPass: Removes pair usage
NFC
2025-07-28 16:14:49 -07:00
Ryan Houdek 96ae671738 64BitAllocator: Removes pair usage in allocator
NFC
2025-07-28 16:07:07 -07:00
Ryan Houdek 16552d1194 Arm64: Remove pair usage in 128-bit loader
NFC
2025-07-28 15:58:11 -07:00
Ryan Houdek 6470c98ee2 OpcodeDispatcher: Remove pair usage from DecodeNZCVCondition
NFC
2025-07-28 15:50:31 -07:00
Ryan Houdek 126c4efa0e FEXCore: Remove unused refcount_shared_mutex 2025-07-28 15:43:44 -07:00
Ryan Houdek f6fd9e18d9 Merge pull request #4727 from bylaws/meopd
Windows: Lock invalidation tracking for the entire duration of memory ops
2025-07-28 11:59:23 -07:00
Ryan Houdek 1f554867e0 CMake: Work around QCom disabling SVE in their chips
Fixes #4124

clang feature checking can't check beyond MIDR, so compiling for a
specific cortex version means compiling SVE on these CPUs that disabled
them.

Just detect the particular situation in-which SVE isn't inside
proc/cpuinfo and is one of the snapdragon cores that are supposed to
support SVE. Then compile for Cortex-a78 instead.
2025-07-28 11:57:02 -07:00
Ryan Houdek d226331ccc Merge pull request #4726 from bylaws/ijwidjn
WOW64: Wrap BTCpuSimulate to ensure correct unwinding
2025-07-28 11:50:34 -07:00
Ryan Houdek 123f8c93b5 Merge pull request #4721 from alyssarosenzweig/ir/pool-as-you-go
Pool constants as we go
2025-07-28 11:50:08 -07:00
Ryan Houdek 6e93b0a9df Merge pull request #4725 from bylaws/sver
Windows: Force-disable SVE usage for now
2025-07-28 11:16:20 -07:00
Billy Laws a5d6d8c928 Windows: Lock invalidation tracking for the entire duration of memory ops 2025-07-28 17:46:14 +01:00
Billy Laws 561ee6601b WOW64: Wrap BTCpuSimulate to ensure correct unwinding
Some compiler versions generated FP-relative operations before loading
it from the stack, which would crash wine when APCs were used.
2025-07-28 17:24:49 +01:00
Billy Laws f44a7c95f6 Windows: Force-disable SVE usage for now 2025-07-28 17:21:57 +01:00
Tony Wasserka ac420347d1 Merge pull request #4712 from Sonicadvance1/move_legacy_binfmt
cmake: Move legacy binfmt arch-specific targets to combined
2025-07-28 09:46:10 +02:00
Tony Wasserka 493a7ccdb6 Merge pull request #4717 from Sonicadvance1/i_dislike_tuple_6
ArchHelpers: Remove pair usage in unaligned handler
2025-07-28 09:40:09 +02:00
Alyssa Rosenzweig 387db78120 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig 1338a99add IR: cap constant pool
If we have more constants than registers, something will be rematerialized. Use
a simple round-robin heuristic to pick instead of the better-but-slower approach
with RA. This is a heuristic to reduce JIT time with minimal impact on code
quality. In Instcountci, the only impact is a block in oblivion only increasing
instruction count by 0.2%. And moves of constants are free for cycles at least
on Firestorm, so this isn't where we want to spend piles of JIT time anyway.

Difference at 95.0% confidence
        -0.00138911 +/- 0.00104724
        -0.418608% +/- 0.315587%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 11:42:55 -04:00
Alyssa Rosenzweig f9edd20bf6 IR: pool constants in the emitter
This is slightly worse for x87 blocks since we can't share constants between the
x87 and the main code, but otherwise should be comparable and this avoids an
expensive remapping operation.

Difference at 95.0% confidence
	-0.00474273 +/- 0.00119189
	-1.40908% +/- 0.354114%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 10:50:17 -04:00
Alyssa Rosenzweig 6b3c7319c4 IR: wrap _Constant as Constant
flag day rename/wrapping. no functional change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:33:15 -04:00
Alyssa Rosenzweig bf51fc7c36 OpcodeDispatcher: do not use Constant as an identifier
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:32:42 -04:00
Alyssa Rosenzweig e910a81c12 OpcodeDispatcher: remove sized constant use
instcountci squashed for visibility.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:26:43 -04:00
Alyssa Rosenzweig 027bd93df9 IREmitter: remove unused
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-26 09:03:00 -04:00
LC 57c0d0f25e Merge pull request #4718 from Sonicadvance1/i_dislike_tuple_7
ELFContainer: Remove tuple usage
2025-07-25 23:04:52 -04:00
Ryan Houdek 24ff6bfde5 ELFContainer: Remove tuple usage 2025-07-25 19:33:32 -07:00
LC 78770683fc Merge pull request #4716 from Sonicadvance1/i_dislike_tuple_5
FEXCore: Replace CustomIREntry tuple with struct
2025-07-25 21:38:55 -04:00
LC c0008af877 Merge pull request #4714 from Sonicadvance1/i_dislike_tuple_3
IR: Remove tuple usage from NodeIterator
2025-07-25 21:37:56 -04:00
LC bffb81241a Merge pull request #4715 from Sonicadvance1/i_dislike_tuple_4
FEXCore: Remove reference SHA implementation
2025-07-25 21:34:29 -04:00
Ryan Houdek 53ff6d54c3 Merge pull request #4720 from tobhe/readme
Readme.md: Mention Ubuntu 25.04 as supported
2025-07-25 16:40:48 -07:00
Tobias Heider 2278e3334e Readme.md: Mention Ubuntu 25.04 as supported
Support was added in 1786c2f157
2025-07-26 00:38:26 +02:00
Ryan Houdek b0c61e2b69 ArchHelpers: Remove pair usage in unaligned handler
It's being treated like optional, where a value means it has been
handled, and no value means it hasn't been handled. Stop using pair in
this case.
2025-07-25 13:01:18 -07:00
Ryan Houdek a402c308ad FEXCore: Replace CustomIREntry tuple with struct 2025-07-25 12:33:25 -07:00
Ryan Houdek 6d83a195c6 InstcountCI: Remove sha 2025-07-25 12:22:13 -07:00
Ryan Houdek 0d69e88d53 FEXCore: Remove reference SHA implementation
Due to us only enabling the CPUID extension in the case that the host
hardware supports SHA or not, this has actually been largely unused now.
Also the only hardware that doesn't support the crypto extension has
been some old Pi hardware and some other things we don't really care
about.

This code was a phenomenal reference point for implementing the SHA
versions of the instructions and would have been significantly more
difficult to implement had this not been available. Kudos to @lioncash
for having written it!

But now as we are no longer utilizing it, it is time to remove it.
2025-07-25 12:18:13 -07:00
Ryan Houdek 7309a5c3e1 HostFeatures: Fixes SHA check for vixl sim
Oops, was accidentally checking for rand.
2025-07-25 12:17:45 -07:00
Ryan Houdek 9e3c7aa9b3 IR: Remove tuple usage from NodeIterator 2025-07-25 12:10:39 -07:00
LC 30fb992cf7 Merge pull request #4708 from Sonicadvance1/geekbench_microtest
unittests/instcountci: Adds a long-lived ymm_high test
2025-07-24 21:31:04 -04:00
LC 8e5db99ba1 Merge pull request #4709 from Sonicadvance1/lock_xadd
unittests/instcountci: Adds missing lock xadd tests
2025-07-24 21:30:22 -04:00
Ryan Houdek 8103cdbec4 Merge pull request #4711 from bylaws/rwxinf
InvalidationTracker: Fix queries in the non-intersecting RWX case
2025-07-24 16:23:11 -07:00
Ryan Houdek b11835aab5 cmake: Move legacy binfmt arch-specific targets to combined
This matches the systemd path, no more _32 and _64 versions, just
binfmt_misc.

Having a mixed install where one program does one architecture and
another is weird and unsupported anyway.
2025-07-24 16:18:35 -07:00
Ryan Houdek 5b855d9baf unittests/instcountci: Adds missing lock xadd tests
Realized we were missing these, Looks like some moves could be
eliminated.
2025-07-24 15:20:02 -07:00
Billy Laws 56b6b96def InvalidationTracker: Fix queries in the non-intersecting RWX case
This needs to return the base of the non-RWX interval, with a size
that when added to the base is the start of the RWX interval.
2025-07-24 22:58:02 +01:00
Ryan Houdek 3744cad3d8 unittests/instcountci: Adds a long-lived ymm_high test
Spills in to spill-slots when it should instead spill in to the context.
2025-07-24 10:24:53 -07:00
Ryan Houdek bf8b5ed9ad Merge pull request #4670 from bylaws/callret
Implement call-ret stack optimisations
2025-07-24 10:24:30 -07:00
Billy Laws 49482fe963 InstCountCI: Update 2025-07-24 14:53:09 +01:00
Billy Laws 3497870a45 JIT: Guard lookupcache locks with the code invalidation mutex
Avoids issues with forking, as the code invalidation mutex is fork-safe.
2025-07-24 14:53:09 +01:00
Billy Laws 9a1efacefe OpcodeDispatcher: Allow for direct linking of non-multiblock direct jumps
Using an add here prevents ExitFunction from taking the direct path.

Reported by chengmingtang on Discord.
2025-07-24 14:53:09 +01:00
Billy Laws 3efb2379ee Dispatcher: Keep the call-ret stack balanced for thunk callbacks 2025-07-24 14:53:09 +01:00
Billy Laws 20efadbe66 OpcodeDispatcher: Treat ThunkOp ExitFunction as a return
ThunkOp acts as an implicit return, mark it as such so the call-ret
stack entry from the caller is popped
2025-07-24 14:53:09 +01:00
Billy Laws 31f8abfa61 Linux: Manage the call-ret stack 2025-07-24 14:53:09 +01:00
Billy Laws cf4cc71010 Dispatcher: Opportunistically perform a call-ret stack return on EC entry 2025-07-24 14:53:09 +01:00
Billy Laws 44107757a3 JIT: Rewrite block linking to support direct ExitFunction calls
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.

To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset

If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.

Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.

(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.

(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).

(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).

(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).

(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).

(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
2025-07-24 14:53:09 +01:00
Billy Laws 45ba1af388 BranchOps: Use the call-ret stack to optimise indirect ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws ba9884a26a JIT: Emit entrypoint code for call return target blocks
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
2025-07-24 14:53:09 +01:00
Billy Laws a40d53497b OpcodeDispatcher: Emit hints for call/ret instructions 2025-07-24 14:53:09 +01:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Billy Laws c604893944 WOW64: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 2f16d25ab8 ARM64EC: Implement call-ret stack management 2025-07-24 14:53:09 +01:00
Billy Laws 132de3160a Windows: Implement common call-ret stack helpers
Allocate the stack with uncommited guard pages on either side, if
any faulting accesses to these occur then the call-ret SP value is
reset to the default and execution resumed.
2025-07-24 14:53:09 +01:00
Billy Laws d67b1645c9 FEXCore: Hold a frontend allocation for the call-ret stack
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
2025-07-24 14:53:09 +01:00
Billy Laws 68270ad425 LookupCache: Return whether Erase removed any cache entries 2025-07-24 14:53:09 +01:00
Billy Laws 963a8c2f08 FEXCore: Save and restore the call/ret SP from CPUState
For simplicity in cases like signal handling, always load it in
fill and store in spill, even the SP is stored in a callee save
register.
2025-07-24 14:53:09 +01:00
Billy Laws a80581e52b Arm64Emitter: Allocate a register for the call-ret SP
The host stack pointer can't be reused since explicit bounds checks
would be far too expensive, and on-stack signals prevent implicit ones
using guard pages from working.
2025-07-24 14:53:09 +01:00
Billy Laws 882cdaaeda Frontend: Move to initializing persistent sets directly 2025-07-24 14:52:54 +01:00
Billy Laws fd58f17dbe FEXCore: Track block executable ranges prior to adding cache entries
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.

This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
2025-07-24 14:52:54 +01:00
Billy Laws 7e5c0d7174 LookupCache: Move CodePages to GuestToHostMap
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.

The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
2025-07-24 14:52:54 +01:00
Ryan Houdek 525462e982 Merge pull request #4642 from pmatos/feature/x87-invalid-operation-bit
Implement x87 invalid operation bit on F80 mode
2025-07-23 11:05:28 -07:00
Ryan Houdek 507b3a69f0 Merge pull request #4699 from bylaws/srawow64
Inline SMC fixes and support for WOW64
2025-07-23 11:03:18 -07:00
Billy Laws b15e49910a WOW64: Support inline SMC
Matches the ARM64EC handling
2025-07-23 18:46:28 +01:00
Billy Laws 573d0858ac ARM64EC: Fix inline SMC handling for writes in the same page as the current block
Consider a page with two blocks in it, A and B. Block A performs SMC on B then A.
With the previous logic, the SMC write of B would unprotect the page and then
the inline SMC touching A would not be detected and a single-step would not be forced.
2025-07-23 18:46:28 +01:00
Billy Laws c3c2789cdf Windows: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:28 +01:00
Billy Laws e9f14e32dd Linux: Specify the single-inst argument when calling the dispatcher 2025-07-23 18:46:26 +01:00
Billy Laws 9b65d819f4 Dispatcher: Add an argument to request a single-step on SRA fill
This already exists for ARM64EC, but this path is required for linux
and WOW64
2025-07-23 18:29:12 +01:00
Billy Laws 287344986c FEXCore: Switch ENTRY_FILL_SRA_SINGLE_INST_REG to TMP2
Will allow this to be taken as the second dispatcher argument
2025-07-23 18:29:12 +01:00
Paulo Matos 7386b4b037 instcountci: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 9a2f62c479 asm_tests: Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:16 +02:00
Paulo Matos 726656d0bf Implement x87 invalid operation bit on F80 mode 2025-07-23 16:01:10 +02:00
Ryan Houdek fe61aabc52 Merge pull request #4701 from bylaws/noexecwin
Windows: Correctly report noexec faults
2025-07-22 12:54:28 -07:00
Ryan Houdek 564e966730 Merge pull request #4698 from lioncash/float2
ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
2025-07-22 12:36:16 -07:00
LC 98abfa4395 Merge pull request #4700 from bylaws/geiex
WinAPI: Avoid sign-extension of processor count in GetSystemInfo
2025-07-22 15:02:17 -04:00
Billy Laws 8e6fbe7615 Windows: Correctly report noexec faults 2025-07-22 19:30:31 +01:00
Billy Laws 27fd5505be OpcodeDispatcher: Set correct trap number for noexec faults 2025-07-22 19:30:23 +01:00
Billy Laws e383cc89eb unittests: Test for the noexec pagefault trap number 2025-07-22 19:29:17 +01:00
Alexandre Julliard a71a0ed137 WinAPI: Avoid sign-extension of processor count in GetSystemInfo 2025-07-22 19:28:36 +01:00
Lioncache ab6fc279ce ASIMDOps/SVEOps: Use IsStandardFloatSize() even more
Forgot in #4697 to take into account that there are still parts of the
emitter that qualify with its own namespace.

This also removes those where applicable to be more in line with the rest
of the cases.
2025-07-22 08:25:35 -04:00
Ryan Houdek dd5c17291a Merge pull request #4697 from lioncash/float
SVEOps: Make use of IsStandardFloatSize() more
2025-07-21 12:37:03 -07:00
Ryan Houdek 5fe06b7413 Merge pull request #4696 from lioncash/opcode
OpcodeDispatcher: Mark functions as const/static where applicable
2025-07-21 12:36:22 -07:00
Lioncache 2a8efa9891 SVEOps: Make use of IsStandardFloatSize() more
Initially introduced in #4687 as a utility for ASIMD ops, this can be used
elsewhere as well to reduce the verbosity of some other assertions.
2025-07-21 13:51:37 -04:00
LC cf43e5eaaf Merge pull request #4694 from Sonicadvance1/fix_moffset_instcountci
InstcountCI: Fix bad encoded moffset instructions
2025-07-21 10:53:26 -04:00
Lioncache cafc906b9e OpcodeDispatcher: Mark functions as const/static where applicable
Not a functional change, but makes it more obvious how these rely on object state.
2025-07-21 10:48:55 -04:00
Tony Wasserka 3387f51e24 Merge pull request #4663 from Sonicadvance1/update_format_requires
External/code-format-helper: Update requirements
2025-07-21 10:17:52 +02:00
Ryan Houdek ccf1bb26af InstcountCI: Fix bad encoded moffset instructions
These were actually encoded incorrectly, and I noticed they changed in
PR #4670 for some reason. So I dove in to the nasm source to figure out
what it takes to encode these sanely rather than with raw bytes, turns
out it's fairly trivial, just annoying to track in their parser through
iwdq->mem_offset->MEM_OFFS->qword.

Should remove the change in #4670 once it rebases.
2025-07-18 12:32:56 -07:00
Ryan Houdek fc1ca01a0b Merge pull request #4693 from neobrain/fix_libfwd_wl
LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options
2025-07-18 10:57:50 -07:00
Tony Wasserka e508f0900c LibraryForwarding/wayland: Add method signatures required by steam-runtime-launch-options 2025-07-18 11:38:46 +02:00
LC 5123ca52ba Merge pull request #4692 from Sonicadvance1/reintroduce_cssc
FEXCore: Reintroduce support for CSSC
2025-07-17 21:21:46 -04:00
Ryan Houdek 7e1ee5bb07 FEXCore: Reintroduce support for CSSC
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.

Been a while since I last looked at this, added a new instcountci file
to show the improvement.
2025-07-17 15:09:27 -07:00
LC b7ac641aa0 Merge pull request #4691 from Sonicadvance1/ignore_jemalloc_config_sources
External: Update jemallocs
2025-07-17 11:29:20 -04:00
Ryan Houdek 251ac145d9 Merge pull request #4650 from pmatos/feature/ReformatChanged
Add --changed flag to reformat.sh script
2025-07-16 23:33:29 -07:00
Ryan Houdek e17c2c8c73 Merge pull request #4661 from pmatos/feature/reformat-clang-format-19
Whole-tree reformat with clang-format-19
2025-07-16 23:33:09 -07:00
Paulo Matos 4473054d33 Add reformat sha to ignore revs for git-blame 2025-07-17 08:11:08 +02:00
Paulo Matos 5267cde60e Whole-tree reformat with clang-format-19 2025-07-17 08:10:00 +02:00
Paulo Matos 3c7ece2ef3 Update .clang-format 2025-07-17 08:09:25 +02:00
Paulo Matos 3b322d8f37 Add --changed flag to reformat.sh script 2025-07-17 08:04:41 +02:00
LC 1d4b6c6a74 Merge pull request #4689 from Sonicadvance1/disable_gcs_protection
Disable GCS in simulator and userspace
2025-07-16 20:25:16 -04:00
Billy Laws f5efe1d251 Merge pull request #4690 from Sonicadvance1/moar_padding
JIT: Add more padding
2025-07-17 00:58:37 +01:00
Ryan Houdek 09db3aba0a External: Update jemallocs
Ignore any additional jemalloc configuration options, can come from
`/etc/malloc.conf` or even the `MALLOC_CONF` environment variable.

Ensure nothing can override our options.
2025-07-16 16:35:48 -07:00
Ryan Houdek aec7deca9d Merge pull request #4688 from lioncash/hfloat2
ASIMDOps: Merge half-float 3-reg same with single/double variants
2025-07-16 15:42:08 -07:00
Ryan Houdek 8f1d4bc710 FEXLoader: Check for GCS being enabled
There is a ELF note for this but currently clang doesn't support a
`no-gcs` flag. The best we can do is check if the kernel has GCS enabled
for the current process and early exit.

Then continue to use the kernel's locking functionality to disable it if
the guest happens to try, ensuring safety.
2025-07-16 15:41:23 -07:00
Ryan Houdek 7edc8417b6 JIT: Add more padding
Go to a whole page of additional padding, #4670 adds some more size to a
block and overran the padding causing a unittest to fail.
2025-07-16 15:28:05 -07:00
Ryan Houdek 5b01682641 Arm64Emitter: Disable GCS in the simulator
FEX isn't going to be compatible with this.
PR #4670 requires this
2025-07-16 15:23:43 -07:00
Lioncache a615f8970a ASIMDOps: Constrain 3-reg same float arguments with IsQOrDRegister
These are only intended to be used with QRegisters or DRegisters, so we can constrain these
so that the proper types are enforced at compile-time.
2025-07-16 17:46:28 -04:00
Lioncache 708ea9d50a ASIMDOps: Merge half-float 3-reg same with single/double variants 2025-07-16 17:36:24 -04:00
Ryan Houdek e36a64ecc0 Merge pull request #4681 from Sonicadvance1/i_dislike_tuple_2
FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
2025-07-16 14:23:18 -07:00
Ryan Houdek 82be09b6a3 FEXCore/OpcodeDispatcher: Removes tuple usage for Dispatch tables
NFC
2025-07-16 13:57:16 -07:00
Ryan Houdek eb7655d93c External/code-format-helper: Update requirements
Latest of everything, let's see what happens.
To get rid of dependabot alerts.
2025-07-16 13:54:43 -07:00
Ryan Houdek c1fe841bf3 Merge pull request #4687 from lioncash/hfloat
ASIMDOps: Merge half-float 2-reg misc with single/double variants
2025-07-16 13:50:53 -07:00
Lioncache 82ffb379c4 ASIMDOps: Merge half-float 2-reg misc with single/double variants
Unifies the interface, so that there's no need for a stark difference.

Previously, calling the single/double variant didn't require explicit template
arguments, but the half-float version did, which is inconsistent.

Technically it also made the interface more cumbersome to use in the event the
element size isn't able to be determined as a constant ahead of time.
2025-07-16 10:16:49 -04:00
Tony Wasserka befae52993 Merge pull request #4685 from pmatos/fix/clang-format-19-ignore
Ensure .clang-format-ignore is compatible with clang-format-19
2025-07-16 15:32:09 +02:00
Paulo Matos 48597d7682 Ensure .clang-format-ignore is compatible with clang-format-19
Since clang-format-19 doesn't support globstar yet, add .clang-format
to disable formatting inside External/.
2025-07-16 15:20:03 +02:00
Lioncache 05d45ccb7f Emitter: Add helper for sanitizing floating point element sizes
Will be used in a following change to reduce the amount of duplication
made in some floating point instructions.
2025-07-16 09:19:45 -04:00
LC e0ca04bfec Merge pull request #4684 from Sonicadvance1/spurious_syscall_crash
FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
2025-07-15 21:44:51 -04:00
LC 34566861ba Merge pull request #4683 from Sonicadvance1/waitpkg_mostly_nop
FEXCore: Implement a mostly NOP implementation of waitpkg
2025-07-15 21:43:11 -04:00
Ryan Houdek dfe3b4502a FEXCore/OpcodeDispatcher: Fix a spurious crash that can occur with multiblock
If multiblock discovers a codepath with an `int 0x80` then it would
crash the emulator even if it never gets executed. Ensure that this
ERROR_AND_DIE_FMT instead just gets handled as an UnhandledOp to ensure
the Core early terminates the block.

Found by having steamwebhelper spuriously crash when it hit this.
Adds a simple unittest to ensure discovery doesn't break again.
2025-07-15 17:43:16 -07:00
Ryan Houdek 4884aeef20 Config: Adds option to disable wfxt 2025-07-15 16:38:15 -07:00
Ryan Houdek 219ab44a82 FEXCore: Implement a mostly NOP implementation of waitpkg
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.

Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.

This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
2025-07-15 16:13:29 -07:00
Ryan Houdek 29fccd0e1b Merge pull request #4682 from bylaws/asahi
Windows: Support enabling hardware TSO on Asahi Linux
2025-07-15 14:25:03 -07:00
Ryan Houdek 35e4ac5a64 Merge pull request #4668 from Sonicadvance1/implement_nx
FEXCore: Implement support for NX bit.
2025-07-15 12:40:48 -07:00
Ryan Houdek 43bba77840 FEXCore: Implement support for NX bit.
Long time coming but thanks to bylaw's changes in #4474, this is now
trivial to implement.

Fixes #2175
2025-07-15 10:41:24 -07:00
Billy Laws 0361274366 Windows: Support enabling hardware TSO on Asahi
Relies on a wine-side patch to expose the prctl to the PE-side
2025-07-15 17:45:21 +01:00
Ryan Houdek 5b0e703811 Merge pull request #4678 from neobrain/refactor_unuse_unused
Drop unnecessary uses of maybe_unused
2025-07-15 09:03:55 -07:00
Ryan Houdek 05742a8e28 Merge pull request #4677 from neobrain/refactor_misc_logs
Miscellaneous log changes
2025-07-15 09:03:12 -07:00
Ryan Houdek 170fbb596c Merge pull request #4679 from neobrain/fix_signed_overflow
Arm64Emitter: Fix signed overflow
2025-07-15 09:02:29 -07:00
Ryan Houdek b72bdf9948 Merge pull request #4680 from lioncash/vfaddv
VectorOps: Correct benign VFAddV op cast
2025-07-15 09:00:41 -07:00
Lioncache 678e03b470 VectorOps: Correct benign VFAddV op cast
This was using the non-float VAddV variant, but had the same behavior,
since the fields were named the same. So this is just a correctness fix.
2025-07-15 11:41:28 -04:00
Tony Wasserka 0f45f3a24d Arm64Emitter: Fix signed integer overflows 2025-07-15 17:17:48 +02:00
Tony Wasserka 9b503f3702 IRDumper: Remove unnecessary use of maybe_unused 2025-07-15 17:10:37 +02:00
Tony Wasserka 7f3f619518 GDBJIT: Remove unnecessary use of maybe_unused 2025-07-15 17:10:27 +02:00
Tony Wasserka f7d59880cc ELFCodeLoader: Drop unused maybe_unused attribute 2025-07-15 17:07:36 +02:00
Tony Wasserka ec9e266c3c FDUtils: Mark get_fdpath as nodiscard 2025-07-15 17:07:36 +02:00
Tony Wasserka ed62c02494 AllocatorOverride: Make error message more prominent 2025-07-15 16:55:18 +02:00
Tony Wasserka d5b166bdd9 OpcodeDispatcher: Strengthen error message to fatal 2025-07-15 16:55:18 +02:00
Tony Wasserka 648c726bb1 LogManager: Use assert log level for ERROR_AND_DIE_FMT 2025-07-15 16:55:18 +02:00
Tony Wasserka 88f30afb26 Merge pull request #4672 from neobrain/refactor_format_formatter
code-format-helper: Fix formatting for the format helper
2025-07-15 14:27:22 +02:00
LC e6f319d6b3 Merge pull request #4674 from neobrain/refactor_todrop
Drop TODO defines
2025-07-15 07:12:35 -04:00
LC a80d8b15c2 Merge pull request #4673 from neobrain/refactor_value_nor
XXFileHash: Drop unnecessary use of value_or
2025-07-15 07:09:47 -04:00
LC 2a372895e4 Merge pull request #4671 from Sonicadvance1/xsaveopt
FEXCore: Implement xsaveopt
2025-07-15 07:08:15 -04:00
Tony Wasserka 3b67e573b5 Revert "FexHeaderUtils: Add TodoDefines"
This reverts commit ad1fd7f54b.
2025-07-15 10:04:01 +02:00
Tony Wasserka a3a55d19b8 Revert "FEX_TODO: Convert some XXX to FEX_TODO"
This reverts commit 256df76674.
2025-07-15 10:03:56 +02:00
Tony Wasserka 4e25bce616 FEXServer: Drop use of FEX_TODO macro 2025-07-15 10:03:49 +02:00
Tony Wasserka 135477e539 XXFileHash: Drop unnecessary use of value_or 2025-07-15 09:47:52 +02:00
Tony Wasserka 147b1f2293 code-format-helper: Replace spurious tab with spaces 2025-07-15 09:43:55 +02:00
Ryan Houdek 401dc69586 InstcountCI: Add xsaveopt
Just in-case it changes.
2025-07-14 17:32:45 -07:00
Ryan Houdek 0dff992fab FEXCore: Implement xsaveopt
It's the same as xsave because we don't track a hidden "xinuse" hardware
mask. So this is a trivial implementation.
2025-07-14 17:31:47 -07:00
LC 4f605739f1 Merge pull request #4667 from Sonicadvance1/i_dislike_tuple
XXFileHash: Remove a tuple usage
2025-07-14 18:21:54 -04:00
LC e71e10adcb Merge pull request #4669 from Sonicadvance1/handful_cpuinfo_missing
EmulatedFiles/cpuinfo: Add a few missing flags
2025-07-14 18:21:03 -04:00
Ryan Houdek 9f1bc05644 EmulatedFiles/cpuinfo: Add a few missing flags
bus_lock_detect might be interesting to expose if applications ever
change behaviour depending on the underlying fault behaviour.
2025-07-14 14:56:24 -07:00
Ryan Houdek 603878b3e9 XXFileHash: Remove a tuple usage 2025-07-14 13:09:27 -07:00
Ryan Houdek 70ce323537 Merge pull request #4666 from pmatos/fix/git-clang-format-19
Set clang-format-19 as the version git-clang-format should run in wor…
2025-07-14 10:38:14 -07:00
Paulo Matos b5f8dd93f7 Set clang-format-19 as the version git-clang-format
Unfortunately git-clang-format-19 will not call clang-format-19 but the system clang-format
so we need to hardcode it here.
2025-07-14 18:37:47 +02:00
Ryan Houdek 2e7b86e452 Merge pull request #4474 from bylaws/mentry
Support multiple entrypoints into a multiblock and executable permission tracking
2025-07-11 19:05:42 -07:00
Ryan Houdek b2eeaf75a9 Merge pull request #4659 from bylaws/sysfix
ARM64EC: Rely on syscall export sorting for deriving their IDs
2025-07-11 18:09:07 -07:00
Ryan Houdek 5d4bddcfb9 Merge pull request #4662 from pmatos/patch-2 2025-07-11 08:49:45 -07:00
Paulo Matos 13ef676115 Do not reformat files in External/ 2025-07-11 15:54:02 +02:00
Ryan Houdek 67bfd3877c Merge pull request #4651 from pmatos/feature/ClangFormat19Upgrade
Upgrade to clang-format-19
2025-07-11 01:03:49 -07:00
Ryan Houdek 2be619b011 Merge pull request #4658 from bylaws/fmt
External: Update libfmt to master to support clang 21
2025-07-10 11:08:33 -07:00
Ryan Houdek ebd0a3ceaf Merge pull request #4660 from lioncash/null
VixlUtils: Fix null pointer dereference vector in IsImmLogical()
2025-07-10 11:07:01 -07:00
Lioncache d8da14d550 VixlUtils: Fix null pointer dereference vector in IsImmLogical()
Technically we can end up doing null pointer dereferencing here since checks were being
chained with || instead of &&. So we can just separate the checks out.
2025-07-10 12:09:20 -04:00
Billy Laws 076c156cbb Frontend: Warn on invalid instructions in entry blocks 2025-07-10 16:51:34 +01:00
Billy Laws 6eaaab8dc2 IntervalList: Also return the full matching interval on query 2025-07-10 16:51:34 +01:00
Billy Laws 46bc8e499a Frontend: Always treat FEXCore X86 callbacks as executable 2025-07-10 16:51:34 +01:00
Billy Laws 21865d09a2 SyscallsSMCTracking: Handle READ_IMPLIES_EXEC for executable mapping queries 2025-07-10 16:51:34 +01:00
Billy Laws 31903d0c0b Frontend: Treat instructions in non-executable memory as invalid 2025-07-10 16:51:33 +01:00
Paulo Matos c7b7cbcac1 Upgrade to clang-format-19
Fixes #4577
2025-07-10 17:11:44 +02:00
Billy Laws d0af858a9c LinuxEmulation: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 2e5c6283f5 CodeSizeValidation: Add dummy QueryGuestExecutableRange impl 2025-07-10 16:00:24 +01:00
Billy Laws 261df26cff DummyHandlers: Stub SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 96bd6ce3e5 Windows: Implement SyscallHandler executable mapping queries 2025-07-10 16:00:24 +01:00
Billy Laws 14eed89bb6 SyscallHandler: Add method to query executable memory ranges 2025-07-10 16:00:24 +01:00
Billy Laws 3feb354186 Frontend: Keep the associated thread object as a member
Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.
2025-07-10 16:00:24 +01:00
Billy Laws d41cb3b69d FEXCore: Support multiple entrypoints into a multiblock
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
2025-07-10 16:00:24 +01:00
Billy Laws cdef1ed0c5 Frontend: Explore after call instructions with multiblock
Each instruction after a call instruction can be treated as an
additional entrypoint to the multiblock
2025-07-10 16:00:24 +01:00
Billy Laws 4f79acc32a X86Tables: Mark call instructions with a flag 2025-07-10 16:00:24 +01:00
Billy Laws 9328b099c3 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 16:00:24 +01:00
Billy Laws 378d0351cf External: Update libfmt to master to support clang 21 2025-07-10 16:00:24 +01:00
Billy Laws 793e3adbb4 ARM64EC: Rely on syscall export sorting for deriving their IDs
Rather than relying on wine-specific alphabetical behaviour, that
has since been changed. Rely on their addresses being sorted which
is more stable behaviour also present in Windows.
2025-07-10 16:00:07 +01:00
Billy Laws 2f4f4253e7 External: Update libfmt to master to support clang 21 2025-07-10 15:52:17 +01:00
Ryan Houdek 10c69d561e Merge pull request #4656 from bylaws/mingw-new
CMake: Set CMAKE_AR for MinGW toolchains
2025-07-09 19:26:30 -07:00
Billy Laws bbf6e8e0c5 CMake: Set CMAKE_AR for MinGW toolchains
Avoids the need for the toolchain to override the system AR.
2025-07-10 00:33:37 +01:00
Ryan Houdek d8a4d03501 Merge pull request #4655 from alyssarosenzweig/bug/divisor-mask
Fix divisor masking
2025-07-09 12:15:16 -07:00
Alyssa Rosenzweig 8ecc8ef75d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig cfd4318080 unittests: add 32-bit masking divide unit test
fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Alyssa Rosenzweig d23bb01e96 JIT: fix divisor masking
oversight. should fix Steam.

Fixes: de4becc26 ("OpcodeDispatcher: mask certain divisors")
Closes: #4652
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-09 15:00:42 -04:00
Ryan Houdek 60f63ca467 Merge pull request #4649 from bylaws/rexvex
Frontend: Raise #UD on invalid VEX/REX encodings
2025-07-09 10:37:50 -07:00
LC 40378cd4d8 Merge pull request #4654 from Sonicadvance1/vcvtsd2si_test
unittests: Update vcvtsd2si test
2025-07-09 12:16:02 -04:00
Ryan Houdek f44af48751 Merge pull request #4648 from bylaws/sfmodreg
SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops
2025-07-09 09:02:12 -07:00
Ryan Houdek d01a48b773 unittests: Update vcvtsd2si test
This would have failed prior to #4647 getting merged.
2025-07-09 08:50:19 -07:00
Ryan Houdek bd136eae71 Merge pull request #4647 from bylaws/sizrfix
OpcodeDispatcher: Fix several cases where incorrectly sized loads/stores could be used
2025-07-09 08:49:38 -07:00
LC e669370629 Merge pull request #4646 from bylaws/uaffix
PoolBufferWithTimedRetirement: Unclaim in dtor
2025-07-09 10:14:13 -04:00
Billy Laws 26215d5425 VEXTables: Complete decode flag information 2025-07-09 00:53:23 +01:00
Billy Laws baa5c6f79b Frontend: Require VEX.vvvv is 0 when unused 2025-07-09 00:53:23 +01:00
Billy Laws c0fcb27aa5 Frontend: Add instruction flags to specify valid REX.W encodings 2025-07-09 00:53:23 +01:00
Billy Laws 02b7938b93 Frontend: Add instruction flags to specify valid VEX.L encodings 2025-07-09 00:53:23 +01:00
Billy Laws 00fec3d51f Update InstCountCI 2025-07-09 00:52:59 +01:00
Billy Laws 0f6ceaac05 OpcodeDispatcher: Load only the element size from memory for VFMAScalarImpl 2025-07-09 00:52:59 +01:00
Billy Laws 635816c07f OpcodeDispatcher: Force ElementSize loads for UCOMISxOp 2025-07-09 00:52:59 +01:00
Billy Laws 2fac5c23ff OpcodeDispatcher: Always use 32-bit load/store for {LD,ST}MXCSR 2025-07-09 00:52:59 +01:00
Billy Laws ff4c1cf0d5 X86Tables: Fix (V)CV(T)TSD2SI size flags
This led to incorrect OOB handling for 32-bit dests.
2025-07-09 00:52:45 +01:00
Billy Laws c2e2d1e92f OpcodeDispatcher: Only read at most ElementSize in CVTFPR_To_GPR 2025-07-09 00:50:47 +01:00
Billy Laws 0ab56bad95 SecondaryGroupTables: Specify FLAGS_SF_MOD_REG_ONLY for more ops 2025-07-08 23:45:32 +01:00
Billy Laws 407c5a0f78 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-08 23:38:43 +01:00
Ryan Houdek 3ba84ad06a Docs: Update for release FEX-2507 2025-07-07 23:49:56 -07:00
Ryan Houdek c6aae9e05a Merge pull request #4634 from Sonicadvance1/fix_horizon
EmulatedFiles: Emulate `current_clocksource`
2025-07-07 21:16:54 -07:00
Ryan Houdek 95b4618833 Merge pull request #4644 from ChanthMiao/fix/sigframe_mistake
Fix: wrong magic value in fpstate.
2025-07-06 17:07:31 -07:00
Changwei Miao 1fe17d55d9 Fix: wrong magic value in fpstate.
FEX should only set fpx_sw_bytes.magic1 with FP_XSTATE_MAGIC when
avx is enabled. Otherwise it may cause segfault in ntdll::save_context,
which requires access to extended xstate info if magic1 equals FP_XSTATE_MAGIC.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-07-06 18:44:37 +08:00
Ryan Houdek ead73371d9 EmulatedFiles: Emulate current_clocksource
WINE uses this to determine TSC frequency and because it doesn't say
`tsc` on ARM devices, it was ignoring TSC and instead using CPU maximum
frequency.

This was causing Horizon to think the TSC ran at whatever the max
frequency of a core was  (1.8Ghz to 2.6Ghz depending?) This was causing
all of Horizon Zero Dawn's physics to run at slower than real time
speeds because our 1Ghz (on Orion) TSC is significantly lower than the
max clock speeds of the cores.

This is still a bug in Wine that it is using the maximum CPU clock speed
in the case of current_clocksource not being TSC, but that's a battle
for a different time.
2025-07-05 21:10:53 -07:00
Ryan Houdek 6f089a4323 Merge pull request #4643 from tstellar/llvm-21
Fix build with LLVM >= 21
2025-07-05 16:51:38 -07:00
Ryan Houdek 640f024551 Merge pull request #4641 from Sonicadvance1/optimize_sincos
JIT: Optimize x87 FSINCOS
2025-07-05 16:26:37 -07:00
Tom Stellard 99920f89dd Fix build with LLVM >= 21 2025-07-05 16:50:34 +00:00
Ryan Houdek 8ea276267f InstcountCI: Update 2025-07-03 18:05:52 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek fa0a54deb9 Merge pull request #4640 from alyssarosenzweig/bug/fix-hades
Fix Hades
2025-07-03 14:53:01 -07:00
Alyssa Rosenzweig 046043090f unittests: add move merging test
this hits a nasty case with post-RA merging. fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:56 -04:00
Alyssa Rosenzweig 61150a18cc RegisterAllocationPass: fix bookkeeping with merging
this fixes Hades.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:53 -04:00
Ryan Houdek 95ca20cfee Merge pull request #4636 from alyssarosenzweig/bug/ra-invariant
RegisterAllocationPass: assert an invariant in post-RA prop
2025-07-02 18:17:08 -07:00
Alyssa Rosenzweig 5a536d47fd RegisterAllocationPass: assert an invariant in post-RA prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:53:47 -04:00
Ryan Houdek afbc7da027 Merge pull request #4629 from alyssarosenzweig/opt/cpuid-basic
Optimize some constant cpuid/xgetbv cases
2025-07-02 10:25:01 -07:00
Alyssa Rosenzweig c093c08c40 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 360d8c629e RegisterAllocationPass: optimize cpuid
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)

however this special case we can (and should) optimize for now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig abb41d39e4 RegisterAllocationPass: optimize xgetbv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig d4eb4ef594 IR: plumb CPUID into RA pass
for cpuid folding.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek 62de1004df Merge pull request #4635 from neobrain/fix_libfwd_wl_more
LibraryForwarding/wayland: Add new interface objects
2025-07-02 09:01:11 -07:00
Tony Wasserka 7f216ca02f LibraryForwarding/wayland: Add new interface objects 2025-07-02 16:51:03 +02:00
LC 492b0fdda8 Merge pull request #4632 from neobrain/feature_nix
Build: Add nix-based helpers to facilitate cross-compilation
2025-07-01 16:23:58 -04:00
LC bb072c0112 Merge pull request #4633 from Sonicadvance1/noexec_test
unittests: Adds unittest for no-exec testing
2025-07-01 16:20:44 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Ryan Houdek 38e0fc2434 unittests: Adds unittest for no-exec testing
In preparation for #4474
2025-06-30 13:19:33 -07:00
Alyssa Rosenzweig 16a70eafc6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:59:27 -04:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Ryan Houdek 94af96df8f Merge pull request #4631 from neobrain/fix_libfwd_wl_cutter
LibraryForwarding: Fix various Wayland issues
2025-06-30 11:38:13 -07:00
Ryan Houdek e685ab818e Merge pull request #4619 from neobrain/refactor_drop_config_h_in
Remove code generation build step for install prefix
2025-06-30 11:35:26 -07:00
Ryan Houdek 5f2a72b65b Merge pull request #4624 from Sonicadvance1/static_analysis_fixes
Some static code analysis fixes
2025-06-30 11:28:06 -07:00
Tony Wasserka c1842a6167 Build: Add nix-based helpers to manage toolchains for ARM64EC/WOW64
The cmake_configure_woa*.sh scripts will automatically install any required
cross-toolchains required to enable either ARM64EC or WOW64 builds of FEX,
and it will initialize the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell WineOnArm/shell.nix`,
which will make the toolchain available via environment variables. This also
generates a meson crossfile for building VKD3D or vkd3d-proton.
2025-06-30 16:50:13 +02:00
Tony Wasserka 3f3907b5d1 Build: Add nix-based helpers to manage toolchains for FEXLinuxTests
The cmake_enable_flt.sh script will automatically install the required
cross-toolchains required to build FEXLinuxTests, and it will reconfigure
the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell FEXLinuxTests/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 1899465390 Build: Add nix-based helpers to manage toolchains for library forwarding
The cmake_enable_libfwd.sh script will automatically install any required
cross- toolchains and development headers required to enable library
forwarding, and it will reconfigure the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell LibraryForwarding/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 18360d4ccb LibraryForwarding/wayland: Add more method signatures 2025-06-27 10:57:10 +02:00
Tony Wasserka a691c3cd99 LibraryForwarding/wayland: Fix mprotect call when crossing page boundaries 2025-06-27 10:57:10 +02:00
Tony Wasserka 70bc561bbf LibraryForwarding/wayland: Fix wl_proxy_marshal_array_constructor 2025-06-27 10:57:10 +02:00
Tony Wasserka b9222d8431 LibraryForwarding/unittests: Fix build with clang 20 2025-06-27 10:57:10 +02:00
Tony Wasserka a5fad89e57 Merge pull request #4621 from neobrain/feature_update_readme
Update Readme.md
2025-06-26 21:54:53 +02:00
Tony Wasserka a14360b89d Update Readme.md 2025-06-26 21:41:23 +02:00
Tony Wasserka c9aaedd217 CPack: Update package description 2025-06-26 21:41:23 +02:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 62410c4381 InstCountCI: add another udiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:22 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00
Ryan Houdek 4a74bea7ab InstcountCI: Don't use a global static initializer for CodeSizeValidation
Relies on fmt facet initialization order which isn't guaranteed to have
correct initialization order.

CID 482003
2025-06-19 16:51:26 -07:00
Ryan Houdek c6d8e60ef8 OpcodeDispatcher: Make sure to initialize ArithRef
CID 482002
2025-06-19 16:46:15 -07:00
Ryan Houdek 1212cd526a VDSOEmulation: Sanitize sysconf result
Unlikely to fail but make sure.

CID 482021
2025-06-19 16:43:53 -07:00
Ryan Houdek 79a8ed53b6 CPUBackend: Make sure to zero initialize variable
CID 482022
2025-06-19 16:41:24 -07:00
Ryan Houdek a6ce115d9c FEXLoader: Handle bad LogFile path
Just go silent but print a log about it.

CID 482031
CID 482023
2025-06-19 16:40:03 -07:00
Ryan Houdek 646a5a7f9e CodeEmitter/ASIMD: Removes redundant ternary selection
Redundant and unnecessary.

CID 482017
2025-06-19 16:31:08 -07:00
Ryan Houdek 7f71b6f1b2 Passes: Use move instead of copy semantics
To initialize in-place.

CID 482035
2025-06-19 16:22:39 -07:00
Ryan Houdek 16d5ca447f FEXCore: DebugData is never null now
This is a required data structure to exist.

CID 482036
2025-06-19 16:21:19 -07:00
Ryan Houdek 3d0c20a263 Merge pull request #4615 from neobrain/feature_logging_qol
Improve log message formatting
2025-06-19 12:50:44 -07:00
Ryan Houdek 1c38b8b046 Merge pull request #4614 from alyssarosenzweig/opt/easy-mov-elim
Merge moves that are immediately consumed
2025-06-19 12:26:07 -07:00
Alyssa Rosenzweig e1124480be InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 5243f50ed1 RegisterAllocationPass: merge 32-bit mov + 64-bit and
mov wA, wB
  and xA, xA, ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig cde805147f RegisterAllocationPass: merge 32-bit moves
mov wA, wB
  op wA, wA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 7e39eb3df2 RegisterAllocationPass: merge full size moves
mov xA, xB
  op xA, xA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig af366d4480 RegisterAllocationPass: skip inlineconstant in RA
similar reasoning as guestopcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 33ef98aae7 RegisterAllocationPass: refactor push/pop merge
to make way for move merging.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig c16db2db4a RegisterAllocationPass: simplify an expression
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Ryan Houdek 0072b289bb Merge pull request #4622 from neobrain/feature_fexconfig_logo
FEXConfig: Add icon
2025-06-18 01:46:30 -07:00
Tony Wasserka e2bd79087e FEXConfig: Add icon 2025-06-18 10:33:07 +02:00
Ryan Houdek 6cd78fc90d Merge pull request #4618 from neobrain/feature_infer_32bit_libfwd_paths
LibraryForwarding: Infer folders for 32-bit wrappers automatically
2025-06-17 17:55:38 -07:00
Ryan Houdek c346aca241 Merge pull request #4620 from neobrain/refactor_data_files
Move data files to Data/
2025-06-17 17:54:28 -07:00
Ryan Houdek bf83569f0b Merge pull request #4623 from neobrain/feature_tracy_0_12
External: Update Tracy submodule to version 0.12.1
2025-06-17 17:53:47 -07:00
Tony Wasserka 1502f04a8a External: Update Tracy submodule to version 0.12.1
Notably, this update adds flame graph functionality for aggregated data.
2025-06-17 17:51:52 +02:00
Tony Wasserka 7bd9d0ae23 Data: Move Dockerfile 2025-06-17 16:40:42 +02:00
Tony Wasserka cf57afdf26 Data: Move CI folder to Data/CI 2025-06-17 16:40:42 +02:00
Tony Wasserka 43e6aebc7a Data: Move CMake support scripts to Data/CMake 2025-06-17 16:40:42 +02:00
Tony Wasserka 22780993e1 Data: Move CPack files to Data/CMake/ 2025-06-17 16:40:42 +02:00
Tony Wasserka 578dcee9af Data: Move toolchain files to a central location 2025-06-17 16:40:42 +02:00
Tony Wasserka 61d77e3f9b LibraryForwarding: Infer folders for 32-bit wrappers automatically
There's no need to bother the user to select these paths manually.
Instead, just use the same folder names with _32 appended.

Fixes #4588.
2025-06-17 11:57:35 +02:00
Tony Wasserka 4ce0acba80 Remove now unused Config.h.in 2025-06-17 11:40:32 +02:00
Tony Wasserka d137212222 FEXGetConfig: Infer install prefix from executable path 2025-06-17 11:40:32 +02:00
Tony Wasserka 9ad4e3a6a0 FEXBash: Clean up and fix FEXInterpreter lookup
Previously, the first attempt to look up a FEXInterpreter would always
fail due to a missing path separator.

Additionally, fallback lookup now uses /proc/self/exe to find a path relative
to the FEXBash executable. The previous use of FindContainerPrefix does not
seem to be required anymore in current Steam versions.
2025-06-17 11:40:32 +02:00
Tony Wasserka 57627d4fcf LogManager: Drop unused STDOUT/STDERR log levels 2025-06-16 13:54:03 +02:00
Tony Wasserka 23b69271eb Use consistent log message formatting for all modules 2025-06-16 13:54:03 +02:00
LC 9d2f557666 Merge pull request #4616 from alyssarosenzweig/ici/32bit-div
InstructionCountCI: add 32-bit division cases
2025-06-13 13:18:55 -04:00
Alyssa Rosenzweig 4f9e352ff0 InstructionCountCI: add 32-bit division cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-13 13:05:35 -04:00
Tony Wasserka 3e85e60a30 LogManager: Use colors for logging when possible 2025-06-13 15:14:47 +02:00
Tony Wasserka febce21b21 FEXServer: Separate PID and TID with a pipe instead of a period
Since most software considers the former a word boundary but not the latter,
this allows the individual values to be copy-pasted more conveniently.
2025-06-13 14:54:22 +02:00
Tony Wasserka 755364e2df FEXServer: Clean up time display for logging
This is now relative to the time of the first message. Furthermore, display
precision is limited to milliseconds (which are actually zero-padded now!).
2025-06-13 14:54:13 +02:00
Tony Wasserka df63979773 LogManager: Avoid ^C being printed when quitting foreground FEXServer 2025-06-13 14:31:25 +02:00
Tony Wasserka 992d86bbc1 LogManager: Shorten debug level strings to a single letter 2025-06-13 14:31:25 +02:00
Ryan Houdek 534b338161 Merge pull request #4612 from alyssarosenzweig/ir/merge-divrem-2
IR: merge integer division & remainder
2025-06-12 14:29:02 -07:00
Alyssa Rosenzweig ef250f936c InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig e57130e364 JIT: drop UDiv extensions
we already extend in the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig eedcb35270 IR: merge ldiv/lrem handlers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
LC 1f15a4e35b Merge pull request #4611 from Sonicadvance1/ubisoft_ptrace
Linux: Implement enough of ptrace to allow Ubisoft launcher
2025-06-10 11:35:49 -04:00
Ryan Houdek 7ed9bea16b Linux: Implement enough of ptrace to allow Ubisoft launcher
Wine does some minimal attaching, peeking, poking, and detaching to have
a different process inspect another's TEB region. Ubisoft's launcher
does this to check if a debugger is attached and rejects it in the case
that it is.

While this implementation isn't all-encompassing, it is good enough for
this family of games.
2025-06-09 16:32:42 -07:00
Ryan Houdek 8b1383d235 Docs: Update for release FEX-2506 2025-06-04 10:48:24 -07:00
LC a73fab3bb5 Merge pull request #4609 from Sonicadvance1/add_plucky_install
Scripts/InstallFEX: Add plucky
2025-06-03 16:30:45 -04:00
Ryan Houdek 7bd64d9c53 Merge pull request #4608 from alyssarosenzweig/ici/case
InstructionCountCI: fix sdiv case
2025-06-03 12:50:36 -07:00
Ryan Houdek 1786c2f157 Scripts/InstallFEX: Add plucky
This was added to the PPA last month.
2025-06-03 12:31:40 -07:00
Alyssa Rosenzweig 958b671736 InstructionCountCI: fix sdiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:30:42 -04:00
Ryan Houdek 6549b66cf6 Merge pull request #4606 from alyssarosenzweig/opt/cdq
Optimize CDQ
2025-06-03 12:24:15 -07:00
Alyssa Rosenzweig 0c855a5ce3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:39 -04:00
Alyssa Rosenzweig 6864d48dcf OpcodeDispatcher: optimize cdq
prereq to optimizing sign-ext+ldiv in a reasonable way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:06 -04:00
Ryan Houdek 99816a23a8 Merge pull request #4583 from neobrain/refactor_syscalls_unify
LinuxSyscalls: Reduce code duplication between 32-bit and 64-bit paths
2025-06-03 09:54:09 -07:00
Ryan Houdek 2ff9546523 Merge pull request #4599 from pmatos/FEXServerFind
Fixes FEXServer path search
2025-06-03 09:52:39 -07:00
Ryan Houdek 06541f21d6 Merge pull request #4605 from neobrain/fix_cmake_full_libdir2
Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values
2025-06-03 09:52:26 -07:00
Tony Wasserka 45a37edd4a Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values 2025-06-03 17:45:29 +02:00
Ryan Houdek 2a713a1f51 Merge pull request #4604 from neobrain/fix_cmake_install_prefix
CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
2025-06-03 08:43:24 -07:00
Tony Wasserka 3e104de377 CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
This allows changes to CMAKE_INSTALL_PREFIX to be automatically picked up and
propagated properly. Previously, you had to update 4 variables in 3 files by
hand to do so.
2025-06-03 16:58:06 +02:00
Paulo Matos b22b316e70 Removes handling of FEXServer from the FEXBash wrapper
This was not necessary since FEXInterpreter is already doing it,
and doing it properly. The previous implementation if FEXBash was
incomplete.
2025-06-03 12:24:03 +02:00
Tony Wasserka df461546c5 LinuxSyscalls: Make error return values consistent 2025-06-03 11:10:08 +02:00
Tony Wasserka 1c1c43cd86 LinuxSyscalls: Fix formatting 2025-06-03 11:10:08 +02:00
Tony Wasserka 3eac9f937e LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:08 +02:00
Tony Wasserka f13a0d8e84 LinuxSyscalls: Unify shmdt implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ca12dc9213 LinuxSyscalls: Unify shmat implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 365ed2cd70 LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 7efbfed0bf LinuxSyscalls: Unify mmap and munmap implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ed502738c6 LinuxSyscalls: Make Get32BitAllocator interface virtual 2025-06-03 11:10:07 +02:00
Ryan Houdek ba162bb058 Merge pull request #4603 from neobrain/fix_cmake_full_libdir
LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
2025-06-02 15:49:22 -07:00
Tony Wasserka f194d35913 LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
This fixes issues in the nix build, where CMAKE_INSTALL_LIBDIR is an
absolute path instead of the typical "lib(64)".
2025-06-03 00:37:23 +02:00
Ryan Houdek e81c84e00a Merge pull request #4602 from Sonicadvance1/add_instcountci_tests
InstcountCI: Adds tests for instructions discovered by #4597
2025-06-02 12:45:51 -07:00
Ryan Houdek 68dc9030bc Merge pull request #4601 from Sonicadvance1/fix_futimesat_flake
unittests/futimesat: Fixes flake
2025-06-02 12:45:26 -07:00
Ryan Houdek fedad275e7 Merge pull request #4597 from alyssarosenzweig/opt/drop-pile-of-constprop
Constant fold on the fly
2025-06-02 12:34:53 -07:00
Ryan Houdek f6b4c76d76 InstcountCI: Adds tests for instructions discovered by #4597
Apparently I completely missed that cpuid, xgetbv, syscall,
l{u,}{div,rem} were failing to hit their optimized cases for inlining
and avoiding 128-bit software divide.

The divisions are a clear performance regression for 64-bit applications
since that is the only real way to do a 64-bit division on x86, I added
those specifically because it sped up games.

CPUID depends heavily on the game, since some games use that as a
serialization instruction fairly heavily.

XGETBV is trivial since it matches behaviour of CPUID (and is basically
an extension of it).

Syscall inlining can save a decent amount of time, again heavily depends
on game.

Adds multi-inst tests for all of these situations so that once it gets
fixed (Apparently broken once RCLSE got stripped out), we can see that
they keep working. Obviously tests couldn't have existed in instcountCI
before since we didn't support multi-instruction tests.
2025-06-02 12:10:39 -07:00
Alyssa Rosenzweig 27854aa091 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig d966ae145e ConstProp: merge inline + pooling
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44474704    0.46750433    0.45455258    0.45446569  0.0044727894
+  50    0.43149892    0.45173984    0.44252267    0.44295575  0.0045621814
Difference at 95.0% confidence
	-0.0115099 +/- 0.00179263
	-2.53263% +/- 0.394447%
	(Student's t, pooled s = 0.00451771)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig 4cb37e6a1b ConstProp: drop constant folding and algebraic opts
No longer needed.

The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.

    N           Min           Max        Median           Avg        Stddev
x 100     0.4472467    0.46646308    0.45708424    0.45713057  0.0040838243
+ 100    0.44707586    0.46581227    0.45479448    0.45509309  0.0037548573
Difference at 95.0% confidence
	-0.00203748 +/- 0.00108734
	-0.445711% +/- 0.237862%
	(Student's t, pooled s = 0.00392279)

...in addition to a net deletion of 144 lines of code.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:24:58 -04:00
Alyssa Rosenzweig d9da81e99b ConstProp: drop dead cross-instr opts
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 7d734740be Addressing: avoid a Bfe
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 06299cca4b OpcodeDispatcher: optimize out Bfi for storereg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig af5aaab38b OpcodeDispatcher: optimize a few ALU-with-constant ops
instead of relying on ConstProp for this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 77bb01d384 OpcodeDispatcher: don't generate pointless Xor for AF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig a9eb1bd4e6 OpcodeDispatcher: don't generate pointless Bfe for moves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 83ef2da95f OpcodeDispatcher: avoid zero shift in SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ba13ccadb8 OpcodeDispatcher: use ArithRef for 8/16-bit imul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 4a2dee873d OpcodeDispatcher: use ArithRef for small rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 589906e6e6 OpcodeDispatcher: generalize AF on constants optimization
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig f0fcf6d9e6 OpcodeDispatcher: optimize SVE vmovmskpd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 1591ced5a7 OpcodeDispatcher: use ArithRef for ADC/SBB flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 0d235d63d0 OpcodeDispatcher: use ArithRef for bit tests
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 820d5c1447 OpcodeDispatcher: use ArithRef for rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 8093e8f5f1 OpcodeDispatcher: optimize LoadDir
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 802857e411 OpcodeDispatcher: add ArithRef helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig bf3a1839a2 OpcodeDispatcher: transform DF ourselves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ad7844d7da OpcodeDispatcher: do not generate useless VMov
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 3ff9128a8a unittests: add asm test for mov ah, 0
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Ryan Houdek c865eb98ea Merge pull request #4598 from alyssarosenzweig/idc
Stop using hashmap in DCE
2025-06-02 10:26:08 -07:00
Ryan Houdek 217a9228f8 unittests/futimesat: Fixes flake
Due to interactions between file times and the lack of granularity in
futimesat, if we don't remove the nanoseconds then this test can flake.

Easy enough.
2025-06-02 10:21:42 -07:00
Alyssa Rosenzweig 105d0a36ad RedundantFlagCalculationElimination: dont use hashmap
combined results on node from this and the previous commit:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44433582    0.46988457    0.45344824    0.45320041  0.0043288623
+  50     0.4385365    0.46615359    0.45026179    0.44996216  0.0045742233
Difference at 95.0% confidence
	-0.00323825 +/- 0.00176704
	-0.71453% +/- 0.389903%
	(Student's t, pooled s = 0.00445323)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Alyssa Rosenzweig 263279d5dd IR: index blocks
this will let us avoid a costly hashmap in DCE.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Ryan Houdek 861ecbec0f Merge pull request #4600 from neobrain/fix_libfwd_cross_target
LibraryForwarding: Change target env to gnu
2025-06-02 09:21:46 -07:00
Ryan Houdek 5ff9bb3669 Merge pull request #4596 from pmatos/JSONOpts
Error when JSON config includes unknown options
2025-06-02 08:50:47 -07:00
Ryan Houdek e794584bb5 Merge pull request #4479 from neobrain/feature_codebuffer_sharing
Reduce JIT time by 25% by sharing code buffers between threads
2025-06-02 08:50:17 -07:00
Tony Wasserka a792dd0703 CPUBackend: Clean up CodeBuffer size limits 2025-06-01 22:45:50 +02:00
Tony Wasserka a9a6a645bf Arm64Emitter: Disable PC-relative constant encoding
This no longer works since the JIT output is now relocated before execution.
2025-06-01 22:45:50 +02:00
Tony Wasserka 7c93becd5f JIT: Increase estimate for CodeBuffer space use
The previous bound was exceeded during Steam startup before.
2025-06-01 22:45:50 +02:00
Tony Wasserka 4bbaef58e9 JIT: Re-enable parallel compilation by compiling to a temporary buffer 2025-06-01 22:45:50 +02:00
Tony Wasserka 95791a985a Core: Minimize the time CodeBufferWriteMutex is held 2025-06-01 22:45:50 +02:00
Tony Wasserka 0dfefe9730 Core: Re-check LookupCache before running compiler backend
This further reduces lock contention by skipping the backend phase in case
another thread raced the active one for the same block.
2025-06-01 22:45:49 +02:00
Tony Wasserka 8481c797df FEXCore/JIT: Extend LookupCache lock to all of ExitFunctionLink 2025-06-01 22:44:49 +02:00
Tony Wasserka a503b5e20b Rename CodeBufferManager reference 2025-06-01 22:44:49 +02:00
Tony Wasserka 4078840ef1 Core: Reduce JIT time by sharing CodeBuffers between threads
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
  be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
  in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
  the modifying thread. Old versions of the CodeBuffer persist as read-only
  data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
  decrease the reference count and eventually trigger deallocation of the
  old version
2025-06-01 22:44:49 +02:00
Tony Wasserka ab51958b26 Move AllocateNewCodeBuffer from CPUBackend to a new CodeBufferManager interface 2025-06-01 22:42:55 +02:00
Tony Wasserka 3376587b6a CPUBackend: Manage CodeBuffers using shared_ptr
This is required for sharing CodeBuffers between threads anyway, but it also
allows use of the constructor/destructor to manage memory automatically.
2025-06-01 22:42:55 +02:00
Tony Wasserka 6681d7dcf9 LookupCache: Split L3 cache into a dedicated interface
This data isn't really a cache, since the JIT is directly responsible of
writing its contents. Instead it be considered the source to populate the
L1/L2 caches from.

Furthermore, splitting off this data allows it to be shared across threads
in the future without affecting L1/L2 caches.
2025-06-01 22:42:55 +02:00
Tony Wasserka 34224481c6 LookupCache: Prefer empty() over a size check 2025-06-01 22:42:55 +02:00
Tony Wasserka 1837aaabe4 fextl: Add shared_ptr and make_shared 2025-06-01 22:42:55 +02:00
Tony Wasserka 5811914a78 LibraryForwarding: Change target env to gnu
This is required for clang to find architecture-specific libstdc++ headers as
distributed by NixOS.
2025-06-01 21:40:33 +02:00
Paulo Matos a109a4efa1 Error when JSON config includes unknown options 2025-06-01 11:59:51 +02:00
Alyssa Rosenzweig a08a6ce5de Merge pull request #4576 from Sonicadvance1/fix_vma_race
Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
2025-05-30 08:14:38 -04:00
Ryan Houdek 3dc8a3ddc1 Merge pull request #4595 from neobrain/refactor_config_templates
Config: Clean up use of templates
2025-05-29 12:26:34 -07:00
Ryan Houdek 5e103365f7 Merge pull request #4594 from neobrain/fix_self_move
Async: Don't destruct on self-moves
2025-05-29 12:26:24 -07:00
Ryan Houdek cdea8d7f74 Merge pull request #4593 from alyssarosenzweig/ici/zeroing-sub-regs
InstructionCountCI: add more cases for mov 0/~0
2025-05-29 12:26:15 -07:00
Ryan Houdek ef6dc3d802 Merge pull request #4592 from alyssarosenzweig/opt/x87-tag
OpcodeDispatcher: optimize X87FTWTag
2025-05-29 12:26:04 -07:00
Ryan Houdek e6edb349ba Linux: Fixes some vestigial mmap handling
Missed this in previous commits, had to double check that I got them
all.
2025-05-29 12:15:26 -07:00
Ryan Houdek ad132267ec Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.

This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.

The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>)             = 0
<...>
41227 <... mmap resumed>)               = 0xba87b000
```

While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```

The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.

This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.

The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.

Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
2025-05-29 12:15:26 -07:00
Ryan Houdek 6f837281ef Merge pull request #4591 from Sonicadvance1/fix_warning2
FEXCore/CPUID: Remove warning
2025-05-29 11:21:49 -07:00
Tony Wasserka 2a1d29d2df Config: Clean up use of templates 2025-05-29 18:38:35 +02:00
Tony Wasserka bdceb4ca89 Async: Don't destruct on self-moves
FEXServer's logger performs such self-moves for the log pipe posix_descriptor
of short-lived clients. This resulted in a double-close previously, which
could interfere with file operations on other threads (typically crashing
FEXServer in effect).
2025-05-29 18:28:56 +02:00
Alyssa Rosenzweig 656bb928cf InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig 0fbe69ebcf OpcodeDispatcher: optimize xor-with-self flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig ece817c691 OpcodeDispatcher: optimize logical flags
seems to be strictly better.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:34:50 -04:00
Alyssa Rosenzweig 3fbc8204b7 OpcodeDispatcher: clean up logical flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:28:31 -04:00
Alyssa Rosenzweig 44a5481254 InstructionCountCI: add more cases for mov 0/~0
some obvious opportunities to improve here! although mostly I want this to
regression test my constprop rework.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:26:12 -04:00
Alyssa Rosenzweig cf2ff90f87 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:14:06 -04:00
Alyssa Rosenzweig b8dd5d95b0 OpcodeDispatcher: optimize X87FTWTag
using bit twiddling tricks :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:13:16 -04:00
Ryan Houdek 7cd52febc2 FEXCore/CPUID: Remove warning
This is currently only used on win32 builds because Linux doesn't
understand TPIDRRO.
2025-05-28 09:28:41 -07:00
Ryan Houdek 8e079c1965 Merge pull request #4590 from Sonicadvance1/remove_stack_set_arm
TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
2025-05-27 09:21:14 -07:00
Ryan Houdek 4928af5a64 Merge pull request #4585 from alyssarosenzweig/opt/pair-push-pop
Pair push/pop
2025-05-27 09:17:47 -07:00
Alyssa Rosenzweig 289df740cd InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:13 -04:00
Alyssa Rosenzweig 6ad7392cd7 RegisterAllocationPass: pair push/pop
as a simple post-RA peephole. much much easier to do post-RA than pre-RA.

This isn't a post-RA /pass/ in the traditional sense... it's done while
assigning registers to coalesce the passes over the IR, since we pay per-pass
and we can merge the walks over the IR.

Closes: #4480
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:12 -04:00
Alyssa Rosenzweig b15d5f299c RegisterAllocationPass: drop dead IP increment
written but never read.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:32:33 -04:00
Ryan Houdek 5c7c959dd9 TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
This matches behaviour with the x86 host runner that RSP isn't
guaranteed to be set to a valid memory location. Make sure when running
tests on ARM that it gets the same behaviour.
2025-05-26 13:02:07 -07:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 88682a457a IR: add RemovePostRA helper
Remove blows up because of use tracking, but we can do a much simpler version
for post-RA and elide lots of checks from trying to make Remove more general.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 0e67f30103 OpcodeDispatcher: make Push do the right thing and use it
this both optimizes and bug-fixes pusha while deleting a snotton of code.

Closes: https://github.com/FEX-Emu/FEX/issues/4589
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig babd6e9a7b RegisterAllocationPass: allow Copy on the input IR
useful for pusha.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 76caa2c6e3 unittests: add pop-to-same-reg test
an earlier version of this PR failed this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
Ryan Houdek ded8b3284a Merge pull request #4584 from neobrain/feature_assert_source_location
LogManager: Print source location when failing assertions
2025-05-24 01:15:11 -07:00
Tony Wasserka 2e24ee7a5f LogManager: Print source location when failing assertions 2025-05-24 09:35:11 +02:00
Tony Wasserka 048ae597d2 CodeEmitter: Fix incorrect macro parameter passing
Macros aren't aware of C++ templates, so the comma is considered a macro
argument separator unless the argument is wrapped in parentheses.
2025-05-24 09:35:11 +02:00
Ryan Houdek 0147f7aa19 Merge pull request #4587 from stanfordzhang/main
Fix callee saved floating-point arguments order issue
2025-05-23 20:13:58 -07:00
StanfordZhang 23cda2c961 Update Arm64Emitter.cpp
fix callee saved floating-point arguments issue
2025-05-23 22:14:37 +08:00
Alyssa Rosenzweig feb67658e1 RegisterAllocationPass: exploit new IR
Now that we can just set registers directly, we can simplify RA a lot. All the
Map/Unmap nonsense - it all goes away. We just assign registers as we go and
everything clicks into place naturally.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 49f8332c5b JIT: use registers directly from the IR
This is the flag day change from the series, using all the new shiny
infrastructre we added.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig afce108ed7 JIT: make almost all the DEF_OPs common
this deduplicates a bunch of #defines, letting us change the signature easier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 30f3b545af RegisterAllocationPass: use Header spill slots instead
Removes even more RAData dependence.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig bf1597920c IR: use post-RA flag
rather than implicitly depending on the RA data.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig f66bf3811e RegisterAllocationPass: set PostRA flag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 3136a5e2f8 IR: add helpers to extract physical registers from IR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig c4f00a05df IR: add space for registers right in OrderedNode *
This again follows the same idea of eliminating the RAData sideband.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig e9a8f8a9ff IR: generate builders that take OrderedNodeWrapper
when we need to materialize instructions inside RA without having a
corresponding OrderedNode* source, we want to just pass a OrderedNodeWrapper
with an encoded register. generate appropriate builders so this is possible.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 261ae7a195 IR: add immediates into OrderedNodeWrapper
this will let us encode registers directly inside OrderedNodeWrapper, rather
than pointers to OrderedNode *. that will let us speed up RA & onwards.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 7eaf5ae9e0 IR: drop FillRegister original source
this is now unused, and it's problematic with future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig 10a02449b1 IR: drop RA validation
There's no reasonable way to keep this around without adding significant
complexity to RA. This series prefers to drop complexity from RA, lessening the
need for validation in the first place.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig debc57e8c7 Merge pull request #4575 from alyssarosenzweig/cleanup/no-implicit-size
IR: drop DestSize inference
2025-05-16 15:17:53 -04:00
Tony Wasserka 7a4fff8e5b Merge pull request #4578 from cjacek/libgcc
Use -print-libgcc-file-name to get the compiler-rt file name
2025-05-16 13:40:25 +02:00
Jacek Caban ec0a2a8671 Use -print-libgcc-file-name to get the compiler-rt file name
Avoid hardcoding the file name. Upstream Clang currently expects the aarch64 variant of
compiler-rt. Changing this to use arm64ec in the file name could be problematic in the future,
if the MinGW toolchain gains support for ARM64X. Querying Clang for the correct file name is
the most forward-compatible approach.
2025-05-16 11:24:57 +02:00
LC b8b516f7b6 Merge pull request #4563 from Sonicadvance1/minor_vma_cleanup2
SyscallsVMATracking: Minor cleanup
2025-05-15 19:15:20 -04:00
Alyssa Rosenzweig 3b1e91b1fc IR: drop DestSize inference
no more users, and it's problematic for upcoming work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 15:00:03 -04:00
Alyssa Rosenzweig 8532593d91 IR: specify DestSize for RMWHandle
seems to just have been an oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:58:05 -04:00
Alyssa Rosenzweig 55bd16e2c8 IR: make FillRegister sizes explicit
instead of hacking around it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:57:51 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Ryan Houdek 9b573effd1 Merge pull request #4573 from alyssarosenzweig/cleanup/jit-id-2
JIT: use .ID() even less
2025-05-14 13:36:33 -07:00
Alyssa Rosenzweig dfd1aedae5 JIT: use .ID() even less
oops, missed a "/g" with th sed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 16:23:22 -04:00
Ryan Houdek 221ae2d7b4 Merge pull request #4571 from alyssarosenzweig/opt/constprop-xor-1
ConstProp: optimize XOR with all-1
2025-05-14 12:59:12 -07:00
Ryan Houdek 9127d206b5 Merge pull request #4572 from alyssarosenzweig/cleanup/jit-id
JIT: stop using .ID() pattern
2025-05-14 12:59:01 -07:00
Alyssa Rosenzweig ecc6fea54e JIT: stop using .ID() pattern
sed -ie 's/.ID()//' *

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:46:32 -04:00
Alyssa Rosenzweig 99446da7c1 JIT: add Reg helpers taking OrderedNodeWrappers
more ergonomic and will give us freedom to migrate things easier soon.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:43:20 -04:00
Alyssa Rosenzweig ad75563a26 IR: drop irrelevant reference to x86
we don't run the JIT on x86 anymore.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:32:07 -04:00
Alyssa Rosenzweig 197af972d9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Alyssa Rosenzweig fbd706c191 ConstProp: optimize XOR with all-1
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Ryan Houdek 6b85fa5611 SignalScopeGuards: Review 2025-05-14 10:56:44 -07:00
Ryan Houdek fae66a921f SyscallsVMATracking: Moves MappedResources implementation details to private 2025-05-14 10:56:44 -07:00
Ryan Houdek e6fc462e9d SyscallsVMATracking: Rename VMA tracking functions
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.

NFC, just helps my brain wrap this more easily.
2025-05-14 10:56:44 -07:00
Ryan Houdek d11a265a53 SyscallsVMATracking: Adds validation that thread has ownership of VMA mutex
This way we can capture any programming bugs. Can only check the
functions that actually require unique locks rather than shared locks.
2025-05-14 10:56:44 -07:00
Ryan Houdek 395d870814 SignalScopeGuards: Add checks for locks being held by the calling thread
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
2025-05-14 10:56:44 -07:00
Ryan Houdek fec1ffaa6b Merge pull request #4566 from neobrain/feature_pool_alloc_size
ThreadPoolAllocator: Add support for updating the size of managed data
2025-05-14 10:53:34 -07:00
Tony Wasserka d4fdb28e72 PoolBufferWithTimedRetirement: Add support for updating the buffer size
This should be done at low frequency since it may unclaim the buffer.
2025-05-14 13:37:46 +02:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
Tony Wasserka f41501444d ThreadPoolAllocator: Add a dedicated interface to try reowning a buffer without fallback 2025-05-14 13:29:44 +02:00
Tony Wasserka 86b26b80ce FixedSizePooledAllocation: Clean up documentation 2025-05-14 13:29:44 +02:00
Ryan Houdek ae9a5b1125 Merge pull request #4487 from pmatos/GCCTargetTestsN1
Run gcc target tests with block size 1
2025-05-13 04:45:52 -07:00
LC 38579807c2 Merge pull request #4567 from neobrain/fix_glibcxx_debug
ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
2025-05-09 15:17:48 -04:00
Tony Wasserka c19119bcd6 ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
Default-constructed iterators can't be copied.
2025-05-09 15:02:49 +02:00
LC 47f1ad693d Merge pull request #4561 from Sonicadvance1/minor_vma_cleanup
LinuxEmulation: Minor cleanup by separating VMA definitions
2025-05-08 15:21:44 -04:00
Ryan Houdek 2164d7bb96 Merge pull request #4565 from neobrain/fix_async_timeout
FEXServer: Don't time out while clients are still connected
2025-05-08 12:14:56 -07:00
Tony Wasserka c326e2d669 FEXServer: Don't time out while clients are still connected 2025-05-08 11:56:19 +02:00
Tony Wasserka 8eaf45414c Async: Add run_one interface to enable more fine-grained event loop control 2025-05-08 11:56:19 +02:00
Ryan Houdek b3297d106e Merge pull request #4564 from bylaws/arm64ec
CMake: Allow disabling explicit -mcpu usage
2025-05-07 15:30:04 -07:00
Ryan Houdek 2d0e19e6a7 Merge pull request #4562 from WhatAmISupposedToPutHere/main
Windows: Fix building with llvm-libcxx
2025-05-07 15:06:22 -07:00
Billy Laws fd2ee4dc46 CMake: Allow disabling explicit -mcpu usage 2025-05-07 22:48:05 +01:00
Sasha Finkelstein b1bbc37c59 Windows: Fix building with llvm-libcxx
Libcxx uses GetSystemTimePreciseAsFileTime if _WIN32_WINNT specifies
a version new enough to have it.
2025-05-07 23:47:24 +02:00
Ryan Houdek 56409d4f2b LinuxEmulation: Minor cleanup by separating VMA definitions
NFC, just moving this to its own header. It's already a huge PITA to
read. I just want to try and preserve some sanity while attempting to
fix #4557
2025-05-07 13:55:44 -07:00
Paulo Matos ec0683b729 Run gcc target tests with block size 1
info files for the test runner like Disabled_Tests, Known_Failures,
etc, receive not just the filename but the test name (which is the test name
and potentially its running config).

Increase the timeout of gcc tests to 30secs.
2025-05-07 16:14:53 +02:00
LC 89e5041e70 Merge pull request #4559 from Sonicadvance1/we_require_more_machicolations!
LinuxEmulation: Implement custom longjump that is fortification safe
2025-05-07 09:07:12 -04:00
LC 8b3e7312e0 Merge pull request #4560 from Sonicadvance1/fix_bad_define_check
LinuxEmulation: Fix bad compile time definition check
2025-05-06 22:28:41 -04:00
Ryan Houdek 3c76a9176d LinuxEmulation: Fix bad compile time definition check
We don't want this to be compiled out if the definition doesn't exist.
Actually define it in that case.
2025-05-06 18:39:58 -07:00
Ryan Houdek a37def2c22 LinuxEmulation: Implement custom longjump that is fortification safe
With fortifications enabled, glibc long jump has some additional checks
in place that break because we do a stack pivot. The only way around
this is to do our own long jumps. Luckily this is trivial.

Fixes #4558
2025-05-06 15:29:42 -07:00
363 changed files with 31035 additions and 29330 deletions

No files matched your search

+2 -2
View File
@@ -32,7 +32,7 @@ AttributeMacros:
BinPackArguments: true
BinPackParameters: true
BitFieldColonSpacing: Both
BreakAfterAttributes: Always # clang 16 required
BreakAfterAttributes: Leave
BreakBeforeBraces: Attach
BreakBeforeBinaryOperators: None
BreakBeforeInlineASMColon: OnlyMultiline # clang 16 required
@@ -60,7 +60,7 @@ IndentRequires: false
IndentWidth: 2
InsertBraces: true
KeepEmptyLinesAtTheStartOfBlocks: true
LambdaBodyIndentation: OuterScope
LambdaBodyIndentation: Signature
LineEnding: LF # clang 16 required
MaxEmptyLinesToKeep: 2
NamespaceIndentation: Inner
-4
View File
@@ -1,8 +1,4 @@
# This file is used to ignore files and directories from clang-format
# Ignore all files in the External directory
External/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
+4
View File
@@ -16,3 +16,7 @@
# Reformat of CodeEmitter inl files
8760c593ece92d7e9fa94c40da0368fd367c9cad
# Whole-tree reformat with clang-format-19
5267cde60e7642852d18f20ae8568643bb5293d5
+1 -1
View File
@@ -78,7 +78,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+5 -11
View File
@@ -40,11 +40,8 @@ jobs:
echo "Formatting files:"
echo "$CHANGED_FILES"
- name: Check for correct clang-format version
run: clang-format --version | grep -qF '16.0.6'
- name: Check git-clang-format-16 exists
run: which git-clang-format-16
- name: Check git-clang-format-19 exists
run: which git-clang-format-19
- name: Setup Python env
uses: actions/setup-python@v4
@@ -58,19 +55,16 @@ jobs:
- name: Run code formatter
env:
CLANG_FORMAT_PATH: 'git-clang-format-16'
CLANG_FORMAT_PATH: 'git-clang-format-19'
GITHUB_PR_NUMBER: ${{ github.event.pull_request.number }}
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
# TODO(pmatos): Once we adopt v18, we should be able
# to take advantage of the new --diff_from_common_commit option
# explicitly in code-format-helper.py and not have to diff starting at
# the merge base.
# Using --diff_from_common_commit option available in clang-format-19
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
--issue-number $GITHUB_PR_NUMBER \
--start-rev $(git merge-base $START_REV $END_REV) \
--start-rev $START_REV \
--end-rev $END_REV \
--changed-files "$CHANGED_FILES"
+88
View File
@@ -0,0 +1,88 @@
name: Wine DLL artifacts
on:
push:
branches:
- main
env:
BUILD_TYPE: Release
jobs:
wine_dll_artifacts:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, mingw]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean install directory
run: |
rm -Rf ${{runner.workspace}}/build_install
mkdir ${{runner.workspace}}/build_install
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build_arm64ec
rm -Rf ${{runner.workspace}}/build_wow64
- name: Create Build Environment arm64ec
run: |
cmake -E make_directory ${{runner.workspace}}/build_arm64ec
cmake -E make_directory ${{runner.workspace}}/build_wow64
- name: Configure CMake arm64ec
shell: bash
working-directory: ${{runner.workspace}}/build_arm64ec
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Configure CMake wow64
shell: bash
working-directory: ${{runner.workspace}}/build_wow64
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Build arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Build wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: wine_dll_artifacts
path: ${{runner.workspace}}/build_install/usr/lib/wine/aarch64-windows/lib*.dll
retention-days: 60
compression-level: 9
+23 -11
View File
@@ -35,10 +35,19 @@ set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use fo
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
set (DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set (HOSTLIBS_DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
if (NOT DATA_DIRECTORY)
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
endif()
include(GNUInstallDirs)
if (NOT HOSTLIBS_DATA_DIRECTORY)
set(HOSTLIBS_DATA_DIRECTORY "${CMAKE_INSTALL_FULL_LIBDIR}/fex-emu")
endif()
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
@@ -94,7 +103,7 @@ endif()
# uninstall target
if(NOT TARGET uninstall)
configure_file(
"${CMAKE_CURRENT_SOURCE_DIR}/CMakeFiles/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_BINARY_DIR}/CMakeFiles/cmake_uninstall.cmake"
IMMEDIATE @ONLY)
@@ -403,6 +412,13 @@ if (TUNE_CPU STREQUAL "native")
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/NeedDisabledSVE.py"
RESULT_VARIABLE NEEDS_SVE_DISABLED)
if (NEEDS_SVE_DISABLED)
message(STATUS "Platform has bugged SVE. Disabling")
set(AARCH64_CPU "cortex-a78")
endif()
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${AARCH64_CPU}")
@@ -414,7 +430,7 @@ if (TUNE_CPU STREQUAL "native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=native")
endif()
endif()
else()
elseif (NOT TUNE_CPU STREQUAL "none")
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${TUNE_CPU}")
@@ -433,10 +449,6 @@ endif()
add_compile_options(-Wall)
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/ConfigDefines.h)
include(CTest)
if (BUILD_TESTS)
message(STATUS "Unit tests are enabled")
@@ -619,12 +631,12 @@ set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.com>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/CPack/Description.txt")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/CPack/triggers")
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
-3
View File
@@ -1,3 +0,0 @@
x86 and x86-64 Linux emulator
FEX is very much work in progress, so expect things to change.
File diff suppressed because it is too large. Load diff
+8
View File
@@ -86,6 +86,14 @@ constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t {8} << FEXCore::ToUnderlying(size);
}
// Many floating point operations constrain their element sizes to the
// main three float sizes half, single, and double precision. This just
// combines all the checks together for brevity.
[[nodiscard]]
constexpr bool IsStandardFloatSize(SubRegSize size) {
return size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit;
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
+2 -2
View File
@@ -1648,8 +1648,8 @@ public:
template<typename T>
void ASIMDLoadStoreSinglePost(uint32_t Op, uint32_t Q, uint32_t L, uint32_t R, uint32_t opcode, uint32_t S, uint32_t size,
ARMEmitter::Register rm, ARMEmitter::Register rn, T rt) {
LOGMAN_THROW_A_FMT(std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>, "Only supports 128-bit and "
"64-bit vector registers.");
LOGMAN_THROW_A_FMT((std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>), "Only supports 128-bit and "
"64-bit vector registers.");
uint32_t Instr = Op;
Instr |= Q << 30;
+21 -34
View File
@@ -60,8 +60,7 @@ public:
}
void fcmla(SubRegSize size, ZRegister zda, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcmla can only use p0 to p7");
uint32_t Op = 0b0110'0100'0000'0000'0000'0000'0000'0000;
@@ -76,8 +75,7 @@ public:
}
void fcadd(SubRegSize size, ZRegister zd, PRegisterMerge pv, ZRegister zn, ZRegister zm, Rotation rot) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pv <= PReg::p7.Merging(), "fcadd can only use p0 to p7");
LOGMAN_THROW_A_FMT(rot == Rotation::ROTATE_90 || rot == Rotation::ROTATE_270, "fcadd rotation may only be 90 or 270 degrees");
LOGMAN_THROW_A_FMT(zd == zn, "fcadd zd and zn must be the same register");
@@ -815,16 +813,12 @@ public:
// SVE Integer Misc - Unpredicated
// SVE floating-point trig select coefficient
void ftssel(SubRegSize size, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "ftssel may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "ftssel may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b00, zm.Idx(), FEXCore::ToUnderlying(size), zd, zn);
}
// SVE floating-point exponential accelerator
void fexpa(SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "fexpa may only have "
"16-bit, 32-bit, or 64-bit "
"element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "fexpa may only use 16/32/64-bit element sizes");
SVEIntegerMiscUnpredicated(0b10, 0b00000, FEXCore::ToUnderlying(size), zd, zn);
}
// SVE constructive prefix (unpredicated)
@@ -1503,9 +1497,9 @@ public:
}
// SVE broadcast floating-point immediate (unpredicated)
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(size == ARMEmitter::SubRegSize::i16Bit || size == ARMEmitter::SubRegSize::i32Bit || size == ARMEmitter::SubRegSize::i64Bit,
"Unsupported fmov size");
void fdup(SubRegSize size, ZRegister zd, float Value) {
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fmov size");
uint32_t Imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -1518,7 +1512,7 @@ public:
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
}
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
void fmov(SubRegSize size, ZRegister zd, float Value) {
fdup(size, zd, Value);
}
@@ -3401,8 +3395,8 @@ private:
}
void SVEBroadcastFloatImmPredicated(SubRegSize size, ZRegister zd, PRegister pg, float value) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported fcpy/fmov "
"size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported fcpy/fmov size");
uint32_t imm {};
if (size == SubRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
@@ -3578,7 +3572,7 @@ private:
// SVE2 floating-point pairwise operations
void SVEFloatPairwiseArithmetic(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zd needs to equal zn");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0100'0001'0000'1000'0000'0000'0000;
@@ -3592,7 +3586,7 @@ private:
// SVE floating-point arithmetic (unpredicated)
void SVEFloatArithmeticUnpredicated(uint32_t opc, SubRegSize size, ZRegister zm, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
uint32_t Instr = 0b0110'0101'0000'0000'0000'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -3700,7 +3694,7 @@ private:
// SVE floating-point arithmetic (predicated)
void SVEFloatArithmeticPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zd, ZRegister zn, ZRegister zm) {
LOGMAN_THROW_A_FMT(zd == zn, "zn needs to equal zd");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Invalid float size");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Invalid float size");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'1000'0000'0000'0000;
@@ -3728,9 +3722,7 @@ private:
}
void SVEFPRecursiveReduction(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "FP reduction operation can "
"only use 16-bit, 32-bit, "
"or 64-bit element sizes");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "FP reduction operation can only use 16/32/64-bit element sizes");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "FP reduction operation can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'0000'0010'0000'0000'0000;
@@ -4112,7 +4104,7 @@ private:
// 0b111 - I - Current
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'0000'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4721,7 +4713,7 @@ private:
void SVEFloatUnary(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zn, ZRegister zd) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "Unsupported size in {}", __func__);
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "Unsupported size in {}", __func__);
uint32_t Instr = 0b0110'0101'0000'1100'1010'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4809,8 +4801,7 @@ private:
}
void SVEFPUnaryOpsUnpredicated(uint32_t opc, SubRegSize size, ZRegister zd, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
uint32_t Instr = 0b0110'0101'0000'1000'0011'0000'0000'0000;
Instr |= FEXCore::ToUnderlying(size) << 22;
@@ -4821,8 +4812,7 @@ private:
}
void SVEFPSerialReductionPredicated(uint32_t opc, SubRegSize size, VRegister vd, PRegister pg, VRegister vn, ZRegister zm) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(vd == vn, "vn must be the same as vd");
@@ -4836,8 +4826,7 @@ private:
}
void SVEFPCompareWithZero(uint32_t eqlt, uint32_t ne, SubRegSize size, PRegister pd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0001'0000'0010'0000'0000'0000;
@@ -4852,8 +4841,7 @@ private:
void SVEFPMultiplyAdd(uint32_t opc, SubRegSize size, ZRegister zd, PRegister pg, ZRegister zn, ZRegister zm) {
// NOTE: opc also includes the op0 bit (bit 15) like op0:opc, since the fields are adjacent
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0010'0000'0000'0000'0000'0000;
@@ -4867,8 +4855,7 @@ private:
}
void SVEFPMultiplyAddIndexed(uint32_t op, SubRegSize size, ZRegister zda, ZRegister zn, ZRegister zm, uint32_t index) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit, "SubRegSize must be 16-bit, "
"32-bit, or 64-bit");
LOGMAN_THROW_A_FMT(IsStandardFloatSize(size), "SubRegSize must be 16-bit, 32-bit, or 64-bit");
LOGMAN_THROW_A_FMT((size <= SubRegSize::i32Bit && zm <= ZReg::z7) || (size == SubRegSize::i64Bit && zm <= ZReg::z15),
"16-bit and 32-bit indexed variants may only use Zm between z0-z7\n"
"64-bit variants may only use Zm between z0-z15");
+5 -7
View File
@@ -27,8 +27,6 @@ struct EmitterOps : Emitter {
public:
// Advanced SIMD scalar copy
void dup(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
constexpr uint32_t Op = 0b0101'1110'0000'0000'0000'01 << 10;
const uint32_t SizeImm = FEXCore::ToUnderlying(size);
const uint32_t IndexShift = SizeImm + 1;
const uint32_t ElementSize = 1U << SizeImm;
@@ -38,10 +36,10 @@ public:
const uint32_t imm5 = (Index << IndexShift) | ElementSize;
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
ASIMDScalarCopy(1, 1, imm5, 0b0000, rd, rn);
}
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
void mov(ScalarRegSize size, VRegister rd, VRegister rn, uint32_t Index) {
dup(size, rd, rn, Index);
}
@@ -1282,10 +1280,10 @@ public:
private:
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
void ASIMDScalarCopy(uint32_t Q, uint32_t b28, uint32_t imm5, uint32_t imm4, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0000'1110'0000'0000'0000'01U << 10;
Instr |= Q << 30;
Instr |= b28 << 28;
Instr |= imm5 << 16;
Instr |= imm4 << 11;
Instr |= Encode_rn(rn);
+5 -1
View File
@@ -224,9 +224,13 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// 11110s 2 UInt(s)
//
// So we 'or' (2 * -d) with our computed s to form imms.
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
if (n != nullptr) {
*n = out_n;
}
if (imm_s != nullptr) {
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
}
if (imm_r != nullptr) {
*imm_r = r;
}
File renamed without changes.
File renamed without changes.
File renamed without changes.
+3
View File
@@ -0,0 +1,3 @@
x86 and x86-64 Linux emulator
FEX allows you to run x86 applications on ARM64 Linux devices. It offers broad compatibility with both 32-bit and 64-bit binaries, and it can be used alongside Wine/Proton to play Windows games.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
@@ -1,6 +1,6 @@
# This is a reference AArch64 cross compile script
# Pass in to cmake when building:
# eg: cmake -DCMAKE_TOOLCHAIN_FILE=../CMakeToolchains/AArch64.cmake ..
# eg: cmake --toolchain ../Data/CMake/toolchain_aarch64.cmake ..
if (NOT DEFINED ENV{SYSROOT})
message(FATAL_ERROR "Need to have SYSROOT environment variable set")
endif()
@@ -4,6 +4,7 @@ set(CMAKE_RC_COMPILER ${MINGW_TRIPLE}-windres)
set(CMAKE_C_COMPILER ${MINGW_TRIPLE}-clang)
set(CMAKE_CXX_COMPILER ${MINGW_TRIPLE}-clang++)
set(CMAKE_DLLTOOL ${MINGW_TRIPLE}-dlltool)
set(CMAKE_AR ${MINGW_TRIPLE}-ar)
# Compile everything as static to avoid requiring the MinGW runtime libraries, force page aligned sections so that
# debug symbols work correctly, and disable loop alignment to workaround an LLVM bug
File renamed without changes.
File renamed without changes.
File renamed without changes.
View File
File renamed without changes.
+35
View File
@@ -0,0 +1,35 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
gcc32 = pkgs.writeText "toolchain_nix_gcc_x86_32.txt" ''
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-g++)
'';
gcc64 = pkgs.writeText "toolchain_nix_gcc_x86_64.txt" ''
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-g++)
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "toolchain32: ${gcc32}"
echo "toolchain64: ${gcc64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${gcc32} -DX86_64_TOOLCHAIN_FILE=${gcc64}";
}
+83
View File
@@ -0,0 +1,83 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
devRootFS = pkgs.buildEnv {
name = "fex-dev-rootfs";
paths = [
pkgsCross64.stdenv.cc.libc_dev
pkgsCross32.stdenv.cc.libc_dev
pkgsCross64.stdenv.cc.cc
pkgsCross32.stdenv.cc.cc
pkgs.alsa-lib.dev
pkgs.libdrm.dev
pkgs.libGL.dev
pkgs.wayland.dev
pkgs.xorg.libX11.dev
pkgs.xorg.libxcb.dev
pkgs.xorg.libXrandr.dev
pkgs.xorg.libXrender.dev
pkgs.xorg.xorgproto
];
ignoreCollisions = true;
pathsToLink = [
"/include"
"/lib"
];
postBuild = ''
mkdir -p $out/usr
ln -s $out/include $out/usr/
'';
};
toolchain32 = pkgs.writeText "toolchain_nix_x86_32.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target i686-linux-gnu -msse2 -mfpmath=sse --sysroot=${devRootFS} -iwithsysroot/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
toolchain64 = pkgs.writeText "toolchain_nix_x86_64.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target x86_64-linux-gnu --sysroot=${devRootFS} -iwithsysroot/usr/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "Set up dev RootFS at ${devRootFS}"
echo "toolchain32: ${toolchain32}"
echo "toolchain64: ${toolchain64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${toolchain32} -DX86_64_TOOLCHAIN_FILE=${toolchain64} -DX86_DEV_ROOTFS=${devRootFS}";
ROOTFS = "${devRootFS}";
}
+52
View File
@@ -0,0 +1,52 @@
{ pkgs ? import <nixpkgs> { } }:
let
toolchain = pkgs.fetchzip {
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250305/llvm-mingw-20250305-ucrt-ubuntu-20.04-aarch64.tar.xz";
sha256 = "sha256-cA03/ab9O61eO9+S2JzIXD4V0HzTXK5/AYyxW2d73Po=";
};
cmakeToolchainFile = pkgs.substitute {
# Use absolute paths that are discoverable outside of the nix shell
src = ../../CMake/toolchain_mingw.cmake;
substitutions = ["--replace-fail" "\${MINGW_TRIPLE}-" "${toolchain}/bin/\${MINGW_TRIPLE}-"];
};
mesonCrossFile = pkgs.writeText "crossfile_llvm_mingw.txt" ''
[binaries]
ar = '${toolchain}/bin/arm64ec-w64-mingw32-ar'
c = '${toolchain}/bin/arm64ec-w64-mingw32-gcc'
cpp = '${toolchain}/bin/arm64ec-w64-mingw32-g++'
ld = '${toolchain}/bin/arm64ec-w64-mingw32-ld'
windres = '${toolchain}/bin/arm64ec-w64-mingw32-windres'
strip = '${toolchain}/bin/strip'
widl = '${toolchain}/bin/arm64ec-w64-mingw32-widl'
pkgconfig = 'aarch64-linux-gnu-pkg-config'
[host_machine]
system = 'windows'
cpu_family = 'aarch64'
cpu = 'aarch64'
endian = 'little'
'';
in
pkgs.mkShell {
buildInputs = [
toolchain
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "llvm-mingw set up at ${toolchain}."
echo ""
echo "To configure DXVK/vkd3d-proton: meson setup \$FEX_MESON_CROSSFILE"
echo ""
echo "To configure 32-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_WOW64"
echo "To configure 64-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_ARM64EC"
fi
'';
# E.g. cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False
FEX_CMAKE_TOOLCHAIN_ARM64EC = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_CMAKE_TOOLCHAIN_WOW64 = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_MESON_CROSSFILE = "--cross-file ${mesonCrossFile}";
}
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 32-bit applications in Wine/Proton.
# The required cross-toolchains will be set up and managed by nix.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_WOW64 -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 64-bit applications in Wine/Proton
# Nix is used to install and manage the required cross-toolchains.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+17
View File
@@ -0,0 +1,17 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash FEXLinuxTests/shell.nix
# Helper script to configure CMake for building FEXLinuxTests.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf unittests/FEXLinuxTests
set -o xtrace
cmake . $FEX_CMAKE_TOOLCHAINS -DBUILD_TESTS=ON -DBUILD_FEX_LINUX_TESTS=ON
+22
View File
@@ -0,0 +1,22 @@
# Helper script to configure CMake for library forwarding in FEX.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf guest-libs guest-libs-32 Guest Guest_32
# Set clang executable path manually since the one from the nix store
# will be picked up otherwise
CLANG_EXEC_PATH=""
if ! grep -q CLANG_EXEC_PATH CMakeCache.txt
then
CLANG_EXEC_PATH="-DCLANG_EXEC_PATH=`which clang`"
fi
nix-shell `dirname -- "$0"`/LibraryForwarding/shell.nix \
--run "set -o xtrace; cmake . \$FEX_CMAKE_TOOLCHAINS -DBUILD_THUNKS=ON $CLANG_EXEC_PATH; set +o xtrace"
+1
View File
@@ -0,0 +1 @@
DisableFormat: true
+4 -8
View File
@@ -169,14 +169,9 @@ View the diff from {self.name} here.
class ClangFormatHelper(FormatHelper):
name = "clang-format"
name = "git-clang-format"
friendly_name = "C/C++ code formatter"
@property
def cformat_wrapper_path(self) -> str:
relpath = "../../Scripts/clang-format.py"
curpath = os.path.dirname(os.path.abspath(__file__))
return os.path.abspath(os.path.normpath(os.path.join(curpath, relpath)))
@property
def instructions(self) -> str:
@@ -199,7 +194,7 @@ class ClangFormatHelper(FormatHelper):
def clang_fmt_path(self) -> str:
if "CLANG_FORMAT_PATH" in os.environ:
return os.environ["CLANG_FORMAT_PATH"]
return "git-clang-format"
return "git-clang-format-19"
def has_tool(self) -> bool:
cmd = [self.clang_fmt_path, "-h"]
@@ -217,8 +212,9 @@ class ClangFormatHelper(FormatHelper):
cf_cmd = [
self.clang_fmt_path,
f"--binary={self.cformat_wrapper_path}",
"--binary=clang-format-19",
"--diff",
"--diff_from_common_commit",
]
if args.start_rev and args.end_rev:
+378 -38
View File
@@ -1,52 +1,392 @@
#
# This file is autogenerated by pip-compile with Python 3.11
# This file is autogenerated by pip-compile with Python 3.13
# by the following command:
#
# pip-compile --output-file=llvm/utils/git/requirements_formatting.txt llvm/utils/git/requirements_formatting.txt.in
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==23.9.1
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
# via
# -r llvm/utils/git/requirements_formatting.txt.in
# -r requirements_formatting.txt.in
# darker
certifi==2023.7.22
# via requests
cffi==1.15.1
certifi==2025.7.14 \
--hash=sha256:6b31f564a415d79ee77df69d757bb49a5bb53bd9f756cbbe24394ffd6fc1f4b2 \
--hash=sha256:8ea99dbdfaaf2ba2f9bac77b9249ef62ec5218e7c2b2e903378ed5fccf765995
# via
# -r requirements_formatting.txt.in
# requests
cffi==1.15.1 \
--hash=sha256:00a9ed42e88df81ffae7a8ab6d9356b371399b91dbdf0c3cb1e84c03a13aceb5 \
--hash=sha256:03425bdae262c76aad70202debd780501fabeaca237cdfddc008987c0e0f59ef \
--hash=sha256:04ed324bda3cda42b9b695d51bb7d54b680b9719cfab04227cdd1e04e5de3104 \
--hash=sha256:0e2642fe3142e4cc4af0799748233ad6da94c62a8bec3a6648bf8ee68b1c7426 \
--hash=sha256:173379135477dc8cac4bc58f45db08ab45d228b3363adb7af79436135d028405 \
--hash=sha256:198caafb44239b60e252492445da556afafc7d1e3ab7a1fb3f0584ef6d742375 \
--hash=sha256:1e74c6b51a9ed6589199c787bf5f9875612ca4a8a0785fb2d4a84429badaf22a \
--hash=sha256:2012c72d854c2d03e45d06ae57f40d78e5770d252f195b93f581acf3ba44496e \
--hash=sha256:21157295583fe8943475029ed5abdcf71eb3911894724e360acff1d61c1d54bc \
--hash=sha256:2470043b93ff09bf8fb1d46d1cb756ce6132c54826661a32d4e4d132e1977adf \
--hash=sha256:285d29981935eb726a4399badae8f0ffdff4f5050eaa6d0cfc3f64b857b77185 \
--hash=sha256:30d78fbc8ebf9c92c9b7823ee18eb92f2e6ef79b45ac84db507f52fbe3ec4497 \
--hash=sha256:320dab6e7cb2eacdf0e658569d2575c4dad258c0fcc794f46215e1e39f90f2c3 \
--hash=sha256:33ab79603146aace82c2427da5ca6e58f2b3f2fb5da893ceac0c42218a40be35 \
--hash=sha256:3548db281cd7d2561c9ad9984681c95f7b0e38881201e157833a2342c30d5e8c \
--hash=sha256:3799aecf2e17cf585d977b780ce79ff0dc9b78d799fc694221ce814c2c19db83 \
--hash=sha256:39d39875251ca8f612b6f33e6b1195af86d1b3e60086068be9cc053aa4376e21 \
--hash=sha256:3b926aa83d1edb5aa5b427b4053dc420ec295a08e40911296b9eb1b6170f6cca \
--hash=sha256:3bcde07039e586f91b45c88f8583ea7cf7a0770df3a1649627bf598332cb6984 \
--hash=sha256:3d08afd128ddaa624a48cf2b859afef385b720bb4b43df214f85616922e6a5ac \
--hash=sha256:3eb6971dcff08619f8d91607cfc726518b6fa2a9eba42856be181c6d0d9515fd \
--hash=sha256:40f4774f5a9d4f5e344f31a32b5096977b5d48560c5592e2f3d2c4374bd543ee \
--hash=sha256:4289fc34b2f5316fbb762d75362931e351941fa95fa18789191b33fc4cf9504a \
--hash=sha256:470c103ae716238bbe698d67ad020e1db9d9dba34fa5a899b5e21577e6d52ed2 \
--hash=sha256:4f2c9f67e9821cad2e5f480bc8d83b8742896f1242dba247911072d4fa94c192 \
--hash=sha256:50a74364d85fd319352182ef59c5c790484a336f6db772c1a9231f1c3ed0cbd7 \
--hash=sha256:54a2db7b78338edd780e7ef7f9f6c442500fb0d41a5a4ea24fff1c929d5af585 \
--hash=sha256:5635bd9cb9731e6d4a1132a498dd34f764034a8ce60cef4f5319c0541159392f \
--hash=sha256:59c0b02d0a6c384d453fece7566d1c7e6b7bae4fc5874ef2ef46d56776d61c9e \
--hash=sha256:5d598b938678ebf3c67377cdd45e09d431369c3b1a5b331058c338e201f12b27 \
--hash=sha256:5df2768244d19ab7f60546d0c7c63ce1581f7af8b5de3eb3004b9b6fc8a9f84b \
--hash=sha256:5ef34d190326c3b1f822a5b7a45f6c4535e2f47ed06fec77d3d799c450b2651e \
--hash=sha256:6975a3fac6bc83c4a65c9f9fcab9e47019a11d3d2cf7f3c0d03431bf145a941e \
--hash=sha256:6c9a799e985904922a4d207a94eae35c78ebae90e128f0c4e521ce339396be9d \
--hash=sha256:70df4e3b545a17496c9b3f41f5115e69a4f2e77e94e1d2a8e1070bc0c38c8a3c \
--hash=sha256:7473e861101c9e72452f9bf8acb984947aa1661a7704553a9f6e4baa5ba64415 \
--hash=sha256:8102eaf27e1e448db915d08afa8b41d6c7ca7a04b7d73af6514df10a3e74bd82 \
--hash=sha256:87c450779d0914f2861b8526e035c5e6da0a3199d8f1add1a665e1cbc6fc6d02 \
--hash=sha256:8b7ee99e510d7b66cdb6c593f21c043c248537a32e0bedf02e01e9553a172314 \
--hash=sha256:91fc98adde3d7881af9b59ed0294046f3806221863722ba7d8d120c575314325 \
--hash=sha256:94411f22c3985acaec6f83c6df553f2dbe17b698cc7f8ae751ff2237d96b9e3c \
--hash=sha256:98d85c6a2bef81588d9227dde12db8a7f47f639f4a17c9ae08e773aa9c697bf3 \
--hash=sha256:9ad5db27f9cabae298d151c85cf2bad1d359a1b9c686a275df03385758e2f914 \
--hash=sha256:a0b71b1b8fbf2b96e41c4d990244165e2c9be83d54962a9a1d118fd8657d2045 \
--hash=sha256:a0f100c8912c114ff53e1202d0078b425bee3649ae34d7b070e9697f93c5d52d \
--hash=sha256:a591fe9e525846e4d154205572a029f653ada1a78b93697f3b5a8f1f2bc055b9 \
--hash=sha256:a5c84c68147988265e60416b57fc83425a78058853509c1b0629c180094904a5 \
--hash=sha256:a66d3508133af6e8548451b25058d5812812ec3798c886bf38ed24a98216fab2 \
--hash=sha256:a8c4917bd7ad33e8eb21e9a5bbba979b49d9a97acb3a803092cbc1133e20343c \
--hash=sha256:b3bbeb01c2b273cca1e1e0c5df57f12dce9a4dd331b4fa1635b8bec26350bde3 \
--hash=sha256:cba9d6b9a7d64d4bd46167096fc9d2f835e25d7e4c121fb2ddfc6528fb0413b2 \
--hash=sha256:cc4d65aeeaa04136a12677d3dd0b1c0c94dc43abac5860ab33cceb42b801c1e8 \
--hash=sha256:ce4bcc037df4fc5e3d184794f27bdaab018943698f4ca31630bc7f84a7b69c6d \
--hash=sha256:cec7d9412a9102bdc577382c3929b337320c4c4c4849f2c5cdd14d7368c5562d \
--hash=sha256:d400bfb9a37b1351253cb402671cea7e89bdecc294e8016a707f6d1d8ac934f9 \
--hash=sha256:d61f4695e6c866a23a21acab0509af1cdfd2c013cf256bbf5b6b5e2695827162 \
--hash=sha256:db0fbb9c62743ce59a9ff687eb5f4afbe77e5e8403d6697f7446e5f609976f76 \
--hash=sha256:dd86c085fae2efd48ac91dd7ccffcfc0571387fe1193d33b6394db7ef31fe2a4 \
--hash=sha256:e00b098126fd45523dd056d2efba6c5a63b71ffe9f2bbe1a4fe1716e1d0c331e \
--hash=sha256:e229a521186c75c8ad9490854fd8bbdd9a0c9aa3a524326b55be83b54d4e0ad9 \
--hash=sha256:e263d77ee3dd201c3a142934a086a4450861778baaeeb45db4591ef65550b0a6 \
--hash=sha256:ed9cb427ba5504c1dc15ede7d516b84757c3e3d7868ccc85121d9310d27eed0b \
--hash=sha256:fa6693661a4c91757f4412306191b6dc88c1703f780c8234035eac011922bc01 \
--hash=sha256:fcd131dd944808b5bdb38e6f5b53013c5aa4f334c5cad0c72742f6eba4b73db0
# via
# cryptography
# pynacl
charset-normalizer==3.2.0
charset-normalizer==3.2.0 \
--hash=sha256:04e57ab9fbf9607b77f7d057974694b4f6b142da9ed4a199859d9d4d5c63fe96 \
--hash=sha256:09393e1b2a9461950b1c9a45d5fd251dc7c6f228acab64da1c9c0165d9c7765c \
--hash=sha256:0b87549028f680ca955556e3bd57013ab47474c3124dc069faa0b6545b6c9710 \
--hash=sha256:1000fba1057b92a65daec275aec30586c3de2401ccdcd41f8a5c1e2c87078706 \
--hash=sha256:1249cbbf3d3b04902ff081ffbb33ce3377fa6e4c7356f759f3cd076cc138d020 \
--hash=sha256:1920d4ff15ce893210c1f0c0e9d19bfbecb7983c76b33f046c13a8ffbd570252 \
--hash=sha256:193cbc708ea3aca45e7221ae58f0fd63f933753a9bfb498a3b474878f12caaad \
--hash=sha256:1a100c6d595a7f316f1b6f01d20815d916e75ff98c27a01ae817439ea7726329 \
--hash=sha256:1f30b48dd7fa1474554b0b0f3fdfdd4c13b5c737a3c6284d3cdc424ec0ffff3a \
--hash=sha256:203f0c8871d5a7987be20c72442488a0b8cfd0f43b7973771640fc593f56321f \
--hash=sha256:246de67b99b6851627d945db38147d1b209a899311b1305dd84916f2b88526c6 \
--hash=sha256:2dee8e57f052ef5353cf608e0b4c871aee320dd1b87d351c28764fc0ca55f9f4 \
--hash=sha256:2efb1bd13885392adfda4614c33d3b68dee4921fd0ac1d3988f8cbb7d589e72a \
--hash=sha256:2f4ac36d8e2b4cc1aa71df3dd84ff8efbe3bfb97ac41242fbcfc053c67434f46 \
--hash=sha256:3170c9399da12c9dc66366e9d14da8bf7147e1e9d9ea566067bbce7bb74bd9c2 \
--hash=sha256:3b1613dd5aee995ec6d4c69f00378bbd07614702a315a2cf6c1d21461fe17c23 \
--hash=sha256:3bb3d25a8e6c0aedd251753a79ae98a093c7e7b471faa3aa9a93a81431987ace \
--hash=sha256:3bb7fda7260735efe66d5107fb7e6af6a7c04c7fce9b2514e04b7a74b06bf5dd \
--hash=sha256:41b25eaa7d15909cf3ac4c96088c1f266a9a93ec44f87f1d13d4a0e86c81b982 \
--hash=sha256:45de3f87179c1823e6d9e32156fb14c1927fcc9aba21433f088fdfb555b77c10 \
--hash=sha256:46fb8c61d794b78ec7134a715a3e564aafc8f6b5e338417cb19fe9f57a5a9bf2 \
--hash=sha256:48021783bdf96e3d6de03a6e39a1171ed5bd7e8bb93fc84cc649d11490f87cea \
--hash=sha256:4957669ef390f0e6719db3613ab3a7631e68424604a7b448f079bee145da6e09 \
--hash=sha256:5e86d77b090dbddbe78867a0275cb4df08ea195e660f1f7f13435a4649e954e5 \
--hash=sha256:6339d047dab2780cc6220f46306628e04d9750f02f983ddb37439ca47ced7149 \
--hash=sha256:681eb3d7e02e3c3655d1b16059fbfb605ac464c834a0c629048a30fad2b27489 \
--hash=sha256:6c409c0deba34f147f77efaa67b8e4bb83d2f11c8806405f76397ae5b8c0d1c9 \
--hash=sha256:7095f6fbfaa55defb6b733cfeb14efaae7a29f0b59d8cf213be4e7ca0b857b80 \
--hash=sha256:70c610f6cbe4b9fce272c407dd9d07e33e6bf7b4aa1b7ffb6f6ded8e634e3592 \
--hash=sha256:72814c01533f51d68702802d74f77ea026b5ec52793c791e2da806a3844a46c3 \
--hash=sha256:7a4826ad2bd6b07ca615c74ab91f32f6c96d08f6fcc3902ceeedaec8cdc3bcd6 \
--hash=sha256:7c70087bfee18a42b4040bb9ec1ca15a08242cf5867c58726530bdf3945672ed \
--hash=sha256:855eafa5d5a2034b4621c74925d89c5efef61418570e5ef9b37717d9c796419c \
--hash=sha256:8700f06d0ce6f128de3ccdbc1acaea1ee264d2caa9ca05daaf492fde7c2a7200 \
--hash=sha256:89f1b185a01fe560bc8ae5f619e924407efca2191b56ce749ec84982fc59a32a \
--hash=sha256:8b2c760cfc7042b27ebdb4a43a4453bd829a5742503599144d54a032c5dc7e9e \
--hash=sha256:8c2f5e83493748286002f9369f3e6607c565a6a90425a3a1fef5ae32a36d749d \
--hash=sha256:8e098148dd37b4ce3baca71fb394c81dc5d9c7728c95df695d2dca218edf40e6 \
--hash=sha256:94aea8eff76ee6d1cdacb07dd2123a68283cb5569e0250feab1240058f53b623 \
--hash=sha256:95eb302ff792e12aba9a8b8f8474ab229a83c103d74a750ec0bd1c1eea32e669 \
--hash=sha256:9bd9b3b31adcb054116447ea22caa61a285d92e94d710aa5ec97992ff5eb7cf3 \
--hash=sha256:9e608aafdb55eb9f255034709e20d5a83b6d60c054df0802fa9c9883d0a937aa \
--hash=sha256:a103b3a7069b62f5d4890ae1b8f0597618f628b286b03d4bc9195230b154bfa9 \
--hash=sha256:a386ebe437176aab38c041de1260cd3ea459c6ce5263594399880bbc398225b2 \
--hash=sha256:a38856a971c602f98472050165cea2cdc97709240373041b69030be15047691f \
--hash=sha256:a401b4598e5d3f4a9a811f3daf42ee2291790c7f9d74b18d75d6e21dda98a1a1 \
--hash=sha256:a7647ebdfb9682b7bb97e2a5e7cb6ae735b1c25008a70b906aecca294ee96cf4 \
--hash=sha256:aaf63899c94de41fe3cf934601b0f7ccb6b428c6e4eeb80da72c58eab077b19a \
--hash=sha256:b0dac0ff919ba34d4df1b6131f59ce95b08b9065233446be7e459f95554c0dc8 \
--hash=sha256:baacc6aee0b2ef6f3d308e197b5d7a81c0e70b06beae1f1fcacffdbd124fe0e3 \
--hash=sha256:bf420121d4c8dce6b889f0e8e4ec0ca34b7f40186203f06a946fa0276ba54029 \
--hash=sha256:c04a46716adde8d927adb9457bbe39cf473e1e2c2f5d0a16ceb837e5d841ad4f \
--hash=sha256:c0b21078a4b56965e2b12f247467b234734491897e99c1d51cee628da9786959 \
--hash=sha256:c1c76a1743432b4b60ab3358c937a3fe1341c828ae6194108a94c69028247f22 \
--hash=sha256:c4983bf937209c57240cff65906b18bb35e64ae872da6a0db937d7b4af845dd7 \
--hash=sha256:c4fb39a81950ec280984b3a44f5bd12819953dc5fa3a7e6fa7a80db5ee853952 \
--hash=sha256:c57921cda3a80d0f2b8aec7e25c8aa14479ea92b5b51b6876d975d925a2ea346 \
--hash=sha256:c8063cf17b19661471ecbdb3df1c84f24ad2e389e326ccaf89e3fb2484d8dd7e \
--hash=sha256:ccd16eb18a849fd8dcb23e23380e2f0a354e8daa0c984b8a732d9cfaba3a776d \
--hash=sha256:cd6dbe0238f7743d0efe563ab46294f54f9bc8f4b9bcf57c3c666cc5bc9d1299 \
--hash=sha256:d62e51710986674142526ab9f78663ca2b0726066ae26b78b22e0f5e571238dd \
--hash=sha256:db901e2ac34c931d73054d9797383d0f8009991e723dab15109740a63e7f902a \
--hash=sha256:e03b8895a6990c9ab2cdcd0f2fe44088ca1c65ae592b8f795c3294af00a461c3 \
--hash=sha256:e1c8a2f4c69e08e89632defbfabec2feb8a8d99edc9f89ce33c4b9e36ab63037 \
--hash=sha256:e4b749b9cc6ee664a3300bb3a273c1ca8068c46be705b6c31cf5d276f8628a94 \
--hash=sha256:e6a5bf2cba5ae1bb80b154ed68a3cfa2fa00fde979a7f50d6598d3e17d9ac20c \
--hash=sha256:e857a2232ba53ae940d3456f7533ce6ca98b81917d47adc3c7fd55dad8fab858 \
--hash=sha256:ee4006268ed33370957f55bf2e6f4d263eaf4dc3cfc473d1d90baff6ed36ce4a \
--hash=sha256:eef9df1eefada2c09a5e7a40991b9fc6ac6ef20b1372abd48d2794a316dc0449 \
--hash=sha256:f058f6963fd82eb143c692cecdc89e075fa0828db2e5b291070485390b2f1c9c \
--hash=sha256:f25c229a6ba38a35ae6e25ca1264621cc25d4d38dca2942a7fce0b67a4efe918 \
--hash=sha256:f2a1d0fd4242bd8643ce6f98927cf9c04540af6efa92323e9d3124f57727bfc1 \
--hash=sha256:f7560358a6811e52e9c4d142d497f1a6e10103d3a6881f18d04dbce3729c0e2c \
--hash=sha256:f779d3ad205f108d14e99bb3859aa7dd8e9c68874617c72354d7ecaec2a054ac \
--hash=sha256:f87f746ee241d30d6ed93969de31e5ffd09a2961a051e60ae6bddde9ec3583aa
# via requests
click==8.1.7
click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==41.0.3
# via pyjwt
darker==1.7.2
# via -r llvm/utils/git/requirements_formatting.txt.in
deprecated==1.2.14
cryptography==45.0.5 \
--hash=sha256:0027d566d65a38497bc37e0dd7c2f8ceda73597d2ac9ba93810204f56f52ebc7 \
--hash=sha256:101ee65078f6dd3e5a028d4f19c07ffa4dd22cce6a20eaa160f8b5219911e7d8 \
--hash=sha256:12e55281d993a793b0e883066f590c1ae1e802e3acb67f8b442e721e475e6463 \
--hash=sha256:14d96584701a887763384f3c47f0ca7c1cce322aa1c31172680eb596b890ec30 \
--hash=sha256:1e1da5accc0c750056c556a93c3e9cb828970206c68867712ca5805e46dc806f \
--hash=sha256:206210d03c1193f4e1ff681d22885181d47efa1ab3018766a7b32a7b3d6e6afd \
--hash=sha256:2089cc8f70a6e454601525e5bf2779e665d7865af002a5dec8d14e561002e135 \
--hash=sha256:3a264aae5f7fbb089dbc01e0242d3b67dffe3e6292e1f5182122bdf58e65215d \
--hash=sha256:3af26738f2db354aafe492fb3869e955b12b2ef2e16908c8b9cb928128d42c57 \
--hash=sha256:3fcfbefc4a7f332dece7272a88e410f611e79458fab97b5efe14e54fe476f4fd \
--hash=sha256:460f8c39ba66af7db0545a8c6f2eabcbc5a5528fc1cf6c3fa9a1e44cec33385e \
--hash=sha256:57c816dfbd1659a367831baca4b775b2a5b43c003daf52e9d57e1d30bc2e1b0e \
--hash=sha256:5aa1e32983d4443e310f726ee4b071ab7569f58eedfdd65e9675484a4eb67bd1 \
--hash=sha256:6ff8728d8d890b3dda5765276d1bc6fb099252915a2cd3aff960c4c195745dd0 \
--hash=sha256:7259038202a47fdecee7e62e0fd0b0738b6daa335354396c6ddebdbe1206af2a \
--hash=sha256:72e76caa004ab63accdf26023fccd1d087f6d90ec6048ff33ad0445abf7f605a \
--hash=sha256:7760c1c2e1a7084153a0f68fab76e754083b126a47d0117c9ed15e69e2103492 \
--hash=sha256:8c4a6ff8a30e9e3d38ac0539e9a9e02540ab3f827a3394f8852432f6b0ea152e \
--hash=sha256:9024beb59aca9d31d36fcdc1604dd9bbeed0a55bface9f1908df19178e2f116e \
--hash=sha256:90cb0a7bb35959f37e23303b7eed0a32280510030daba3f7fdfbb65defde6a97 \
--hash=sha256:91098f02ca81579c85f66df8a588c78f331ca19089763d733e34ad359f474174 \
--hash=sha256:926c3ea71a6043921050eaa639137e13dbe7b4ab25800932a8498364fc1abec9 \
--hash=sha256:982518cd64c54fcada9d7e5cf28eabd3ee76bd03ab18e08a48cad7e8b6f31b18 \
--hash=sha256:9b4cf6318915dccfe218e69bbec417fdd7c7185aa7aab139a2c0beb7468c89f0 \
--hash=sha256:ad0caded895a00261a5b4aa9af828baede54638754b51955a0ac75576b831b27 \
--hash=sha256:b85980d1e345fe769cfc57c57db2b59cff5464ee0c045d52c0df087e926fbe63 \
--hash=sha256:b8fa8b0a35a9982a3c60ec79905ba5bb090fc0b9addcfd3dc2dd04267e45f25e \
--hash=sha256:b9e38e0a83cd51e07f5a48ff9691cae95a79bea28fe4ded168a8e5c6c77e819d \
--hash=sha256:bd4c45986472694e5121084c6ebbd112aa919a25e783b87eb95953c9573906d6 \
--hash=sha256:be97d3a19c16a9be00edf79dca949c8fa7eff621763666a145f9f9535a5d7f42 \
--hash=sha256:c648025b6840fe62e57107e0a25f604db740e728bd67da4f6f060f03017d5097 \
--hash=sha256:d05a38884db2ba215218745f0781775806bde4f32e07b135348355fe8e4991d9 \
--hash=sha256:dd420e577921c8c2d31289536c386aaa30140b473835e97f83bc71ea9d2baf2d \
--hash=sha256:e357286c1b76403dd384d938f93c46b2b058ed4dfcdce64a770f0537ed3feb6f \
--hash=sha256:e6c00130ed423201c5bc5544c23359141660b07999ad82e34e7bb8f882bb78e0 \
--hash=sha256:e74d30ec9c7cb2f404af331d5b4099a9b322a8a6b25c4632755c8757345baac5 \
--hash=sha256:f3562c2f23c612f2e4a6964a61d942f891d29ee320edb62ff48ffb99f3de9ae8
# via
# -r requirements_formatting.txt.in
# pyjwt
darker==2.1.1 \
--hash=sha256:a6e6a682c0604e76fe9aec7650e96a944f517563c69b28fcc076db9d957d98ea \
--hash=sha256:ead701414c45359fc0312bc285614d3285fc135476d43f3bc08d989ee19d9020
# via -r requirements_formatting.txt.in
darkgraylib==1.2.1 \
--hash=sha256:60c59de69842367ce0c78c32c451fa8e9d29500e681312d9864a7416bcdb7792 \
--hash=sha256:a5dd6a2015a470d9047278cdd01a91ccb1d746675f8fd4562b3b5f6b8cbda930
# via
# darker
# graylint
deprecated==1.2.14 \
--hash=sha256:6fac8b097794a90302bdbb17b9b815e732d3c4720583ff1b198499d78470466c \
--hash=sha256:e5323eb936458dccc2582dc6f9c322c852a775a27065ff2b0c4970b9d53d01b3
# via pygithub
idna==3.4
# via requests
mypy-extensions==1.0.0
# via black
packaging==23.1
# via black
pathspec==0.11.2
# via black
platformdirs==3.10.0
# via black
pycparser==2.21
# via cffi
pygithub==1.59.1
# via -r llvm/utils/git/requirements_formatting.txt.in
pyjwt[crypto]==2.8.0
# via pygithub
pynacl==1.5.0
# via pygithub
requests==2.31.0
# via pygithub
toml==0.10.2
graylint==1.1.1 \
--hash=sha256:0fd8e02972ca03d0ef2bf0adea76b5343efcd492d7afb5f658f3e3a724f55a36 \
--hash=sha256:b7e0eab6c159684dbf5ef84e942c3340f6a6549b02a3d11b1a1763cc4f8f0593
# via darker
urllib3==2.0.4
# via requests
wrapt==1.15.0
idna==3.10 \
--hash=sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9 \
--hash=sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3
# via
# -r requirements_formatting.txt.in
# requests
mypy-extensions==1.0.0 \
--hash=sha256:4392f6c0eb8a5668a69e23d168ffa70f0be9ccfd32b5cc2d26a34ae5b844552d \
--hash=sha256:75dbf8955dc00442a438fc4d0666508a9a97b6bd41aa2f0ffe9d2f2725af0782
# via black
packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
--hash=sha256:d7c24979f292f916dc9cbf8648319032f551ea8c49a4c9bf2fb556a02070ec1d
# via black
pycparser==2.21 \
--hash=sha256:8ee45429555515e1f6b185e78100aea234072576aa43ab53aefcae078162fca9 \
--hash=sha256:e644fdec12f7872f86c58ff790da456218b10f863970249516d60a5eaca77206
# via cffi
pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pynacl==1.5.0 \
--hash=sha256:06b8f6fa7f5de8d5d2f7573fe8c863c051225a27b61e6860fd047b1775807858 \
--hash=sha256:0c84947a22519e013607c9be43706dd42513f9e6ae5d39d3613ca1e142fba44d \
--hash=sha256:20f42270d27e1b6a29f54032090b972d97f0a1b0948cc52392041ef7831fee93 \
--hash=sha256:401002a4aaa07c9414132aaed7f6836ff98f59277a234704ff66878c2ee4a0d1 \
--hash=sha256:52cb72a79269189d4e0dc537556f4740f7f0a9ec41c1322598799b0bdad4ef92 \
--hash=sha256:61f642bf2378713e2c2e1de73444a3778e5f0a38be6fee0fe532fe30060282ff \
--hash=sha256:8ac7448f09ab85811607bdd21ec2464495ac8b7c66d146bf545b0f08fb9220ba \
--hash=sha256:a36d4a9dda1f19ce6e03c9a784a2921a4b726b02e1c736600ca9c22029474394 \
--hash=sha256:a422368fc821589c228f4c49438a368831cb5bbc0eab5ebe1d7fac9dded6567b \
--hash=sha256:e46dae94e34b085175f8abb3b0aaa7da40767865ac82c928eeb9e57e1ea8a543
# via pygithub
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
# via
# -r requirements_formatting.txt.in
# pygithub
toml==0.10.2 \
--hash=sha256:806143ae5bfb6a3c6e736a764057db0e6a0e05e338b5630894a5f779cabb4f9b \
--hash=sha256:b3bda1d108d5dd99f4a20d24d9c348e91c4db7ab1b749200bded2f839ccbe68f
# via
# darker
# darkgraylib
typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.5.0 \
--hash=sha256:3fc47733c7e419d4bc3f6b3dc2b4f890bb743906a30d56ba4a5bfa4bbff92760 \
--hash=sha256:e6b01673c0fa6a13e374b50871808eb3bf7046c4b125b216f6bf1cc604cff0dc
# via
# -r requirements_formatting.txt.in
# pygithub
# requests
wrapt==1.15.0 \
--hash=sha256:02fce1852f755f44f95af51f69d22e45080102e9d00258053b79367d07af39c0 \
--hash=sha256:077ff0d1f9d9e4ce6476c1a924a3332452c1406e59d90a2cf24aeb29eeac9420 \
--hash=sha256:078e2a1a86544e644a68422f881c48b84fef6d18f8c7a957ffd3f2e0a74a0d4a \
--hash=sha256:0970ddb69bba00670e58955f8019bec4a42d1785db3faa043c33d81de2bf843c \
--hash=sha256:1286eb30261894e4c70d124d44b7fd07825340869945c79d05bda53a40caa079 \
--hash=sha256:21f6d9a0d5b3a207cdf7acf8e58d7d13d463e639f0c7e01d82cdb671e6cb7923 \
--hash=sha256:230ae493696a371f1dbffaad3dafbb742a4d27a0afd2b1aecebe52b740167e7f \
--hash=sha256:26458da5653aa5b3d8dc8b24192f574a58984c749401f98fff994d41d3f08da1 \
--hash=sha256:2cf56d0e237280baed46f0b5316661da892565ff58309d4d2ed7dba763d984b8 \
--hash=sha256:2e51de54d4fb8fb50d6ee8327f9828306a959ae394d3e01a1ba8b2f937747d86 \
--hash=sha256:2fbfbca668dd15b744418265a9607baa970c347eefd0db6a518aaf0cfbd153c0 \
--hash=sha256:38adf7198f8f154502883242f9fe7333ab05a5b02de7d83aa2d88ea621f13364 \
--hash=sha256:3a8564f283394634a7a7054b7983e47dbf39c07712d7b177b37e03f2467a024e \
--hash=sha256:3abbe948c3cbde2689370a262a8d04e32ec2dd4f27103669a45c6929bcdbfe7c \
--hash=sha256:3bbe623731d03b186b3d6b0d6f51865bf598587c38d6f7b0be2e27414f7f214e \
--hash=sha256:40737a081d7497efea35ab9304b829b857f21558acfc7b3272f908d33b0d9d4c \
--hash=sha256:41d07d029dd4157ae27beab04d22b8e261eddfc6ecd64ff7000b10dc8b3a5727 \
--hash=sha256:46ed616d5fb42f98630ed70c3529541408166c22cdfd4540b88d5f21006b0eff \
--hash=sha256:493d389a2b63c88ad56cdc35d0fa5752daac56ca755805b1b0c530f785767d5e \
--hash=sha256:4ff0d20f2e670800d3ed2b220d40984162089a6e2c9646fdb09b85e6f9a8fc29 \
--hash=sha256:54accd4b8bc202966bafafd16e69da9d5640ff92389d33d28555c5fd4f25ccb7 \
--hash=sha256:56374914b132c702aa9aa9959c550004b8847148f95e1b824772d453ac204a72 \
--hash=sha256:578383d740457fa790fdf85e6d346fda1416a40549fe8db08e5e9bd281c6a475 \
--hash=sha256:58d7a75d731e8c63614222bcb21dd992b4ab01a399f1f09dd82af17bbfc2368a \
--hash=sha256:5c5aa28df055697d7c37d2099a7bc09f559d5053c3349b1ad0c39000e611d317 \
--hash=sha256:5fc8e02f5984a55d2c653f5fea93531e9836abbd84342c1d1e17abc4a15084c2 \
--hash=sha256:63424c681923b9f3bfbc5e3205aafe790904053d42ddcc08542181a30a7a51bd \
--hash=sha256:64b1df0f83706b4ef4cfb4fb0e4c2669100fd7ecacfb59e091fad300d4e04640 \
--hash=sha256:74934ebd71950e3db69960a7da29204f89624dde411afbfb3b4858c1409b1e98 \
--hash=sha256:75669d77bb2c071333417617a235324a1618dba66f82a750362eccbe5b61d248 \
--hash=sha256:75760a47c06b5974aa5e01949bf7e66d2af4d08cb8c1d6516af5e39595397f5e \
--hash=sha256:76407ab327158c510f44ded207e2f76b657303e17cb7a572ffe2f5a8a48aa04d \
--hash=sha256:76e9c727a874b4856d11a32fb0b389afc61ce8aaf281ada613713ddeadd1cfec \
--hash=sha256:77d4c1b881076c3ba173484dfa53d3582c1c8ff1f914c6461ab70c8428b796c1 \
--hash=sha256:780c82a41dc493b62fc5884fb1d3a3b81106642c5c5c78d6a0d4cbe96d62ba7e \
--hash=sha256:7dc0713bf81287a00516ef43137273b23ee414fe41a3c14be10dd95ed98a2df9 \
--hash=sha256:7eebcdbe3677e58dd4c0e03b4f2cfa346ed4049687d839adad68cc38bb559c92 \
--hash=sha256:896689fddba4f23ef7c718279e42f8834041a21342d95e56922e1c10c0cc7afb \
--hash=sha256:96177eb5645b1c6985f5c11d03fc2dbda9ad24ec0f3a46dcce91445747e15094 \
--hash=sha256:96e25c8603a155559231c19c0349245eeb4ac0096fe3c1d0be5c47e075bd4f46 \
--hash=sha256:9d37ac69edc5614b90516807de32d08cb8e7b12260a285ee330955604ed9dd29 \
--hash=sha256:9ed6aa0726b9b60911f4aed8ec5b8dd7bf3491476015819f56473ffaef8959bd \
--hash=sha256:a487f72a25904e2b4bbc0817ce7a8de94363bd7e79890510174da9d901c38705 \
--hash=sha256:a4cbb9ff5795cd66f0066bdf5947f170f5d63a9274f99bdbca02fd973adcf2a8 \
--hash=sha256:a74d56552ddbde46c246b5b89199cb3fd182f9c346c784e1a93e4dc3f5ec9975 \
--hash=sha256:a89ce3fd220ff144bd9d54da333ec0de0399b52c9ac3d2ce34b569cf1a5748fb \
--hash=sha256:abd52a09d03adf9c763d706df707c343293d5d106aea53483e0ec8d9e310ad5e \
--hash=sha256:abd8f36c99512755b8456047b7be10372fca271bf1467a1caa88db991e7c421b \
--hash=sha256:af5bd9ccb188f6a5fdda9f1f09d9f4c86cc8a539bd48a0bfdc97723970348418 \
--hash=sha256:b02f21c1e2074943312d03d243ac4388319f2456576b2c6023041c4d57cd7019 \
--hash=sha256:b06fa97478a5f478fb05e1980980a7cdf2712015493b44d0c87606c1513ed5b1 \
--hash=sha256:b0724f05c396b0a4c36a3226c31648385deb6a65d8992644c12a4963c70326ba \
--hash=sha256:b130fe77361d6771ecf5a219d8e0817d61b236b7d8b37cc045172e574ed219e6 \
--hash=sha256:b56d5519e470d3f2fe4aa7585f0632b060d532d0696c5bdfb5e8319e1d0f69a2 \
--hash=sha256:b67b819628e3b748fd3c2192c15fb951f549d0f47c0449af0764d7647302fda3 \
--hash=sha256:ba1711cda2d30634a7e452fc79eabcadaffedf241ff206db2ee93dd2c89a60e7 \
--hash=sha256:bbeccb1aa40ab88cd29e6c7d8585582c99548f55f9b2581dfc5ba68c59a85752 \
--hash=sha256:bd84395aab8e4d36263cd1b9308cd504f6cf713b7d6d3ce25ea55670baec5416 \
--hash=sha256:c99f4309f5145b93eca6e35ac1a988f0dc0a7ccf9ccdcd78d3c0adf57224e62f \
--hash=sha256:ca1cccf838cd28d5a0883b342474c630ac48cac5df0ee6eacc9c7290f76b11c1 \
--hash=sha256:cd525e0e52a5ff16653a3fc9e3dd827981917d34996600bbc34c05d048ca35cc \
--hash=sha256:cdb4f085756c96a3af04e6eca7f08b1345e94b53af8921b25c72f096e704e145 \
--hash=sha256:ce42618f67741d4697684e501ef02f29e758a123aa2d669e2d964ff734ee00ee \
--hash=sha256:d06730c6aed78cee4126234cf2d071e01b44b915e725a6cb439a879ec9754a3a \
--hash=sha256:d5fe3e099cf07d0fb5a1e23d399e5d4d1ca3e6dfcbe5c8570ccff3e9208274f7 \
--hash=sha256:d6bcbfc99f55655c3d93feb7ef3800bd5bbe963a755687cbf1f490a71fb7794b \
--hash=sha256:d787272ed958a05b2c86311d3a4135d3c2aeea4fc655705f074130aa57d71653 \
--hash=sha256:e169e957c33576f47e21864cf3fc9ff47c223a4ebca8960079b8bd36cb014fd0 \
--hash=sha256:e20076a211cd6f9b44a6be58f7eeafa7ab5720eb796975d0c03f05b47d89eb90 \
--hash=sha256:e826aadda3cae59295b95343db8f3d965fb31059da7de01ee8d1c40a60398b29 \
--hash=sha256:eef4d64c650f33347c1f9266fa5ae001440b232ad9b98f1f43dfe7a79435c0a6 \
--hash=sha256:f2e69b3ed24544b0d3dbe2c5c0ba5153ce50dcebb576fdc4696d52aa22db6034 \
--hash=sha256:f87ec75864c37c4c6cb908d282e1969e79763e0d9becdfe9fe5473b7bb1e5f09 \
--hash=sha256:fbec11614dba0424ca72f4e8ba3c420dba07b4a7c206c8c8e4e73f2e98f4c559 \
--hash=sha256:fd69666217b62fa5d7c6aa88e507493a34dec4fa20c5bd925e4bc12fce586639
# via deprecated
@@ -0,0 +1,8 @@
black~=25.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=43.0.1
urllib3>=2.5.0
requests>=2.32.4
idna>=3.7
certifi>=2024.7.4
+1 -1
+1 -1
+66 -30
View File
@@ -575,7 +575,7 @@ def print_ir_arg_printer():
if arg.IsSSA:
# SSA value
output_file.write("\tPrintArg(out, IR, Op->Header.Args[{}], RAData);\n".format(SSAArgNum))
output_file.write("\tPrintArg(out, IR, Op->Header.Args[{}]);\n".format(SSAArgNum))
SSAArgNum = SSAArgNum + 1
else:
# User defined op that is stored
@@ -587,6 +587,15 @@ def print_ir_arg_printer():
output_file.write("#undef IROP_ARGPRINTER_HELPER\n")
output_file.write("#endif\n")
def print_validation(op):
if op.EmitValidation != None:
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
Sanitized = Validation.replace("\"", "\\\"")
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
# Print out IR allocator helpers
def print_ir_allocator_helpers():
output_file.write("#ifdef IROP_ALLOCATE_HELPERS\n")
@@ -678,7 +687,7 @@ def print_ir_allocator_helpers():
output_file.write("{} {}".format(CType, arg.Name));
elif arg.IsSSA:
# SSA value
output_file.write("OrderedNode *{}".format(arg.Name))
output_file.write("OrderedNodeWrapper {}".format(arg.Name))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
@@ -708,35 +717,16 @@ def print_ir_allocator_helpers():
output_file.write("\t\tauto _Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
if op.SSAArgNum != 0:
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t_Op.first->{} = {}->Wrapped(ListDataBegin);\n".format(arg.Name, arg.Name))
if op.SSAArgNum != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
output_file.write("\t\t_Op.first->{} = {};\n".format(arg.Name, arg.Name))
if len(op.Arguments) != 0:
for arg in op.Arguments:
if not arg.Temporary and not arg.IsSSA:
output_file.write("\t\t_Op.first->{} = {};\n".format(arg.Name, arg.Name))
if (op.HasDest):
# We can only infer a size if we have arguments
if op.DestSize == None:
# We need to infer destination size
output_file.write("\t\tIR::OpSize InferSize = OpSize::iUnsized;\n")
if len(op.Arguments) != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tauto Size{} = GetOpSize({});\n".format(arg.Name, arg.Name))
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tInferSize = std::max(InferSize, Size{});\n".format(arg.Name))
output_file.write("\t\t_Op.first->Header.Size = InferSize;\n")
assert not (op.HasDest and op.DestSize is None)
# Some ops without a destination still need an operating size
# Effectively reusing the destination size value for operation size
@@ -748,18 +738,64 @@ def print_ir_allocator_helpers():
else:
output_file.write("\t\t_Op.first->Header.ElementSize = {};\n".format(op.ElementSize))
# Insert validation here
if op.EmitValidation != None:
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
Sanitized = Validation.replace("\"", "\\\"")
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
# Only validate here if there's no OrderedNode * version. Else
# validation is in that version, see the comment below.
if op.SSAArgNum == 0:
print_validation(op)
output_file.write("\t\treturn _Op;\n")
output_file.write("\t}\n\n")
# Now do the OrderedNode * version if necessary
if op.SSAArgNum:
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
elif arg.IsSSA:
output_file.write("OrderedNode *{}".format(arg.Name))
else:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
if arg.DefaultInitializer != None:
output_file.write(" = {}".format(arg.DefaultInitializer))
if not LastArg:
output_file.write(", ")
output_file.write(") {\n")
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
# Insert validation here. This is skipped for the
# OrderedNodeWrapper version because validation can depend on
# the OrderedNode, but that's ok in practice. Everything pre-RA
# uses the OrderedNode version, and anything RA-onwards is
# dubious to validate.
print_validation(op)
output_file.write(f"\t\treturn _{op.Name}(")
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
output_file.write(arg.Name)
if arg.IsSSA:
output_file.write("->Wrapped(ListDataBegin)")
if not LastArg:
output_file.write(", ")
output_file.write(");\n");
output_file.write("\t}\n\n");
output_file.write("#undef IROP_ALLOCATE_HELPERS\n")
output_file.write("#endif\n")
-2
View File
@@ -1,4 +1,3 @@
include(GNUInstallDirs)
set (MAN_DIR share/man CACHE PATH "MAN_DIR")
set (FEXCORE_BASE_SRCS
@@ -69,7 +68,6 @@ set (SRCS
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/x87StackOptimizationPass.cpp
+53 -1
View File
@@ -157,7 +157,30 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
/*
* FPREM is not an IEEE-754 remainder. From the spec:
* Check for invalid operation cases first - Intel FPREM sets Invalid Operation
* for several cases including infinity dividend and zero divisor.
*/
X80SoftFloat result = 0;
if (HandleInfinityOp(state, lhs, result)) {
return result;
} else if (lhs.Exponent == 0x7FFF && (lhs.Significand & 0x7FFFFFFFFFFFFFFFULL)) { // NaN
// propagate NaN
state->exceptionFlags |= softfloat_flag_invalid;
return lhs;
}
// Check for zero divisor - fprem(x, 0) is invalid operation
if (rhs.Exponent == 0 && rhs.Significand == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return result;
}
/*
* FPREM is not an IEEE-754 remainder. From the Intel spec:
*
* Computes the remainder obtained from dividing the value in the ST(0)
* register (the dividend) by the value in the ST(1) register (the divisor
@@ -390,6 +413,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::tanl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -411,6 +439,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::sinl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -432,6 +465,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat result;
if (HandleInfinityOp(state, lhs, result)) {
return result;
}
BIGFLOAT Src_d = lhs.ToFMax(state);
Src_d = FEXCore::cephes_128bit::cosl(Src_d);
return X80SoftFloat(state, Src_d);
@@ -591,6 +629,20 @@ private:
static constexpr uint64_t IntegerBit = (1ULL << 63);
static constexpr uint64_t Bottom62Significand = ((1ULL << 62) - 1);
static constexpr uint32_t ExponentBias = 16383;
// Helper function to check for infinity and set invalid operation flag.
// Returns true if infinity is dealt with, false otherwise.
FEXCORE_PRESERVE_ALL_ATTR static bool HandleInfinityOp(softfloat_state* state, const X80SoftFloat& arg, X80SoftFloat& result) {
if (arg.Exponent == 0x7FFF && arg.Significand == 0x8000000000000000ULL) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN.
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return true;
}
return false;
}
};
#ifndef _WIN32
+18
View File
@@ -3,14 +3,32 @@
#ifdef _M_X86_64
#include <xmmintrin.h>
#include <immintrin.h>
#endif
namespace FEXCore {
struct VectorScalarF64Pair {
double val[2];
};
#ifdef _M_ARM_64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
struct VectorRegPairType {
VectorRegType val[2];
};
static inline VectorRegPairType MakeVectorRegPair(VectorRegType low, VectorRegType high) {
return VectorRegPairType {low, high};
}
#elif defined(_M_X86_64)
using VectorRegType = __m128i;
using VectorRegPairType = __m256i;
static inline VectorRegPairType MakeVectorRegPair(VectorRegType low, VectorRegType high) {
return _mm256_set_m128i(high, low);
}
#endif
} // namespace FEXCore
+7 -18
View File
@@ -68,7 +68,9 @@
"ENABLESVEBITPERM": "enablesvebitperm",
"DISABLESVEBITPERM": "disablesvebitperm",
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi"
"DISABLEPRESERVEALLABI": "disablepreserveallabi",
"ENABLEWFXT": "enablewfxt",
"DISABLEWFXT": "disablewfxt"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -89,7 +91,8 @@
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it"
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -118,7 +121,7 @@
},
"ThunkHostLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/@CMAKE_INSTALL_LIBDIR@/fex-emu/HostThunks/",
"Default": "@CMAKE_INSTALL_FULL_LIBDIR@/fex-emu/HostThunks",
"ShortArg": "t",
"Desc": [
"Folder to find the host-side thunking libraries."
@@ -126,26 +129,12 @@
},
"ThunkGuestLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks/",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks",
"ShortArg": "j",
"Desc": [
"Folder to find the guest-side thunking libraries."
]
},
"ThunkHostLibs32": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/@CMAKE_INSTALL_LIBDIR@/fex-emu/HostThunks_32/",
"Desc": [
"Folder to find the 32-bit host-side thunking libraries."
]
},
"ThunkGuestLibs32": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks_32/",
"Desc": [
"Folder to find the 32-bit guest-side thunking libraries."
]
},
"ThunkConfig": {
"Type": "str",
"Default": "",
+22 -21
View File
@@ -50,7 +50,6 @@ namespace HLE {
} // namespace FEXCore
namespace FEXCore::IR {
class RegisterAllocationData;
struct IRListCopy;
class IRListView;
namespace Validation {
@@ -60,8 +59,9 @@ namespace Validation {
namespace FEXCore::Context {
struct FEX_PACKED ExitFunctionLinkData {
uint64_t HostBranch;
uint64_t HostCode;
uint64_t GuestRIP;
int64_t CallerOffset;
};
struct CustomIRResult {
@@ -76,7 +76,7 @@ struct CustomIRResult {
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class ContextImpl final : public FEXCore::Context::Context {
class ContextImpl final : public FEXCore::Context::Context, CPU::CodeBufferManager {
public:
// Context base class implementation.
bool InitCore() override;
@@ -165,8 +165,10 @@ public:
IRCaptureCache.WriteFilesWithCode(Writer);
}
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -248,15 +250,7 @@ public:
ContextImpl(const FEXCore::HostFeatures& Features);
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
return Fn(Frame, Record);
}
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
@@ -264,6 +258,10 @@ public:
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
// NOTE: Other threads sharing the same CodeBuffer may reference
// invalidated data ranges through their L1/L2 caches. This is
// not currently a problem since FEX does not repurpose the
// invalidated CodeBuffer memory range currently.
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
@@ -271,34 +269,32 @@ public:
struct GenerateIRResult {
std::optional<IR::IRListView> IRView;
IR::RegisterAllocationData* RAData;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
bool NeedsAddGuestCodeRanges;
};
[[nodiscard]]
GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst);
struct CompileCodeResult {
void* CompiledCode;
CPU::CPUBackend::CompiledCode CompiledCode;
fextl::unique_ptr<FEXCore::Core::DebugData> DebugData;
uint64_t StartAddr;
uint64_t Length;
bool NeedsAddGuestCodeRanges;
};
[[nodiscard]]
CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
IR::OpSize GetGPROpSize() const {
return Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
FEXCore::JITSymbols Symbols;
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator;
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -373,7 +369,12 @@ private:
std::shared_mutex CustomIRMutex;
std::atomic<bool> HasCustomIRHandlers {};
fextl::unordered_map<uint64_t, std::tuple<CustomIREntrypointHandler, void*, void*>> CustomIRHandlers;
struct CustomIRHandlerEntry final {
CustomIREntrypointHandler Handler;
void *Creator;
void *Data;
};
fextl::unordered_map<uint64_t, CustomIRHandlerEntry> CustomIRHandlers;
IntervalList<uint64_t> ForceTSOValidRanges; // The ranges for which ForceTSOInstructions has populated data
fextl::set<uint64_t> ForceTSOInstructions;
};
+15 -7
View File
@@ -11,7 +11,7 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
Ref Tmp = A.Base;
if (A.Offset) {
Ref Offset = IREmit->_Constant(A.Offset);
Ref Offset = IREmit->Constant(A.Offset);
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, Offset) : Offset;
}
@@ -22,7 +22,7 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
if (Tmp) {
Tmp = IREmit->_AddShift(GPRSize, Tmp, A.Index, ShiftType::LSL, Log2);
} else {
Tmp = IREmit->_Lshl(GPRSize, A.Index, IREmit->_Constant(Log2));
Tmp = IREmit->_Lshl(GPRSize, A.Index, IREmit->Constant(Log2));
}
} else {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Index) : A.Index;
@@ -34,14 +34,22 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
//
// If the AddrSize is not the GPRSize then we need to clear the upper bits.
if ((A.AddrSize < GPRSize) && !AllowUpperGarbage && Tmp) {
Tmp = IREmit->_Bfe(GPRSize, IR::OpSizeAsBits(A.AddrSize), 0, Tmp);
uint32_t Bits = IR::OpSizeAsBits(A.AddrSize);
if (A.Base || A.Index) {
Tmp = IREmit->_Bfe(GPRSize, Bits, 0, Tmp);
} else if (A.Offset) {
uint64_t X = A.Offset;
X &= (1ull << Bits) - 1;
Tmp = IREmit->Constant(X);
}
}
if (A.Segment && AddSegmentBase) {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Segment) : A.Segment;
}
return Tmp ?: IREmit->_Constant(0);
return Tmp ?: IREmit->Constant(0);
}
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
@@ -99,7 +107,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->_Constant(A.Offset),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = 1,
};
@@ -142,7 +150,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->_Constant(A.Offset),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexScale = 1,
};
@@ -155,4 +163,4 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
}
}; // namespace FEXCore::IR
}; // namespace FEXCore::IR
@@ -73,13 +73,13 @@ namespace x64 {
ARMEmitter::Reg::r8, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
};
constexpr std::array<ARMEmitter::Register, 8> RA = {
constexpr std::array<ARMEmitter::Register, 7> RA = {
// All these callee saved
ARMEmitter::Reg::r20, ARMEmitter::Reg::r21, ARMEmitter::Reg::r22, ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24, ARMEmitter::Reg::r25, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
ARMEmitter::Reg::r24, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
};
constexpr unsigned RAPairs = 6;
constexpr unsigned RAPairs = 4;
// Dynamic GPRs
constexpr std::array<ARMEmitter::Register, 2> PreserveAll_Dynamic = {
@@ -143,18 +143,18 @@ namespace x64 {
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r8,
};
constexpr std::array<ARMEmitter::Register, 7> RA = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15,
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
constexpr std::array<ARMEmitter::Register, 6> RA = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15, ARMEmitter::Reg::r16, ARMEmitter::Reg::r30,
};
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_Dynamic = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
};
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_Dynamic = {ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r16,
ARMEmitter::Reg::r17, ARMEmitter::Reg::r30};
constexpr std::array<ARMEmitter::Register, 7> NotPreserved_Dynamic = RA;
constexpr std::array<ARMEmitter::Register, 7> NotPreserved_Dynamic = {ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14,
ARMEmitter::Reg::r15, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
ARMEmitter::Reg::r30};
constexpr unsigned RAPairs = 6;
constexpr unsigned RAPairs = 4;
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
ARMEmitter::VReg::v0, ARMEmitter::VReg::v1, ARMEmitter::VReg::v2, ARMEmitter::VReg::v3,
@@ -245,14 +245,12 @@ namespace x32 {
REG_AF,
};
constexpr std::array<ARMEmitter::Register, 15> RA = {
constexpr std::array<ARMEmitter::Register, 14> RA = {
// All these callee saved
ARMEmitter::Reg::r20,
ARMEmitter::Reg::r21,
ARMEmitter::Reg::r22,
ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24,
ARMEmitter::Reg::r25,
// Registers only available on 32-bit
// All these are caller saved (except for r19).
@@ -265,6 +263,7 @@ namespace x32 {
ARMEmitter::Reg::r29,
ARMEmitter::Reg::r30,
ARMEmitter::Reg::r24,
ARMEmitter::Reg::r19,
};
@@ -273,7 +272,7 @@ namespace x32 {
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
};
constexpr unsigned RAPairs = 12;
constexpr unsigned RAPairs = 10;
// All are caller saved
constexpr std::array<ARMEmitter::VRegister, 8> SRAFPR = {
@@ -370,6 +369,8 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr
// Hardcode a 256-bit vector width if we are running in the simulator.
// Allow the user to override this.
Simulator.SetVectorLengthInBits(ForceSVEWidth() ? ForceSVEWidth() : 256);
// FEX doesn't support GCS.
Simulator.DisableGCSCheck();
#endif
#ifdef VIXL_DISASSEMBLER
// Only setup the disassembler if enabled.
@@ -495,20 +496,23 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
uint64_t AlignedPC = PC & ~0xFFFULL;
// Offset from aligned PC
int64_t AlignedOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(AlignedPC);
auto AlignedOffset = std::bit_cast<int64_t>(Constant - AlignedPC);
int NumMoves = 0;
// If the aligned offset is within the 4GB window then we can use ADRP+ADD
// and the number of move segments more than 1
if (RequiredMoveSegments > 1 && ARMEmitter::Emitter::IsInt32(AlignedOffset)) {
// NOTE: JIT output is moved to a different buffer after compilation, so the
// current cursor address doesn't match the runtime instruction address.
// Hence this optimization is disabled until we enable code relocation patches.
if (RequiredMoveSegments > 1 && ARMEmitter::Emitter::IsInt32(AlignedOffset) && false) {
// If this is 4k page aligned then we only need ADRP
if ((AlignedOffset & 0xFFF) == 0) {
adrp(Reg, AlignedOffset >> 12);
} else {
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
auto SmallOffset = std::bit_cast<int64_t>(Constant - PC);
if (ARMEmitter::Emitter::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
} else {
@@ -590,8 +594,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
void Arm64Emitter::PopCalleeSavedRegisters() {
constexpr static std::array< std::tuple<ARMEmitter::DRegister, ARMEmitter::DRegister, ARMEmitter::DRegister, ARMEmitter::DRegister>, 2> FPRs = {{
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
{ARMEmitter::DReg::d8, ARMEmitter::DReg::d9, ARMEmitter::DReg::d10, ARMEmitter::DReg::d11},
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
}};
for (auto& RegQuad : FPRs) {
@@ -691,6 +695,8 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
@@ -786,6 +792,8 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(STATE, TmpReg, CPU_AREA_EMULATOR_DATA_OFFSET);
#endif
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
@@ -43,6 +43,8 @@ constexpr bool TMP_ABIARGS = true;
constexpr auto REG_PF = ARMEmitter::Reg::r26;
constexpr auto REG_AF = ARMEmitter::Reg::r27;
constexpr auto REG_CALLRET_SP = ARMEmitter::XReg::x25;
// Vector temporaries
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
@@ -61,6 +63,8 @@ constexpr bool TMP_ABIARGS = false;
constexpr auto REG_PF = ARMEmitter::Reg::r9;
constexpr auto REG_AF = ARMEmitter::Reg::r24;
constexpr auto REG_CALLRET_SP = ARMEmitter::XReg::x17;
// Vector temporaries
constexpr auto VTMP1 = ARMEmitter::VReg::v16;
constexpr auto VTMP2 = ARMEmitter::VReg::v17;
@@ -84,7 +88,8 @@ constexpr uint64_t EC_CODE_BITMAP_MAX_ADDRESS = 1ULL << 47;
#endif
// Will force one single instruction block to be generated first if set when entering the JIT filling SRA.
constexpr auto ENTRY_FILL_SRA_SINGLE_INST_REG = TMP1;
// FillStaticRegs must preserve this
constexpr auto ENTRY_FILL_SRA_SINGLE_INST_REG = TMP2;
// Predicate to use in the X87 SVE optimization
constexpr ARMEmitter::PRegister PRED_X87_SVEOPT = ARMEmitter::PReg::p2;
+91 -63
View File
@@ -6,6 +6,8 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <cstdint>
#include "LookupCache.h"
#ifndef _WIN32
#include <sys/prctl.h>
#endif
@@ -13,6 +15,10 @@
namespace FEXCore {
namespace CPU {
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX][2] = {
{0x0003'0002'0001'0000ULL, 0x0007'0006'0005'0004ULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX
{0x000B'000A'0009'0008ULL, 0x000F'000E'000D'000CULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX_UPPER
@@ -264,10 +270,9 @@ namespace CPU {
return TotalLUT;
}()};
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState* ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
CPUBackend::CPUBackend(CodeBufferManager& CodeBuffers, FEXCore::Core::InternalThreadState* ThreadState)
: ThreadState(ThreadState)
, InitialCodeSize(InitialCodeSize)
, MaxCodeSize(MaxCodeSize) {
, CodeBuffers(CodeBuffers) {
auto& Common = ThreadState->CurrentFrame->Pointers.Common;
@@ -304,52 +309,63 @@ namespace CPU {
#endif
}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
CPUBackend::~CPUBackend() = default;
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer* {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = CodeBuffers.data();
auto PrevCodeBuffer = CurrentCodeBuffer;
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer = CodeBuffers.StartLargerCodeBuffer();
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
}
return CurrentCodeBuffer;
RegisterForSignalHandler(PrevCodeBuffer);
return CurrentCodeBuffer.get();
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
void CPUBackend::RegisterForSignalHandler(fextl::shared_ptr<CodeBuffer> CodeBuffer) {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter != 0) {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Keep a reference to the old code buffer to delay deallocation
SignalHandlerCodeBuffers.push_back(CodeBuffer);
} else {
SignalHandlerCodeBuffers.clear();
}
}
fextl::shared_ptr<CodeBuffer> CPUBackend::CheckCodeBufferUpdate() {
fextl::shared_ptr<CodeBuffer> OldCodeBuffer;
auto NewCodeBuffer = CodeBuffers.GetLatest();
if (CurrentCodeBuffer != NewCodeBuffer) {
RegisterForSignalHandler(CurrentCodeBuffer);
return std::exchange(CurrentCodeBuffer, NewCodeBuffer);
}
return nullptr;
}
GuestToHostMap& GetLookupCache(const CodeBuffer& Buffer) {
return *Buffer.LookupCache;
}
CodeBuffer::CodeBuffer(size_t Size)
: Size(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
LookupCache = fextl::make_unique<GuestToHostMap>();
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, Size);
}
auto CodeBufferManager::AllocateNew(size_t Size) -> fextl::shared_ptr<CodeBuffer> {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
@@ -375,39 +391,51 @@ namespace CPU {
}
#endif
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
auto Buffer = fextl::make_shared<CodeBuffer>(Size);
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
Latest = Buffer;
LatestOffset = 0;
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
OnCodeBufferAllocated(*Buffer);
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
fextl::shared_ptr<CodeBuffer> CodeBufferManager::GetLatest() {
if (!Latest) {
AllocateNew(INITIAL_CODE_SIZE);
}
return Latest;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::StartLargerCodeBuffer() {
if (!Latest) {
// Allocate initial CodeBuffer and return it
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->Size;
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
// The last page of the code buffer is protected, so we need to exclude it from the valid range
// when checking if the address is in the code buffer.
for (auto& Buffer : CodeBuffers) {
auto CheckCodeBuffer = [](CodeBuffer& Buffer, uintptr_t Address) {
// The last page of the code buffer is protected, so we need to exclude it from the valid range
// when checking if the address is in the code buffer.
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (Address >= reinterpret_cast<uintptr_t>(Buffer.Ptr) && Address < LastPageAddr) {
return (Address >= reinterpret_cast<uintptr_t>(Buffer.Ptr) && Address < LastPageAddr);
};
if (CheckCodeBuffer(*CurrentCodeBuffer, Address)) {
return true;
}
for (auto& Buffer : SignalHandlerCodeBuffers) {
if (CheckCodeBuffer(*Buffer, Address)) {
return true;
}
}
return false;
}
+70 -33
View File
@@ -9,8 +9,11 @@ $end_info$
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/map.h>
#include <cstdint>
@@ -18,7 +21,6 @@ namespace FEXCore {
namespace IR {
class IRListView;
class RegisterAllocationData;
} // namespace IR
namespace Core {
@@ -32,34 +34,68 @@ namespace CodeSerialize {
struct CodeObjectFileSection;
}
struct GuestToHostMap;
namespace CPU {
struct CodeBuffer {
uint8_t* Ptr;
size_t Size;
fextl::unique_ptr<GuestToHostMap> LookupCache;
CodeBuffer(size_t Size);
CodeBuffer(const CodeBuffer&) = delete;
CodeBuffer& operator=(const CodeBuffer&) = delete;
CodeBuffer(CodeBuffer&& oth) = delete;
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
};
/**
* A manager that coordinates access to the CodeBuffer used for compiling new code across threads.
*
* The CodeBuffer is managed as a partially persistent data structure:
* - Exactly one CodeBuffer is now designated as "active", which means data can be appended to it
* - Lossy modifications to the active CodeBuffer will not invalidate any data in use by other threads (which is what enables save CodeBuffer sharing across threads)
* - Instead, such lossy modifications trigger a new "version" of the data in the modifying thread. Old versions of the CodeBuffer persist as read-only data for use by the other threads.
* - The other threads can update their version of the CodeBuffer. This will decrease the reference count and eventually trigger deallocation of the old version
*/
class CodeBufferManager {
public:
// Get the CodeBuffer that was most recently allocated.
// This is the only CodeBuffer that data may be written to.
fextl::shared_ptr<CodeBuffer> GetLatest();
// Allocate a new CodeBuffer with geometric growth up to an internal maximum.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartLargerCodeBuffer();
// Write offset into the latest CodeBuffer
std::size_t LatestOffset {};
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
fextl::shared_ptr<CodeBuffer> AllocateNew(size_t Size);
};
class CPUBackend {
public:
struct CodeBuffer {
uint8_t* Ptr;
size_t Size;
};
/**
* @param InitialCodeSize - Initial size for the code buffers
* @param MaxCodeSize - Max size for the code buffers
*/
CPUBackend(FEXCore::Core::InternalThreadState* ThreadState, size_t InitialCodeSize, size_t MaxCodeSize);
CPUBackend(CodeBufferManager&, FEXCore::Core::InternalThreadState*);
virtual ~CPUBackend();
struct CompiledCode {
// Where this code block begins.
uint8_t* BlockBegin;
/**
* The function entrypoint to this codeblock.
*
* This may or may not equal `BlockBegin` above. Depending on the CPU backend, it may stick data
* prior to the BlockEntry.
*
* Is actually a function pointer of type `void (FEXCore::Core::ThreadState *Thread)`
*/
uint8_t* BlockEntry;
fextl::map<uint64_t, uint8_t*> EntryPoints;
// The total size of the codeblock from [BlockBegin, BlockBegin+Size).
size_t Size;
};
@@ -119,7 +155,7 @@ namespace CPU {
*/
[[nodiscard]]
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, const FEXCore::IR::RegisterAllocationData* RAData, bool CheckTF) = 0;
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
@@ -143,6 +179,11 @@ namespace CPU {
bool IsAddressInCodeBuffer(uintptr_t Address) const;
// Updates the CodeBuffer if needed and returns a reference to the old one.
// The returned reference should be kept alive carefully to avoid early deletion of resources.
[[nodiscard]]
fextl::shared_ptr<CodeBuffer> CheckCodeBufferUpdate();
protected:
// Max spill slot size in bytes. We need at most 32 bytes
// to be able to handle a 256-bit vector store to a slot.
@@ -150,24 +191,20 @@ namespace CPU {
FEXCore::Core::InternalThreadState* ThreadState;
size_t InitialCodeSize, MaxCodeSize;
[[nodiscard]]
CodeBuffer* GetEmptyCodeBuffer();
// This is the current code buffer that we are tracking
CodeBuffer* CurrentCodeBuffer {};
// This is the code buffer containing the main code under execution by this thread.
// CheckCodeBufferUpdate must be used before compiling new code.
fextl::shared_ptr<CodeBuffer> CurrentCodeBuffer;
// Old CodeBuffer generations required to be valid until returning from signal handlers
fextl::vector<fextl::shared_ptr<CodeBuffer>> SignalHandlerCodeBuffers;
CodeBufferManager& CodeBuffers;
private:
CodeBuffer AllocateNewCodeBuffer(size_t Size);
void FreeCodeBuffer(CodeBuffer Buffer);
void EmplaceNewCodeBuffer(CodeBuffer Buffer) {
CurrentCodeBuffer = &CodeBuffers.emplace_back(Buffer);
}
// This is the array of code buffers. Unless signals force us to keep more than
// buffer, there will be only one entry here
fextl::vector<CodeBuffer> CodeBuffers {};
void RegisterForSignalHandler(fextl::shared_ptr<CodeBuffer>);
};
} // namespace CPU
+30 -29
View File
@@ -626,6 +626,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
// Only enable EnhancedREPMOVS if atomic memcpy tso emulation isn't enabled.
const uint32_t SupportsEnhancedREPMOVS = CTX->IsMemcpyAtomicTSOEnabled() == false;
const uint32_t SupportsVPCLMULQDQ = CTX->HostFeatures.SupportsPMULL_128Bit && SupportsAVX();
const uint32_t SupportsWFXT = CTX->HostFeatures.SupportsWFXT;
// Number of subfunctions
Res.eax = 0x0;
@@ -645,39 +646,39 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(1 << 13) | // Deprecates FPU CS and DS
(0 << 14) | // Intel MPX
(0 << 15) | // Intel Resource Directory Technology Allocation
(0 << 16) | // Reserved
(0 << 17) | // Reserved
(0 << 16) | // AVX512-F
(0 << 17) | // AVX512-DQ
(CTX->HostFeatures.SupportsRAND << 18) | // RDSEED
(1 << 19) | // ADCX and ADOX instructions
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(0 << 21) | // AVX512-IFMA
(0 << 22) | // PCOMMIT (deprecated?)
(1 << 23) | // CLFLUSHOPT instruction
(1 << 24) | // CLWB instruction
(0 << 25) | // Intel processor trace
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 28) | // Reserved
(0 << 26) | // AVX512-PF
(0 << 27) | // AVX512-ER
(0 << 28) | // AVX512-CD
(Features.SHA << 29) | // SHA instructions
(0 << 30) | // Reserved
(0 << 31); // Reserved
(0 << 30) | // AVX512-BW
(0 << 31); // AVX512-VL
Res.ecx = (1 << 0) | // PREFETCHWT1
(0 << 1) | // AVX512VBMI
(0 << 2) | // Usermode instruction prevention
(0 << 3) | // Protection keys for user mode pages
(0 << 4) | // OS protection keys
(0 << 5) | // waitpkg
(0 << 6) | // AVX512_VBMI2
(SupportsWFXT << 5) | // waitpkg
(0 << 6) | // AVX512-VBMI2
(0 << 7) | // CET shadow stack
(0 << 8) | // GFNI
(CTX->HostFeatures.SupportsAES256 << 9) | // VAES
(SupportsVPCLMULQDQ << 10) | // VPCLMULQDQ
(0 << 11) | // AVX512_VNNI
(0 << 12) | // AVX512_BITALG
(0 << 11) | // AVX512-VNNI
(0 << 12) | // AVX512-BITALG
(0 << 13) | // Intel Total Memory Encryption
(0 << 14) | // AVX512_VPOPCNTDQ
(0 << 15) | // Reserved
(0 << 14) | // AVX512-VPOPCNTDQ
(0 << 15) | // FZM (TDX)
(0 << 16) | // 5 Level page tables
(0 << 17) | // MPX MAWAU
(0 << 18) | // MPX MAWAU
@@ -685,28 +686,28 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(0 << 23) | // Reserved
(0 << 24) | // Reserved
(0 << 23) | // AES Key Locker
(1 << 24) | // bus-lock-detect
(0 << 25) | // CLDEMOTE
(0 << 26) | // Reserved
(0 << 26) | // MPRR (TDX)
(0 << 27) | // MOVDIRI
(0 << 28) | // MOVDIR64B
(0 << 29) | // Reserved
(0 << 29) | // ENQCMD
(0 << 30) | // SGX Launch configuration
(0 << 31); // Reserved
(0 << 31); // PKS
Res.edx = (0 << 0) | // Reserved
(0 << 1) | // Reserved
(0 << 2) | // AVX512_4VNNIW
(0 << 3) | // AVX512_4FMAPS
Res.edx = (0 << 0) | // SGX-TEM (TDX)
(0 << 1) | // SGX-KEYS
(0 << 2) | // AVX512-4VNNIW
(0 << 3) | // AVX512-4FMAPS
(1 << 4) | // Fast Short Rep Mov
(0 << 5) | // Reserved
(0 << 5) | // UINTR
(0 << 6) | // Reserved
(0 << 7) | // Reserved
(0 << 8) | // AVX512_VP2INTERSECT
(0 << 8) | // AVX512-VP2INTERSECT
(0 << 9) | // SRBDS_CTRL (Special Register Buffer Data Sampling Mitigations)
(0 << 10) | // VERW clears CPU buffers
(0 << 11) | // Reserved
(0 << 11) | // rtm-always-abort
(0 << 12) | // Reserved
(0 << 13) | // TSX Force Abort (TSX will force abort if attempted)
(0 << 14) | // SERIALIZE instruction
@@ -718,7 +719,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 20) | // Intel CET
(0 << 21) | // Reserved
(0 << 22) | // AMX-BF16 - Tile computation on bfloat16
(0 << 23) | // AVX512_FP16 - FP16 AVX512 instructions
(0 << 23) | // AVX512-FP16 - FP16 AVX512 instructions
(0 << 24) | // AMX-tile - If AMX is implemented
(0 << 25) | // AMX-int8 - AMX on 8-bit integers
(0 << 26) | // IBRS_IBPB - Speculation control
@@ -754,7 +755,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
// XFeatureSupportedMask[63:32]
Res.edx = 0; // Upper 32-bits of XFeatureSupportedMask
} else if (Leaf == 1) {
Res.eax = (0 << 0) | // XSAVEOPT
Res.eax = (1 << 0) | // XSAVEOPT
(0 << 1) | // XSAVEC (and XRSTOR)
(0 << 2) | // XGETBV - XGETBV with ECX=1 supported
(0 << 3); // XSAVES - XSAVES, XRSTORS, and IA32_XSS supported
+1 -1
View File
@@ -115,7 +115,7 @@ public:
private:
const FEXCore::Context::ContextImpl* CTX;
bool SupportsCPUIndexInTPIDRRO {};
[[maybe_unused]] bool SupportsCPUIndexInTPIDRRO {};
bool Hybrid {};
uint32_t Cores {};
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
+150 -75
View File
@@ -46,13 +46,13 @@ $end_info$
#include "FEXCore/Utils/SignalScopeGuards.h"
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SHMStats.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TodoDefines.h>
#include <algorithm>
#include <array>
@@ -405,14 +405,14 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
}
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
// TODO: This doesn't make sense when the parent thread doesn't outlive its children
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->CurrentFrame->Pointers.Common.L1Pointer = Thread->LookupCache->GetL1Pointer();
@@ -426,7 +426,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
// Create CPU backend
Thread->PassManager->InsertRegisterAllocationPass();
Thread->PassManager->InsertRegisterAllocationPass(this);
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
Thread->PassManager->Finalize();
@@ -441,6 +441,22 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXC
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = InitialRIP;
// Set up default code segment.
// Default code segment indexes match the numbers that the Linux kernel uses.
Thread->CurrentFrame->State.cs_idx = 6 << 3;
auto &GDT = Thread->CurrentFrame->State.gdt[Thread->CurrentFrame->State.cs_idx >> 3];
Thread->CurrentFrame->State.SetGDTBase(&GDT, 0);
Thread->CurrentFrame->State.SetGDTLimit(&GDT, 0xF'FFFFU);
if (Config.Is64BitMode) {
GDT.L = 1; // L = Long Mode = 64-bit
GDT.D = 0; // D = Default Operand SIze = Reserved
}
else {
GDT.L = 0; // L = Long Mode = 32-bit
GDT.D = 1; // D = Default Operand Size = 32-bit
}
// Copy over the new thread state to the new object
if (NewThreadState) {
memcpy(&Thread->CurrentFrame->State, NewThreadState, sizeof(FEXCore::Core::CPUState));
@@ -495,7 +511,13 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread) {
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
}
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
if (CodeObjectCacheService) {
@@ -503,18 +525,23 @@ void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread) {
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LookupCache->ClearCache();
Thread->CPUBackend->ClearCache();
if (NewCodeBuffer) {
// Allocate new CodeBuffer + L3 LookupCache and clear L1+L2 caches
Thread->CPUBackend->ClearCache();
} else {
// Clear L1+L2 cache of this thread, and clear L3 cache across any threads using it
Thread->LookupCache->ClearCache();
}
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter* IREmitter, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter* IREmitter, uint64_t GuestRIP) {
FEXCore::File::File FD = FEXCore::File::File::GetStdERR();
fextl::stringstream out;
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
FEXCore::IR::Dump(&out, &NewIR);
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
ContextImpl::GenerateIRResult
@@ -535,7 +562,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
TotalInstructionsLength = 1;
std::get<0>(Handler->second)(GuestRIP, Thread->OpDispatcher.get());
Handler->second.Handler(GuestRIP, Thread->OpDispatcher.get());
HasCustomIR = true;
}
}
@@ -547,19 +574,14 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
bool HadDispatchError {false};
bool HadInvalidInst {false};
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, MaxInst,
[Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Thread, Start, Length);
}
});
Thread->FrontendDecoder->DecodeInstructionsAtEntry(Thread, GuestCode, GuestRIP, MaxInst);
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
auto CodeBlocks = &BlockInfo->Blocks;
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount);
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode);
const auto GPRSize = GetGPROpSize();
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
@@ -673,7 +695,11 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
Thread->OpDispatcher->InvalidOp(DecodedInfo);
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::NOEXEC_INST) {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
} else {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
}
}
HadInvalidInst = true;
@@ -686,7 +712,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {{}, nullptr, 0, 0, 0, 0};
return {{}, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
@@ -712,26 +738,24 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
auto ShouldDump = Thread->OpDispatcher->ShouldDumpIR();
// Debug
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
IRDumper(Thread, IREmitter, GuestRIP);
}
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(IREmitter);
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr;
// Debug
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, RAData);
IRDumper(Thread, IREmitter, GuestRIP);
}
return {
.IRView = IREmitter->ViewIR(),
.RAData = RAData,
.TotalInstructions = TotalInstructions,
.TotalInstructionsLength = TotalInstructionsLength,
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
.Length = Thread->FrontendDecoder->DecodedMaxAddress - Thread->FrontendDecoder->DecodedMinAddress,
.NeedsAddGuestCodeRanges = !HasCustomIR,
};
}
@@ -743,10 +767,11 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto CompiledCode = Thread->CPUBackend->RelocateJITObjectCode(GuestRIP, CodeCacheEntry);
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.CompiledCode = {},
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.StartAddr = 0, // Unused
.Length = 0, // Unused
.NeedsAddGuestCodeRanges = false,
};
}
}
@@ -760,31 +785,44 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
// Generate IR + Meta Info
auto [IRView, RAData, TotalInstructions, TotalInstructionsLength, StartAddr, Length] =
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
return {nullptr, nullptr, 0, 0};
return {{}, nullptr, 0, 0, false};
}
// Attempt to get the CPU backend to compile this code
// Re-check if another thread raced us in compiling this block.
// We could lock CodeBufferWriteMutex earlier to prevent this from happening,
// but this would increase lock contention. Redundant frontend runs aren't
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(GuestRIP)) {
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
.StartAddr = 0,
.Length = 0,
.NeedsAddGuestCodeRanges = false};
}
}
auto DebugData = fextl::make_unique<FEXCore::Core::DebugData>();
// If the trap flag is set we generate single instruction blocks that each check to generate a single step exception.
bool TFSet = Thread->CurrentFrame->State.flags[X86State::RFLAG_TF_RAW_LOC];
// Attempt to get the CPU backend to compile this code
auto CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, Length, TotalInstructions == 1, &*IRView, DebugData.get(), RAData, TFSet);
auto CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, Length, TotalInstructions == 1, &*IRView, DebugData.get(), TFSet);
// Release the IR
Thread->OpDispatcher->DelayedDisownBuffer();
return {
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = CompiledCode.BlockEntry,
.CompiledCode = std::move(CompiledCode),
.DebugData = std::move(DebugData),
.StartAddr = StartAddr,
.Length = Length,
.NeedsAddGuestCodeRanges = NeedsAddGuestCodeRanges,
};
}
@@ -804,14 +842,18 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return HostCode;
}
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
auto [CompiledCode, DebugData, StartAddr, Length, NeedsAddGuestCodeRanges] = CompileCode(Thread, GuestRIP, MaxInst);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
return 0;
} else if (!DebugData) {
// DebugData wasn't populated, indicating another thread raced us for compiling this block
return reinterpret_cast<uintptr_t>(CodePtr);
}
// The core managed to compile the code.
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = reinterpret_cast<uint8_t*>(CodePtr);
auto FragmentBasePtr = CompiledCode.BlockBegin;
if (DebugData) {
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
@@ -820,7 +862,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
for (auto& Subblock : DebugData->Subblocks) {
auto BlockBasePtr = FragmentBasePtr + Subblock.HostCodeOffset;
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename,
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, GuestRIP, Subblock.HostCodeSize);
@@ -828,10 +870,10 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
} else {
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename,
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, DebugData->HostCodeSize);
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, CompiledCode.Size);
}
}
}
@@ -844,8 +886,8 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
.GuestRIP = GuestRIP,
.GuestCodeLength = Length,
.GuestCodeHash = 0,
.HostCodeBegin = CodePtr,
.HostCodeLength = DebugData->HostCodeSize,
.HostCodeBegin = CompiledCode.BlockBegin,
.HostCodeLength = CompiledCode.Size,
.HostCodeHash = 0,
.ThreadJobRefCount = &Thread->ObjectCacheRefCounter,
.Relocations = std::move(*DebugData->Relocations),
@@ -855,14 +897,26 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, {}, DebugData.get(), false)) {
if (IRCaptureCache.PostCompileCode(Thread, CompiledCode.BlockBegin, GuestRIP, StartAddr, Length, {}, DebugData.get(), false)) {
// Early exit
return (uintptr_t)CodePtr;
}
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
// Insert to lookup cache
// Pages containing this block are added via AddBlockExecutableRange before each page gets accessed in the frontend
Thread->LookupCache->AddBlockMapping(GuestRIP, CodePtr);
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(GuestAddr, HostAddr);
}
return (uintptr_t)CodePtr;
}
@@ -876,7 +930,8 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
auto [CodePtr, DebugData, StartAddr, Length] = CompileCode(Thread, GuestRIP, 1);
auto [CompiledCode, DebugData, StartAddr, Length, _] = CompileCode(Thread, GuestRIP, 1);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
return 0;
}
@@ -887,22 +942,40 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lower = Thread->LookupCache->CodePages.lower_bound(Start >> 12);
auto upper = Thread->LookupCache->CodePages.upper_bound((Start + Length - 1) >> 12);
auto lk = Thread->LookupCache->AcquireLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (auto Address : it->second) {
ContextImpl::ThreadRemoveCodeEntry(Thread, Address);
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry)) {
InvalidatedAnyEntries = true;
}
}
it->second.clear();
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
@@ -915,18 +988,18 @@ void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
UpdateAtomicTSOEmulationConfig();
if (Config.TSOAutoMigration) {
// Only the lookup cache is cleared here, so that old code can keep running until next compilation
std::lock_guard<std::recursive_mutex> lkLookupCache(Thread->LookupCache->WriteLock);
// Only the lookup cache is cleared here, so that old code can keep running until next compilation.
// This will leak previously compiled blocks until the CodeBuffer is cleared for some other reason.
Thread->LookupCache->ClearCache();
}
}
}
void ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
std::optional<CustomIRResult>
@@ -935,7 +1008,7 @@ ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandl
std::unique_lock lk(CustomIRMutex);
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, std::tuple(Handler, Creator, Data));
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, CustomIRHandlerEntry {Handler, Creator, Data});
HasCustomIRHandlers = true;
if (!InsertedIterator.second) {
@@ -959,20 +1032,21 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
auto Result = AddCustomIREntrypoint(
Entrypoint,
[this, GuestThunkEntrypoint](uintptr_t Entrypoint, FEXCore::IR::IREmitter* emit) {
auto IRHeader = emit->_IRHeader(emit->Invalid(), Entrypoint, 0, 0);
auto Block = emit->CreateCodeNode();
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
auto IRHeader = emit->_IRHeader(emit->Invalid(), Entrypoint, 0, 0, 0, 0);
auto Block = emit->CreateCodeNode(true, 0);
IRHeader.first->Blocks = emit->WrapNode(Block);
emit->SetCurrentCodeBlock(Block);
const auto GPRSize = GetGPROpSize();
const auto GPRSize = this->Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
if (GPRSize == IR::OpSize::i64Bit) {
emit->_StoreRegister(emit->_Constant(Entrypoint), X86State::REG_R11, IR::GPRClass, GPRSize);
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->_Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::GPRFixedClass, X86State::REG_R11).Raw;
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(IR::OpSize::i64Bit, emit->Constant(GuestThunkEntrypoint), IR::BranchHint::None, emit->Invalid(), emit->Invalid());
},
ThunkHandler, (void*)GuestThunkEntrypoint);
@@ -1006,7 +1080,8 @@ void ContextImpl::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(nullptr, Entrypoint, 1);
InvalidatedEntryAccumulator Accumulator;
InvalidateGuestCodeRange(nullptr, Accumulator, Entrypoint, 1);
CustomIRHandlers.erase(Entrypoint);
HasCustomIRHandlers = !CustomIRHandlers.empty();
@@ -81,11 +81,14 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::rsp, 0);
str(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
ARMEmitter::ForwardLabel CompileSingleStep;
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
FillStaticRegs();
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
ARMEmitter::BiDirectionalLabel LoopTop {};
ARMEmitter::ForwardLabel CompileSingleStep;
#ifdef _M_ARM_64EC
b(&LoopTop);
@@ -111,8 +114,21 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
ldr(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
FillSpecialRegs(TMP1, TMP2, false, true);
// As ARM64EC uses this as an entrypoint for both guest calls and host returns, opportunistically try to return
// using the call-ret stack to avoid unbalancing it.
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
ret(TMP2);
// Enter JIT
#endif
@@ -286,6 +302,8 @@ void Dispatcher::EmitDispatcher() {
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
@@ -498,6 +516,7 @@ void Dispatcher::EmitDispatcher() {
// load static regs
FillStaticRegs();
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
// Now go back to the regular dispatcher loop
b(&LoopTop);
@@ -517,14 +536,15 @@ void Dispatcher::EmitDispatcher() {
ldr(ARMEmitter::XReg::x3, R, Offset);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<__uint128_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
// Result is now in x0
// Result is now in x0, x1
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
mov(TMP2, ARMEmitter::XReg::x1);
}
FillStaticRegs();
@@ -539,8 +559,6 @@ void Dispatcher::EmitDispatcher() {
LUDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
LDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
LUREMHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
LREMHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
// Interpreter fallbacks
{
@@ -559,6 +577,8 @@ void Dispatcher::EmitDispatcher() {
FABI_I64_I16_F80_F80_PTR,
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
}};
@@ -606,6 +626,7 @@ void Dispatcher::EmitDispatcher() {
#ifdef VIXL_SIMULATOR
void Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.WriteXRegister(1, 0);
Simulator.RunFrom(reinterpret_cast< const vixl::aarch64::Instruction*>(DispatchPtr));
}
@@ -627,6 +648,22 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
constexpr static auto VABI1 = ARMEmitter::VReg::v0;
constexpr static auto VABI2 = ARMEmitter::VReg::v1;
auto FillF80x2Result = [&]() {
if (!TMP_ABIARGS) {
mov(VTMP1.Q(), VABI1.Q());
mov(VTMP2.Q(), VABI2.Q());
}
FillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, true);
};
auto FillF64x2Result = [&]() {
if (!TMP_ABIARGS) {
fmov(VTMP1.D(), VABI1.D());
fmov(VTMP2.D(), VABI2.D());
}
FillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, true);
};
auto FillF80Result = [&]() {
if (VTMP1 != VABI1) {
mov(VTMP1.Q(), VABI1.Q());
@@ -947,6 +984,52 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
FillF80Result();
} break;
case FABI_F80x2_I16_F80_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
// vtmp1 (v0/v16): vector source 1
// vtmp2 (v1/v16): vector source 2
SpillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, TMP3, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
if (!TMP_ABIARGS) {
mov(VABI1.Q(), VTMP1.Q());
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
// GenerateIndirectRuntimeCall<FEXCore::VectorRegPairType, uint16_t, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
FillF80x2Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
// vtmp1 (v0/v16): vector source 1
// vtmp2 (v1/v16): vector source 2
SpillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, TMP3, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
// GenerateIndirectRuntimeCall<FEXCore::VectorScalarF64Pair, uint16_t, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
FillF64x2Result();
} break;
case FABI_I32_I64_I64_V128_V128_I16: {
// Linux Reg/Win32 Reg:
// stack: FallbackHandler
@@ -1036,8 +1119,6 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
auto& AArch64 = Thread->CurrentFrame->Pointers.AArch64;
AArch64.LUDIVHandler = LUDIVHandlerAddress;
AArch64.LDIVHandler = LDIVHandlerAddress;
AArch64.LUREMHandler = LUREMHandlerAddress;
AArch64.LREMHandler = LREMHandlerAddress;
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers, &ABIPointers[0]);
@@ -76,7 +76,7 @@ public:
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame);
DispatchPtr(Frame, false);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
@@ -109,7 +109,7 @@ public:
protected:
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame);
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame, bool SingleInst);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
@@ -118,8 +118,6 @@ private:
// Long division helpers
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
uint64_t LUREMHandlerAddress {};
uint64_t LREMHandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
+150 -78
View File
@@ -9,6 +9,8 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
#include <algorithm>
@@ -21,6 +23,7 @@ $end_info$
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/fextl/set.h>
namespace FEXCore::Frontend {
@@ -64,33 +67,73 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
}
}
Decoder::Decoder(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, OSABI {ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {}
Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
: Thread {Thread}
, CTX {static_cast<FEXCore::Context::ContextImpl*>(Thread->CTX)}
, OSABI {CTX->SyscallHandler ? CTX->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {CTX->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {}
Decoder::~Decoder() {
PoolObject.UnclaimBuffer();
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
ExecutableRangeEnd = RangeInfo.Base + RangeInfo.Size;
if (RangeInfo.Size == 0) {
return false;
}
uint64_t RangeRemainingSize = ExecutableRangeEnd - Address;
if (Size > RangeRemainingSize) {
Size -= RangeRemainingSize;
Address += RangeRemainingSize;
}
}
return true;
}
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
std::optional<uint8_t> Byte = PeekByte(0);
if (!Byte) {
HitNonExecutableRange = true;
// Pretend we read 0, the main decode loop will see HitNonExecutableRange and rollback the instruction.
return 0;
}
Instruction[InstructionSize] = *Byte;
InstructionSize++;
return Byte;
return *Byte;
}
uint8_t Decoder::PeekByte(uint8_t Offset) const {
uint8_t Byte = InstStream[InstructionSize + Offset];
return Byte;
std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream + InstructionSize + Offset);
if (CheckRangeExecutable(ByteAddress, 1)) {
return InstStream[InstructionSize + Offset];
} else {
return std::nullopt;
}
}
uint64_t Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
uint64_t Address = reinterpret_cast<uint64_t>(InstStream + InstructionSize);
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream[InstructionSize], Size);
} else {
HitNonExecutableRange = true;
// See PeekByte, this specific case may cause some executable memory to read as 0 but it doesn't matter as the entire instruction will be rolled back anyway.
Res = 0;
}
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
for (size_t i = 0; i < Size; ++i) {
@@ -290,7 +333,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
uint8_t DestSize {};
const bool HasWideningDisplacement =
(FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST) != 0 ||
(Options.w && CTX->Config.Is64BitMode);
(Options.w && BlockInfo.Is64BitMode);
const bool HasNarrowingDisplacement =
(FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
@@ -308,7 +351,21 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const bool HasMODRM = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM);
const bool HasREX = !!(DecodeInst->Flags & DecodeFlags::FLAG_REX_PREFIX);
const bool Has16BitAddressing = !CTX->Config.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
if (Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_0)) {
return false;
} else if (!Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_1)) {
return false;
}
if (Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_0)) {
return false;
} else if (!Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_1)) {
return false;
}
const bool UseVEXL = Options.L && !(Info->Flags & InstFlags::FLAGS_VEX_L_IGNORE);
// This is used for ModRM register modification
// For both modrm.reg and modrm.rm(when mod == 0b11) when value is >= 0b100
@@ -339,7 +396,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
DestSize = 2;
} else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
if (Options.L) {
if (UseVEXL) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_256BIT);
DestSize = 32;
} else {
@@ -355,7 +412,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
DestSize = 2;
} else if ((HasXMMDst || HasMMDst || CTX->Config.Is64BitMode) &&
} else if ((HasXMMDst || HasMMDst || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_64BIT);
@@ -372,7 +429,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
} else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_16BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
} else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
if (Options.L) {
if (UseVEXL) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
@@ -384,7 +441,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// See table 1-2. Operand-Size Overrides for this decoding
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
} else if ((HasXMMSrc || HasMMSrc || CTX->Config.Is64BitMode) &&
} else if ((HasXMMSrc || HasMMSrc || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_64BIT);
@@ -484,6 +541,9 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
size_t CurrentSrc = 0;
const auto VEXOperand = Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_SRC_MASK;
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_NO_OPERAND && Options.vvvv) {
return false;
}
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
@@ -651,7 +711,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodedHeader options {};
if ((Byte1 & 0b10000000) == 0) {
if (!CTX->Config.Is64BitMode) {
if (!BlockInfo.Is64BitMode) {
return false;
}
@@ -670,12 +730,12 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
if (!CTX->Config.Is64BitMode) {
if (!BlockInfo.Is64BitMode) {
return false;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
if (BlockInfo.Is64BitMode && (Byte1 & 0b00100000) == 0) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
if (options.w) {
@@ -720,7 +780,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstruction(uint64_t PC) {
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
Instruction.fill(0);
@@ -748,7 +808,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
const bool Has16BitAddressing = !CTX->Config.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
@@ -857,29 +917,25 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->Flags |= DecodeFlags::FLAG_ADDRESS_SIZE;
break;
case 0x26: // ES legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_ES_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_ES_PREFIX;
}
break;
case 0x2E: // CS legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_CS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_CS_PREFIX;
}
break;
case 0x36: // SS legacy prefix
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_SS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_SS_PREFIX;
}
break;
case 0x3E: // DS legacy prefix
// Annoyingly GCC generates NOP ops with these prefixes
// Just ignore them for now
// eg. 66 2e 0f 1f 84 00 00 00 00 00 nop WORD PTR cs:[rax+rax*1+0x0]
if (!CTX->Config.Is64BitMode) {
DecodeInst->Flags |= DecodeFlags::FLAG_DS_PREFIX;
if (!BlockInfo.Is64BitMode) {
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_DS_PREFIX;
}
break;
break;
case 0xF0: // LOCK prefix
DecodeInst->Flags |= DecodeFlags::FLAG_LOCK;
break;
@@ -892,19 +948,16 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->LastEscapePrefix = Op;
break;
case 0x64: // FS prefix
DecodeInst->Flags |= DecodeFlags::FLAG_FS_PREFIX;
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_FS_PREFIX;
break;
case 0x65: // GS prefix
DecodeInst->Flags |= DecodeFlags::FLAG_GS_PREFIX;
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_GS_PREFIX;
break;
default:
[[likely]] { // Default base table
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
if (!CTX->Config.Is64BitMode) {
return false;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
@@ -935,6 +988,31 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
}
}
if (DecodeInst->Dest.IsGPR()) {
return false;
}
return true;
}
Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
return ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST : DecodedBlockStatus::NOEXEC_INST;
} else if (!DecodeInst->TableInfo || !DecodeInst->TableInfo->OpcodeDispatcher) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
}
return DecodedBlockStatus::SUCCESS;
}
void Decoder::BranchTargetInMultiblockRange() {
@@ -944,16 +1022,23 @@ void Decoder::BranchTargetInMultiblockRange() {
// If the RIP setting is conditional AND within our symbol range then it can be considered for multiblock
uint64_t TargetRIP = 0;
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
bool Conditional = true;
const auto InstEnd = DecodeInst->PC + DecodeInst->InstSize;
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_CALL) {
AddBranchTarget(InstEnd);
BlockInfo.EntryPoints.emplace(InstEnd);
return;
}
// Calls are handled above
switch (DecodeInst->OP) {
case 0x70 ... 0x7F: // Conditional JUMP
case 0x80 ... 0x8F: { // More conditional
// Source is a literal
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// auto RIPTargetConst = Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
break;
@@ -963,11 +1048,6 @@ void Decoder::BranchTargetInMultiblockRange() {
TargetRIP = InstEnd + DecodeInst->Src[0].Literal();
Conditional = false;
break;
case 0xE8: // Call - Immediate target, We don't want to inline calls
if (ExternalBranches) {
ExternalBranches->insert(InstEnd);
}
[[fallthrough]];
case 0xC2: // RET imm
case 0xC3: // RET
default: return; break;
@@ -1015,7 +1095,7 @@ bool Decoder::InstCanContinue() const {
}
uint64_t TargetRIP = 0;
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
if (DecodeInst->OP == 0xE8) { // Call - immediate target
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
@@ -1067,7 +1147,7 @@ void Decoder::AddBranchTarget(uint64_t Target) {
.Size = BlockIt->Size - SplitOffset,
.NumInstructions = BlockIt->NumInstructions - SplitIdx,
.DecodedInstructions = BlockIt->DecodedInstructions + SplitIdx,
.HasInvalidInstruction = BlockIt->HasInvalidInstruction,
.BlockStatus = BlockIt->BlockStatus,
};
BlockIt->Size = SplitOffset;
@@ -1107,12 +1187,10 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst,
std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage) {
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
BlocksToDecode.clear();
VisitedBlocks.clear();
// Reset internal state management
DecodedSize = 0;
@@ -1120,9 +1198,15 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// Decode operating mode from thread's CS segment.
const auto CSSegment = Thread->CurrentFrame->State.gdt[Thread->CurrentFrame->State.cs_idx >> 3];
BlockInfo.Is64BitMode = CSSegment.L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
// XXX: Load symbol data
SymbolAvailable = false;
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
@@ -1138,13 +1222,11 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
DecodedMaxAddress = EntryPoint;
// Entry is a jump target
BlocksToDecode.emplace(PC);
BlocksToDecode = {PC};
uint64_t CurrentCodePage = PC & FEXCore::Utils::FEX_PAGE_MASK;
fextl::set<uint64_t> CodePages = {CurrentCodePage};
AddContainedCodePage(PC, CurrentCodePage, FEXCore::Utils::FEX_PAGE_SIZE);
BlockInfo.CodePages = {CurrentCodePage};
if (MaxInst == 0) {
MaxInst = CTX->Config.MaxInstPerBlock;
@@ -1179,6 +1261,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockIt->Entry = RIPToDecode;
BlockIt->Size = 0;
BlockIt->IsEntryPoint = EntryBlock;
uint64_t PCOffset = 0;
uint64_t BlockStartOffset = DecodedSize;
@@ -1200,7 +1283,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
auto OpMinPage = OpAddress & FEXCore::Utils::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FEXCore::Utils::FEX_PAGE_MASK;
if (!EntryBlock && OpMinPage == OpMaxPage && PeekByte(0) == 0 && PeekByte(1) == 0) [[unlikely]] {
if (!EntryBlock && OpMinPage == OpMaxPage && PeekByte(0).value_or(0) == 0 && PeekByte(1).value_or(0) == 0) [[unlikely]] {
// End the multiblock early if we hit 2 consecutive null bytes (add [rax], al) in the same page with the
// assumption we are most likely trying to explore garbage code.
break;
@@ -1208,31 +1291,17 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
CodePages.insert(CurrentCodePage);
BlockInfo.CodePages.insert(CurrentCodePage);
}
if (OpMaxPage != CurrentCodePage) {
CurrentCodePage = OpMaxPage;
CodePages.insert(CurrentCodePage);
BlockInfo.CodePages.insert(CurrentCodePage);
}
bool ErrorDuringDecoding = !DecodeInstruction(OpAddress);
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
if (ErrorDuringDecoding) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
BlockIt->HasInvalidInstruction = true;
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
} else {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
auto TableInfo = DecodeInst->TableInfo;
if (!TableInfo || !TableInfo->OpcodeDispatcher) {
BlockIt->HasInvalidInstruction = true;
}
}
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
DecodedMaxAddress = std::max(DecodedMaxAddress, OpEndAddress);
@@ -1248,7 +1317,7 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockIt->Size += DecodeInst->InstSize;
// Can not continue this block at all on invalid instruction
if (BlockIt->HasInvalidInstruction) [[unlikely]] {
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
@@ -1257,6 +1326,9 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
DecodedSize = BlockStartOffset;
InstStream -= PCOffset;
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" : "NoExec", OpAddress);
}
break;
}
@@ -1296,8 +1368,8 @@ void Decoder::DecodeInstructionsAtEntry(const uint8_t* _InstStream, uint64_t PC,
BlockInfo.TotalInstructionCount = TotalInstructions;
for (auto CodePage : CodePages) {
AddContainedCodePage(PC, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
for (auto& Block : BlockInfo.Blocks) {
Block.IsEntryPoint = BlockInfo.EntryPoints.contains(Block.Entry);
}
}
+32 -8
View File
@@ -2,6 +2,7 @@
#pragma once
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Telemetry.h>
@@ -19,24 +20,32 @@ class ContextImpl;
namespace FEXCore::Frontend {
class Decoder final {
public:
enum class DecodedBlockStatus {
SUCCESS,
INVALID_INST,
NOEXEC_INST,
};
// New Frontend decoding
struct DecodedBlocks final {
uint64_t Entry {};
uint64_t Size {};
uint64_t NumInstructions {};
FEXCore::X86Tables::DecodedInst* DecodedInstructions;
bool HasInvalidInstruction {};
DecodedBlockStatus BlockStatus;
bool IsEntryPoint {};
};
struct DecodedBlockInformation final {
uint64_t TotalInstructionCount;
bool Is64BitMode {};
fextl::vector<DecodedBlocks> Blocks;
fextl::set<uint64_t> EntryPoints;
fextl::set<uint64_t> CodePages; // Start addresses of all pages touching the block
};
Decoder(FEXCore::Context::ContextImpl* ctx);
~Decoder();
void DecodeInstructionsAtEntry(const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst,
std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
Decoder(FEXCore::Core::InternalThreadState* Thread);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
return &BlockInfo;
@@ -56,6 +65,10 @@ public:
PoolObject.DelayedDisownBuffer();
}
void ResetExecutableRangeCache() {
ExecutableRangeBase = ExecutableRangeEnd = 0;
}
private:
// To pass any information from instruction prefixes
// down into the actual instruction handling machinery.
@@ -65,18 +78,22 @@ private:
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Core::InternalThreadState* Thread;
FEXCore::Context::ContextImpl* CTX;
const FEXCore::HLE::SyscallOSABI OSABI {};
bool DecodeInstruction(uint64_t PC);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
bool InstCanContinue() const;
void AddBranchTarget(uint64_t Target);
bool CheckRangeExecutable(uint64_t Address, uint64_t Size);
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset) const;
std::optional<uint8_t> PeekByte(uint8_t Offset);
uint64_t ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) {
InstructionSize += Size;
@@ -87,10 +104,17 @@ private:
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
FEXCore::X86Tables::DecodedInst* DecodedBuffer {};
Utils::FixedSizePooledAllocation<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
Utils::PoolBufferWithTimedRetirement<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint64_t ExecutableRangeBase {};
uint64_t ExecutableRangeEnd {};
bool HitNonExecutableRange {};
const uint8_t* InstStream {};
IR::OpSize GetGPROpSize() const {
return BlockInfo.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
static constexpr size_t MAX_INST_SIZE = 15;
uint8_t InstructionSize {};
@@ -6,7 +6,7 @@
#include "Interface/IR/IR.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SHMStats.h>
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW, bool Force80BitPrecision = false) {
@@ -36,18 +36,48 @@ FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t
return State;
}
FEXCORE_PRESERVE_ALL_ATTR static void HandleX87Exception(const softfloat_state& State, FEXCore::Core::CpuStateFrame* Frame) {
// Check for Invalid Operation exception (bit 0 of X87 status word)
if (State.exceptionFlags & softfloat_flag_invalid) {
Frame->State.flags[FEXCore::X86State::X87FLAG_IE_LOC] = 1;
}
}
// Wrapper for SoftFloat state to handle X87 exceptions
class ScopedSoftFloatState {
public:
FEXCORE_PRESERVE_ALL_ATTR ScopedSoftFloatState(uint16_t FCW, FEXCore::Core::CpuStateFrame* Frame, bool Force80BitPrecision = false)
: State(SoftFloatStateFromFCW(FCW, Force80BitPrecision))
, Frame(Frame) {}
FEXCORE_PRESERVE_ALL_ATTR ~ScopedSoftFloatState() {
HandleX87Exception(State, Frame);
}
// Disable copy and move to ensure RAII semantics
ScopedSoftFloatState(const ScopedSoftFloatState&) = delete;
ScopedSoftFloatState& operator=(const ScopedSoftFloatState&) = delete;
ScopedSoftFloatState(ScopedSoftFloatState&&) = delete;
ScopedSoftFloatState& operator=(ScopedSoftFloatState&&) = delete;
softfloat_state State;
private:
FEXCore::Core::CpuStateFrame* Frame;
};
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle4(uint16_t FCW, float src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(&State.State, src);
}
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle8(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(&State.State, src);
}
};
@@ -55,12 +85,12 @@ template<>
struct OpHandlers<IR::OP_F80CMP> {
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
ScopedSoftFloatState State {FCW, Frame};
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(&State, Src1, Src2, &eq, &lt, &nan);
X80SoftFloat::FCMP(&State.State, Src1, Src2, &eq, &lt, &nan);
if (lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
@@ -78,14 +108,14 @@ template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToF32(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToF32(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToF64(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToF64(&State.State);
}
};
@@ -93,26 +123,26 @@ template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI16(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI16(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI32(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI32(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(src).ToI64(&State);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat(src).ToI64(&State.State);
}
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
auto rv = extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
auto rv = extF80_to_i32(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
@@ -124,14 +154,14 @@ struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i32(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
return extF80_to_i32(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i64(&State, X80SoftFloat(src), softfloat_round_minMag, false);
ScopedSoftFloatState State {FCW, Frame};
return extF80_to_i64(&State.State, X80SoftFloat(src), softfloat_round_minMag, false);
}
};
@@ -152,8 +182,8 @@ template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FRNDINT(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FRNDINT(&State.State, Src1);
}
};
@@ -161,8 +191,8 @@ template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::F2XM1(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::F2XM1(&State.State, Src1);
}
};
@@ -170,8 +200,8 @@ template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FTAN(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FTAN(&State.State, Src1);
}
};
@@ -179,8 +209,8 @@ template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSQRT(&State, Src1);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FSQRT(&State.State, Src1);
}
};
@@ -188,8 +218,8 @@ template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSIN(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FSIN(&State.State, Src1);
}
};
@@ -197,8 +227,17 @@ template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FCOS(&State, Src1);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FCOS(&State.State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SINCOS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegPairType handle(uint16_t FCW, VectorRegType Src1, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame, true};
return FEXCore::MakeVectorRegPair(X80SoftFloat::FSIN(&State.State, Src1), X80SoftFloat::FCOS(&State.State, Src1));
}
};
@@ -222,8 +261,8 @@ template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FADD(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FADD(&State.State, Src1, Src2);
}
};
@@ -231,8 +270,8 @@ template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSUB(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FSUB(&State.State, Src1, Src2);
}
};
@@ -240,8 +279,8 @@ template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FMUL(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FMUL(&State.State, Src1, Src2);
}
};
@@ -249,8 +288,8 @@ template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FDIV(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame};
return X80SoftFloat::FDIV(&State.State, Src1, Src2);
}
};
@@ -258,8 +297,8 @@ template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FYL2X(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FYL2X(&State.State, Src1, Src2);
}
};
@@ -267,8 +306,8 @@ template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FATAN(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FATAN(&State.State, Src1, Src2);
}
};
@@ -276,8 +315,8 @@ template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM1(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FREM1(&State.State, Src1, Src2);
}
};
@@ -285,8 +324,8 @@ template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FREM(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FREM(&State.State, Src1, Src2);
}
};
@@ -294,8 +333,8 @@ template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
softfloat_state State = SoftFloatStateFromFCW(FCW, true);
return X80SoftFloat::FSCALE(&State, Src1, Src2);
ScopedSoftFloatState State {FCW, Frame, true};
return X80SoftFloat::FSCALE(&State.State, Src1, Src2);
}
};
@@ -315,6 +354,21 @@ struct OpHandlers<IR::OP_F64COS> {
}
};
template<>
struct OpHandlers<IR::OP_F64SINCOS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorScalarF64Pair handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
double sin, cos;
#ifdef _WIN32
sin = ::sin(src);
cos = ::cos(src);
#else
sincos(src, &sin, &cos);
#endif
return VectorScalarF64Pair {sin, cos};
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
@@ -380,15 +434,15 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1q, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
X80SoftFloat Src1 = Src1q;
softfloat_state State = SoftFloatStateFromFCW(FCW);
ScopedSoftFloatState State {FCW, Frame};
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(&State, Src1);
Src1 = X80SoftFloat::FRNDINT(&State.State, Src1);
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1.ToI64(&State);
uint64_t Tmp = Src1.ToI64(&State.State);
X80SoftFloat Rv;
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
@@ -50,6 +50,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
Info[Core::OPINDEX_F80SQRT] = {ABIHandlers[FABI_F80_I16_F80_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80SQRT>::handle)};
Info[Core::OPINDEX_F80SIN] = {ABIHandlers[FABI_F80_I16_F80_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80SIN>::handle)};
Info[Core::OPINDEX_F80COS] = {ABIHandlers[FABI_F80_I16_F80_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80COS>::handle)};
Info[Core::OPINDEX_F80SINCOS] = {ABIHandlers[FABI_F80x2_I16_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80SINCOS>::handle)};
Info[Core::OPINDEX_F80XTRACT_EXP] = {ABIHandlers[FABI_F80_I16_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_EXP>::handle)};
Info[Core::OPINDEX_F80XTRACT_SIG] = {ABIHandlers[FABI_F80_I16_F80_PTR],
@@ -82,6 +84,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
// Double Precision Unary
Info[Core::OPINDEX_F64SIN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle)};
Info[Core::OPINDEX_F64COS] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle)};
Info[Core::OPINDEX_F64SINCOS] = {ABIHandlers[FABI_F64x2_I16_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SINCOS>::handle)};
Info[Core::OPINDEX_F64TAN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle)};
Info[Core::OPINDEX_F64F2XM1] = {ABIHandlers[FABI_F64_I16_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle)};
@@ -198,6 +202,12 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
return true; \
}
#define COMMON_UNARYPAIR_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80x2_I16_F80_PTR, Core::OPINDEX_F80##OP}; \
return true; \
}
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80_PTR, Core::OPINDEX_F80##OP}; \
@@ -215,6 +225,12 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
*Info = {FABI_F64_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_UNARYPAIR_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64x2_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_BINARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_I16_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
@@ -228,6 +244,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
COMMON_UNARY_X87_OP(SQRT)
COMMON_UNARY_X87_OP(SIN)
COMMON_UNARY_X87_OP(COS)
COMMON_UNARYPAIR_X87_OP(SINCOS)
COMMON_UNARY_X87_OP(XTRACT_EXP)
COMMON_UNARY_X87_OP(XTRACT_SIG)
COMMON_UNARY_X87_OP(BCDSTORE)
@@ -249,6 +266,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
COMMON_UNARY_F64_OP(TAN)
COMMON_UNARY_F64_OP(SIN)
COMMON_UNARY_F64_OP(COS)
COMMON_UNARYPAIR_F64_OP(SINCOS)
// Double Precision Binary
COMMON_BINARY_F64_OP(FYL2X)
@@ -27,6 +27,8 @@ enum FallbackABI {
FABI_I64_I16_F80_F80_PTR,
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
FABI_UNKNOWN,
File diff suppressed because it is too large. Load diff
+28 -147
View File
@@ -10,18 +10,17 @@ $end_info$
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
LOGMAN_THROW_A_FMT(IROp->ElementSize == IR::OpSize::i32Bit || IROp->ElementSize == IR::OpSize::i64Bit, "Wrong element size");
// Size is the size of each pair element
auto Dst0 = GetReg(Op->OutLo.ID());
auto Dst1 = GetReg(Op->OutHi.ID());
auto Expected0 = GetReg(Op->ExpectedLo.ID());
auto Expected1 = GetReg(Op->ExpectedHi.ID());
auto Desired0 = GetReg(Op->DesiredLo.ID());
auto Desired1 = GetReg(Op->DesiredHi.ID());
auto MemSrc = GetReg(Op->Addr.ID());
auto Dst0 = GetReg(Op->OutLo);
auto Dst1 = GetReg(Op->OutHi);
auto Expected0 = GetReg(Op->ExpectedLo);
auto Expected1 = GetReg(Op->ExpectedHi);
auto Desired0 = GetReg(Op->DesiredLo);
auto Desired1 = GetReg(Op->DesiredHi);
auto MemSrc = GetReg(Op->Addr);
const auto EmitSize = IROp->ElementSize == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
if (CTX->HostFeatures.SupportsAtomics) {
@@ -98,9 +97,9 @@ DEF_OP(CAS) {
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
auto Expected = GetReg(Op->Expected.ID());
auto Desired = GetReg(Op->Desired.ID());
auto MemSrc = GetReg(Op->Addr.ID());
auto Expected = GetReg(Op->Expected);
auto Desired = GetReg(Op->Desired);
auto MemSrc = GetReg(Op->Addr);
auto Dst = GetReg(Node);
if (CTX->HostFeatures.SupportsAtomics) {
@@ -139,115 +138,13 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
if (CTX->HostFeatures.SupportsAtomics) {
staddl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
add(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
staddl(SubEmitSize, TMP2, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
sub(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
stclrl(SubEmitSize, TMP2, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
and_(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicCLR) {
auto Op = IROp->C<IR::IROp_AtomicCLR>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
if (CTX->HostFeatures.SupportsAtomics) {
stclrl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
bic(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
if (CTX->HostFeatures.SupportsAtomics) {
stsetl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
orr(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
@@ -261,21 +158,6 @@ DEF_OP(AtomicXor) {
}
}
DEF_OP(AtomicNeg) {
auto Op = IROp->C<IR::IROp_AtomicNeg>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
@@ -284,8 +166,8 @@ DEF_OP(AtomicSwap) {
"d CAS "
"size");
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = OpSize == IR::OpSize::i64Bit ? ARMEmitter::SubRegSize::i64Bit :
@@ -310,8 +192,8 @@ DEF_OP(AtomicFetchAdd) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
ldaddal(SubEmitSize, Src, GetReg(Node), MemSrc);
@@ -331,8 +213,8 @@ DEF_OP(AtomicFetchSub) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
@@ -353,8 +235,8 @@ DEF_OP(AtomicFetchAnd) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
@@ -375,8 +257,8 @@ DEF_OP(AtomicFetchCLR) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
ldclral(SubEmitSize, Src, GetReg(Node), MemSrc);
@@ -396,8 +278,8 @@ DEF_OP(AtomicFetchOr) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
ldsetal(SubEmitSize, Src, GetReg(Node), MemSrc);
@@ -417,8 +299,8 @@ DEF_OP(AtomicFetchXor) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
ldeoral(SubEmitSize, Src, GetReg(Node), MemSrc);
@@ -438,7 +320,7 @@ DEF_OP(AtomicFetchNeg) {
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto MemSrc = GetReg(Op->Addr);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
@@ -452,7 +334,7 @@ DEF_OP(AtomicFetchNeg) {
DEF_OP(TelemetrySetValue) {
#ifndef FEX_DISABLE_TELEMETRY
auto Op = IROp->C<IR::IROp_TelemetrySetValue>();
auto Src = GetReg(Op->Value.ID());
auto Src = GetReg(Op->Value);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.TelemetryValueAddresses[Op->TelemetryValueIndex]));
@@ -473,5 +355,4 @@ DEF_OP(TelemetrySetValue) {
#endif
}
#undef DEF_OP
} // namespace FEXCore::CPU
+133 -37
View File
@@ -18,7 +18,6 @@ $end_info$
#include <FEXCore/Utils/MathUtils.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(CallbackReturn) {
// spill back to CTX
@@ -59,30 +58,111 @@ DEF_OP(ExitFunction) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef _M_ARM_64EC
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
} else {
#endif
// Align to 16 byte to allow atomic patching of the following 16 byte
// of code (excluding the RIP data) on platforms that support LSE2
Align16B();
// In order to support direct branches without constantly hitting the L1 cache, we emit a call to a block linker,
// this will compile the branch target block when it is hit and replace the branch to the linker at the callsite
// with a direct branch to the destination block. Upon invalidation of the target block the backpatch is undone.
//
// In addition, to avoid needing to lookup in the cache for returns and any indirect branch prediction penalty,
// a shadow stack of <GuestReturnRIP, HostReturnPC> pairs is maintained, acting as a first level cache for any
// return operations. As the guest may not balance calls and returns exactly, an exception handler is expected to
// be installed by the frontend, to reset the shadow stack to the middle of its valid bounds on overflow/underflow.
// This shadow stack is also cleared on block invalidation operations or codebuffer switches, to ensure all pointed-to
// host code is always valid.
// This code will be backpatched by Arm64JITCore_ExitFunctionLink, below is an enumeration of all the possible cases.
// Jump thunks are emitted in JIT.cpp after compilation of the entire multiblock.
//
// Call with known return block - unlinked
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl JmpThunk00
// JmpThunk00:
// 00: b 0x8
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode
// 18: GuestRIP
// 20: CallerOffset
//
// Call with known return block after backpatching - linked in branch immediate range
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl HostCode - MODIFIED
//
// Call with known return block after backpatching - linked out of range
// 00: adr TMP1, 0xC
// 04: stp RetReg, TMP1, [SpReg, -0x10]!
// 08: bl JmpThunk00
// JmpThunk00:
// 00: ldr TMP1, 0x10 - MODIFIED 2nd
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode - MODIFIED 1st
// 18: GuestRIP
// 20: CallerOffset
//
// Jump - unlinked
// 00: b JmpThunk00
// JmpThunk00:
// 00: b 0x8
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode
// 18: GuestRIP
// 20: CallerOffset
//
// Jump after backpatching - linked in branch immediate range
// 00: b HostCode - MODIFIED
//
// Jump after backpatching - linked out of range
// 00: b JmpThunk00
// JmpThunk00:
// 00: ldr TMP1, 0x10 - MODIFIED 2nd
// 04: br TMP1
// 08: ldr TMP1, <Shared exit linker>
// 0c: blr TMP1
// 10: HostCode - MODIFIED 1st
// 18: GuestRIP
// 20: CallerOffset
ARMEmitter::ForwardLabel l_BranchHost;
ldr(TMP1, &l_BranchHost);
blr(TMP1);
ARMEmitter::ForwardLabel l_CallReturn;
if (Op->Hint == IR::BranchHint::Call) {
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
}
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
dc64(NewRIP);
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
Bind(&l_CallReturn);
#ifdef _M_ARM_64EC
}
#endif
} else {
ARMEmitter::ForwardLabel SkipFullLookup;
auto RipReg = GetReg(Op->NewRIP);
ARMEmitter::ForwardLabel FullLookup;
auto RipReg = GetReg(Op->NewRIP.ID());
if (Op->Hint == IR::BranchHint::Return) {
// First try to pop from the call-ret stack, otherwise follow the normal path (but ending in a ret)
ldp<ARMEmitter::IndexType::POST>(TMP1, TMP2, REG_CALLRET_SP, 0x10);
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
}
// L1 Cache
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
@@ -94,24 +174,40 @@ DEF_OP(ExitFunction) {
ubfiz(ARMEmitter::Size::i64Bit, TMP4, RipReg, 4, 20);
add(TMP1, TMP1, TMP4);
// Note: sub+cbnz used over cmp+br to preserve flags.
ldp<ARMEmitter::IndexType::OFFSET>(TMP2, TMP1, TMP1, 0);
sub(TMP1, TMP1, RipReg.X());
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(TMP2);
Bind(&FullLookup);
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
// Note: sub+cbnz used over cmp+br to preserve flags.
sub(TMP1, TMP1, RipReg.X());
cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
br(TMP1);
Bind(&SkipFullLookup);
if (Op->Hint == IR::BranchHint::Call) {
ARMEmitter::ForwardLabel l_CallReturn;
if (!Op->CallReturnBlock.IsInvalid()) {
auto CallReturnAddressReg = GetReg(Op->CallReturnAddress).X();
PendingCallReturnTargetLabel = &CallReturnTargets.try_emplace(Op->CallReturnBlock.ID()).first->second;
adr(TMP1, &l_CallReturn);
stp<ARMEmitter::IndexType::PRE>(CallReturnAddressReg, TMP1, REG_CALLRET_SP, -0x10);
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
blr(TMP2);
Bind(&l_CallReturn);
} else if (Op->Hint == IR::BranchHint::Return) {
ret(TMP2);
} else {
br(TMP2);
}
}
}
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto Target = Op->TargetBlock.ID();
const auto Target = Op->TargetBlock;
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
PendingTargetLabel = &JumpTargets.try_emplace(Target.ID()).first->second;
}
DEF_OP(CondJump) {
@@ -125,10 +221,10 @@ DEF_OP(CondJump) {
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
auto Reg = GetReg(Op->Cmp1.ID());
auto Reg = GetReg(Op->Cmp1);
const auto Size = Op->CompareSize == IR::OpSize::i32Bit ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(isConst, "CondJump: Expected constant source");
if (Op->Cond.Val == FEXCore::IR::COND_EQ) {
@@ -184,7 +280,7 @@ DEF_OP(Syscall) {
if (Op->Header.Args[i].IsInvalid()) {
continue;
}
str(GetReg(Op->Header.Args[i].ID()).X(), ARMEmitter::Reg::rsp, i * 8);
str(GetReg(Op->Header.Args[i]).X(), ARMEmitter::Reg::rsp, i * 8);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj));
@@ -244,7 +340,7 @@ DEF_OP(InlineSyscall) {
break;
}
auto Reg = GetReg(Op->Header.Args[i].ID());
auto Reg = GetReg(Op->Header.Args[i]);
if (Reg == ARMEmitter::Reg::r8 || Reg == ARMEmitter::Reg::r4 || Reg == ARMEmitter::Reg::r5) {
SpillMask |= (1U << Reg.Idx());
@@ -274,7 +370,7 @@ DEF_OP(InlineSyscall) {
break;
}
auto Reg = GetReg(Op->Header.Args[i].ID());
auto Reg = GetReg(Op->Header.Args[i]);
if (SpillMask & (1U << Reg.Idx())) {
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RDX, and RSP. Which have just been spilled
@@ -292,7 +388,7 @@ DEF_OP(InlineSyscall) {
break;
}
mov(EmitSize, RegArgs[i].R(), GetReg(Op->Header.Args[i].ID()));
mov(EmitSize, RegArgs[i].R(), GetReg(Op->Header.Args[i]));
}
}
@@ -325,7 +421,7 @@ DEF_OP(Thunk) {
PushDynamicRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr.ID()));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
@@ -412,8 +508,9 @@ DEF_OP(ThreadRemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
mov(ARMEmitter::Size::i64Bit, TMP2, GetReg(Op->Function.ID()));
mov(ARMEmitter::Size::i64Bit, TMP3, GetReg(Op->Leaf.ID()));
isb();
mov(ARMEmitter::Size::i64Bit, TMP2, GetReg(Op->Function));
mov(ARMEmitter::Size::i64Bit, TMP3, GetReg(Op->Leaf));
PushDynamicRegs(TMP4);
SpillStaticRegs(TMP4);
@@ -446,10 +543,10 @@ DEF_OP(CPUID) {
// Results are in x0, x1
// Results want to be 4xi32 scalars
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutEAX.ID()), TMP1);
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutECX.ID()), TMP2);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEBX.ID()), TMP1, 32, 32);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEDX.ID()), TMP2, 32, 32);
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutEAX), TMP1);
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutECX), TMP2);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEBX), TMP1, 32, 32);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEDX), TMP2, 32, 32);
}
DEF_OP(XGetBV) {
@@ -458,7 +555,7 @@ DEF_OP(XGetBV) {
PushDynamicRegs(TMP4);
SpillStaticRegs(TMP4);
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function));
// x0 = CPUID Handler
// x1 = XCR Function
@@ -479,9 +576,8 @@ DEF_OP(XGetBV) {
PopDynamicRegs();
// Results are in x0, need to split into i32 parts
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutEAX.ID()), TMP1);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEDX.ID()), TMP1, 32, 32);
mov(ARMEmitter::Size::i32Bit, GetReg(Op->OutEAX), TMP1);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Op->OutEDX), TMP1, 32, 32);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -9,7 +9,6 @@ $end_info$
#include "Interface/Context/Context.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
@@ -23,8 +22,8 @@ DEF_OP(VInsGPR) {
const auto ElementsPer128Bit = IR::NumElements(IR::OpSize::i128Bit, ElementSize);
const auto Dst = GetVReg(Node);
const auto DestVector = GetVReg(Op->DestVector.ID());
const auto Src = GetReg(Op->Src.ID());
const auto DestVector = GetVReg(Op->DestVector);
const auto Src = GetReg(Op->Src);
if (HostSupportsSVE256 && Is256Bit) {
const auto ElementSizeBits = IR::OpSizeAsBits(ElementSize);
@@ -88,7 +87,7 @@ DEF_OP(VInsGPR) {
DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
auto Dst = GetVReg(Node);
auto Src = GetReg(Op->Src.ID());
auto Src = GetReg(Op->Src);
switch (Op->Header.ElementSize) {
case IR::OpSize::i8Bit:
@@ -109,8 +108,8 @@ DEF_OP(VLoadTwoGPRs) {
const auto Op = IROp->C<IR::IROp_VLoadTwoGPRs>();
const auto Dst = GetVReg(Node);
const auto SrcLower = GetReg(Op->Lower.ID());
const auto SrcUpper = GetReg(Op->Upper.ID());
const auto SrcLower = GetReg(Op->Lower);
const auto SrcUpper = GetReg(Op->Upper);
fmov(ARMEmitter::Size::i64Bit, Dst.D(), SrcLower);
fmov(ARMEmitter::Size::i64Bit, Dst.D(), SrcUpper, true);
}
@@ -120,7 +119,7 @@ DEF_OP(VDupFromGPR) {
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src = GetReg(Op->Src.ID());
const auto Src = GetReg(Op->Src);
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -141,7 +140,7 @@ DEF_OP(Float_FromGPR_S) {
const uint16_t Conv = (ElementSize << 8) | IR::OpSizeToSize(Op->SrcElementSize);
auto Dst = GetVReg(Node);
auto Src = GetReg(Op->Src.ID());
auto Src = GetReg(Op->Src);
switch (Conv) {
case 0x0204: { // Half <- int32_t
@@ -179,7 +178,7 @@ DEF_OP(Float_FToF) {
const uint16_t Conv = (IR::OpSizeToSize(Op->Header.ElementSize) << 8) | IR::OpSizeToSize(Op->SrcElementSize);
auto Dst = GetVReg(Node);
auto Src = GetVReg(Op->Scalar.ID());
auto Src = GetVReg(Op->Scalar);
switch (Conv) {
case 0x0204: { // Half <- Float
@@ -220,7 +219,7 @@ DEF_OP(Vector_SToF) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B;
scvtf(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
@@ -253,7 +252,7 @@ DEF_OP(Vector_FToZS) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B;
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
@@ -286,7 +285,7 @@ DEF_OP(Vector_FToS) {
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B;
@@ -294,7 +293,7 @@ DEF_OP(Vector_FToS) {
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Dst.Z(), SubEmitSize);
} else {
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (OpSize == IR::OpSize::i64Bit) {
frinti(SubEmitSize, Dst.D(), Vector.D());
fcvtzs(SubEmitSize, Dst.D(), Dst.D());
@@ -317,7 +316,7 @@ DEF_OP(Vector_FToF) {
const auto Conv = (IR::OpSizeToSize(ElementSize) << 8) | IR::OpSizeToSize(Op->SrcElementSize);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
// Curiously, FCVTLT and FCVTNT have no bottom variants,
@@ -381,7 +380,7 @@ DEF_OP(VFCVTL2) {
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
fcvtl2(SubEmitSize, Dst.D(), Vector.D());
}
@@ -392,8 +391,8 @@ DEF_OP(VFCVTN2) {
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto VectorLower = GetVReg(Op->VectorLower.ID());
const auto VectorUpper = GetVReg(Op->VectorUpper.ID());
const auto VectorLower = GetVReg(Op->VectorLower);
const auto VectorUpper = GetVReg(Op->VectorUpper);
auto Lower = VectorLower;
if (Dst != VectorLower) {
@@ -418,7 +417,7 @@ DEF_OP(Vector_FToI) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
@@ -479,7 +478,7 @@ DEF_OP(Vector_FToISized) {
LOGMAN_THROW_A_FMT(CTX->HostFeatures.SupportsFRINTTS, "Need FRINTTS for Vector_FToISized");
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (ElementSize == IROp->Size) {
// See above
@@ -533,7 +532,7 @@ DEF_OP(Vector_F64ToI32) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE128 || HostSupportsSVE256) {
const auto Mask = Is256Bit ? PRED_TMP_32B.Merging() : PRED_TMP_16B.Merging();
// First step is to round the f64 values to integrals (frint*)
@@ -583,5 +582,4 @@ DEF_OP(Vector_F64ToI32) {
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -8,11 +8,10 @@ $end_info$
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(VAESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
aesimc(GetVReg(Node), GetVReg(Op->Vector.ID()));
aesimc(GetVReg(Node), GetVReg(Op->Vector));
}
DEF_OP(VAESEnc) {
@@ -20,9 +19,9 @@ DEF_OP(VAESEnc) {
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
const auto Key = GetVReg(Op->Key);
const auto State = GetVReg(Op->State);
const auto ZeroReg = GetVReg(Op->ZeroReg);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
@@ -45,9 +44,9 @@ DEF_OP(VAESEncLast) {
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
const auto Key = GetVReg(Op->Key);
const auto State = GetVReg(Op->State);
const auto ZeroReg = GetVReg(Op->ZeroReg);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
@@ -68,9 +67,9 @@ DEF_OP(VAESDec) {
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
const auto Key = GetVReg(Op->Key);
const auto State = GetVReg(Op->State);
const auto ZeroReg = GetVReg(Op->ZeroReg);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
@@ -93,9 +92,9 @@ DEF_OP(VAESDecLast) {
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
const auto Key = GetVReg(Op->Key);
const auto State = GetVReg(Op->State);
const auto ZeroReg = GetVReg(Op->ZeroReg);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
@@ -114,9 +113,9 @@ DEF_OP(VAESDecLast) {
DEF_OP(VAESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
const auto Dst = GetVReg(Node);
const auto Src = GetVReg(Op->Src.ID());
const auto Swizzle = GetVReg(Op->KeyGenTBLSwizzle.ID());
auto ZeroReg = GetVReg(Op->ZeroReg.ID());
const auto Src = GetVReg(Op->Src);
const auto Swizzle = GetVReg(Op->KeyGenTBLSwizzle);
auto ZeroReg = GetVReg(Op->ZeroReg);
if (Dst == ZeroReg) {
// Seriously? ZeroReg ended up being the destination register?
@@ -148,8 +147,8 @@ DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
const auto Dst = GetReg(Node);
const auto Src1 = GetReg(Op->Src1.ID());
const auto Src2 = GetReg(Op->Src2.ID());
const auto Src1 = GetReg(Op->Src1);
const auto Src2 = GetReg(Op->Src2);
switch (Op->SrcSize) {
case IR::OpSize::i8Bit: crc32cb(Dst.W(), Src1.W(), Src2.W()); break;
@@ -164,7 +163,7 @@ DEF_OP(VSha1H) {
auto Op = IROp->C<IR::IROp_VSha1H>();
const auto Dst = GetVReg(Node);
const auto Src = GetVReg(Op->Src.ID());
const auto Src = GetVReg(Op->Src);
sha1h(Dst.S(), Src.S());
}
@@ -173,9 +172,9 @@ DEF_OP(VSha1C) {
auto Op = IROp->C<IR::IROp_VSha1C>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Src3 = GetVReg(Op->Src3);
if (Dst == Src1) {
sha1c(Dst, Src2.S(), Src3);
@@ -193,9 +192,9 @@ DEF_OP(VSha1M) {
auto Op = IROp->C<IR::IROp_VSha1M>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Src3 = GetVReg(Op->Src3);
if (Dst == Src1) {
sha1m(Dst, Src2.S(), Src3);
@@ -213,9 +212,9 @@ DEF_OP(VSha1P) {
auto Op = IROp->C<IR::IROp_VSha1P>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Src3 = GetVReg(Op->Src3);
if (Dst == Src1) {
sha1p(Dst, Src2.S(), Src3);
@@ -233,8 +232,8 @@ DEF_OP(VSha1SU1) {
auto Op = IROp->C<IR::IROp_VSha1SU1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
if (Dst == Src1) {
sha1su1(Dst, Src2);
@@ -252,9 +251,9 @@ DEF_OP(VSha256H) {
auto Op = IROp->C<IR::IROp_VSha256H>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Src3 = GetVReg(Op->Src3);
if (Dst == Src1) {
sha256h(Dst, Src2, Src3);
@@ -272,9 +271,9 @@ DEF_OP(VSha256H2) {
auto Op = IROp->C<IR::IROp_VSha256H2>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src3 = GetVReg(Op->Src3.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Src3 = GetVReg(Op->Src3);
if (Dst == Src1) {
sha256h2(Dst, Src2, Src3);
@@ -292,8 +291,8 @@ DEF_OP(VSha256U0) {
auto Op = IROp->C<IR::IROp_VSha256U0>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
if (Dst == Src1) {
sha256su0(Dst, Src2);
@@ -308,8 +307,8 @@ DEF_OP(VSha256U1) {
auto Op = IROp->C<IR::IROp_VSha256U1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
if (Dst != Src1 && Dst != Src2) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
@@ -326,8 +325,8 @@ DEF_OP(PCLMUL) {
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
@@ -346,5 +345,4 @@ DEF_OP(PCLMUL) {
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
+358 -151
View File
@@ -40,36 +40,33 @@ $end_info$
#include <string.h>
#include <limits>
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
namespace {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
struct DivRem {
uint64_t Quotient;
uint64_t Remainder;
};
static struct DivRem LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
return Res;
return {
.Quotient = (uint64_t)(Source / Divisor),
.Remainder = (uint64_t)(Source % Divisor),
};
}
static int64_t LDIV(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
static struct DivRem
LDIV(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
return {
.Quotient = (uint64_t)(Source / Divisor),
.Remainder = (uint64_t)(Source % Divisor),
};
}
static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
return Res;
}
static int64_t LREM(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
static void PrintValue(uint64_t Value) {
static void
PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
@@ -80,13 +77,23 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
namespace FEXCore::CPU {
void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_MSG_A_FMT("Unhandled IR Op: {}", FEXCore::IR::GetName(IROp->Op));
#endif
} else {
auto FillF80x2Result = [&](auto DstLo, auto DstHi) {
mov(DstLo.Q(), VTMP1.Q());
mov(DstHi.Q(), VTMP2.Q());
};
auto FillF64x2Result = [&](auto DstLo, auto DstHi) {
fmov(DstLo.D(), VTMP1.D());
fmov(DstHi.D(), VTMP2.D());
};
auto FillF80Result = [&]() {
const auto Dst = GetVReg(Node);
mov(Dst.Q(), VTMP1.Q());
@@ -125,7 +132,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.S(), Src1.S());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -143,7 +150,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -162,7 +169,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// tmp2 (x1/x11): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetReg(IROp->Args[0].ID());
const auto Src1 = GetReg(IROp->Args[0]);
// Need to sign or zero extend this for the dispatcher handler.
if (Info.ABI == FABI_F80_I16_I16_PTR) {
@@ -186,7 +193,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -204,7 +211,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -222,7 +229,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): vector source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -232,6 +239,30 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
FillF64Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
// vtmp1 (v0/v16): vector source
// vtmp2 (v1/v16): vector source
#ifdef VIXL_SIMULATOR
LOGMAN_THROW_A_FMT(CTX->Config.DisableVixlIndirectCalls, "Vector register pairs unsupported by simulator currently");
#endif
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0]);
const auto DstLo = GetVReg(IROp->Args[1]);
const auto DstHi = GetVReg(IROp->Args[2]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
FillF64x2Result(DstLo, DstHi);
} break;
case FABI_F64_I16_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
@@ -241,8 +272,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp2 (v1/v17): vector source 2
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
const auto Src2 = GetVReg(IROp->Args[1]);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
@@ -262,7 +293,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -280,7 +311,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -298,7 +329,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): source
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -317,8 +348,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp2 (v1/v17): vector source 2
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
const auto Src2 = GetVReg(IROp->Args[1]);
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
@@ -337,7 +368,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp1 (v0/v16): vector source 1
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
@@ -348,6 +379,31 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
FillF80Result();
} break;
case FABI_F80x2_I16_F80_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
// vtmp1 (v0/v16): vector source 1
// vtmp2 (v1/v16): vector source 2
#ifdef VIXL_SIMULATOR
LOGMAN_THROW_A_FMT(CTX->Config.DisableVixlIndirectCalls, "Vector register pairs unsupported by simulator currently");
#endif
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0]);
const auto DstLo = GetVReg(IROp->Args[1]);
const auto DstHi = GetVReg(IROp->Args[2]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
FillF80x2Result(DstLo, DstHi);
} break;
case FABI_F80_I16_F80_F80_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
@@ -356,8 +412,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
// vtmp2 (v1/v17): vector source 2
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
const auto Src1 = GetVReg(IROp->Args[0]);
const auto Src2 = GetVReg(IROp->Args[1]);
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
@@ -385,16 +441,16 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
stp<ARMEmitter::IndexType::PRE>(TMP1, ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
const auto SrcRAX = GetReg(Op->RAX);
const auto SrcRDX = GetReg(Op->RDX);
const auto Control = Op->Control;
mov(TMP1, SrcRAX.X());
mov(TMP2, SrcRDX.X());
movz(ARMEmitter::Size::i32Bit, TMP3, Control);
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Src1 = GetVReg(Op->LHS);
const auto Src2 = GetVReg(Op->RHS);
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
@@ -414,8 +470,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Src1 = GetVReg(Op->LHS);
const auto Src2 = GetVReg(Op->RHS);
const auto Control = Op->Control;
mov(VTMP1.Q(), Src1.Q());
@@ -439,71 +495,132 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
// Emit new 16 bytes of code to a temporary patch, then atomically apply it
__uint128_t Patch;
ARMEmitter::Emitter emit((uint8_t*)&Patch, sizeof(Patch));
emit.ldr(TMP1, 8); // PC-relative value pointing to constant after blr
emit.blr(TMP1);
emit.dc64(Frame->Pointers.Common.ExitFunctionLinker);
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
auto branch = reinterpret_cast<__uint128_t*>((uintptr_t)Record - 8);
std::atomic_ref<__uint128_t>(*branch).store(Patch, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache((void*)branch, sizeof(*branch));
// Replace the patched callsite with a branch to the jump thunk.
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
if (Call) {
BranchEmit.bl(BranchOffset);
} else {
BranchEmit.b(BranchOffset);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
// No need to reset HostCode here as the exit linker pointer is stored separately, and if the block is relinked it will be updated.
}
uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
bool TFSet = Thread->CurrentFrame->State.flags[X86State::RFLAG_TF_RAW_LOC];
uintptr_t HostCode {};
auto GuestRip = Record->GuestRIP;
if (!TFSet) {
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (TFSet || !HostCode) {
if (TFSet) {
// If TF is set, the cache must be skipped as different code needs to be generated.
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
} else {
{
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (!HostCode) {
// Hold a reference to the code buffer, to avoid linking unmapped code if compilation triggers a recreation.
auto CodeBuffer = static_cast<Arm64JITCore*>(Thread->CPUBackend.get())->CurrentCodeBuffer;
HostCode = static_cast<Context::ContextImpl*>(Thread->CTX)->CompileBlock(Frame, GuestRip, 0);
if (Thread->LookupCache->Shared != CodeBuffer->LookupCache.get()) {
return HostCode;
}
}
}
uintptr_t branch = (uintptr_t)(Record)-8;
LOGMAN_THROW_A_FMT((branch % 16) == 0, "Incorrect alignment for block linking record");
// See ExitFunction in BranchOps.cpp for an assembly level view of the handled cases.
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = HostCode / 4 - CallerAddress / 4;
auto offset = HostCode / 4 - branch / 4;
if (ARMEmitter::Emitter::IsInt26(offset)) {
// This is the optimal case, where the target can be encoded in a single instruction.
// Atomically patch the code with a relative branch.
const uint32_t Patch = (0b0001'01 << 26) | (offset & ((1u << 26) - 1));
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(branch)).store(Patch, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache((void*)branch, 4);
uint32_t ExpectedKnownCallMarkerInst = 0;
ARMEmitter::Emitter ExpectedKnownCallMarkerEmit(reinterpret_cast<uint8_t*>(&ExpectedKnownCallMarkerInst), 4);
ExpectedKnownCallMarkerEmit.adr(TMP1, 0xC);
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
// Lock here is necessary to prevent simultaneous linking and delinking
auto lk = Thread->LookupCache->AcquireLock();
// For non-calls, this would extend into the block's code, however that's fine as an out-of-range adr would never
// be generated avoiding any false positives.
uintptr_t KnownCallMarkerAddr = CallerAddress - 0x8;
uint32_t KnownCallMarkerInst = *reinterpret_cast<uint32_t*>(KnownCallMarkerAddr);
if (ARMEmitter::Emitter::IsInt26(BranchOffset)) {
// Directly patch the callsite with the appropriate branch instruction.
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, true);
});
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(GuestRip, Record, [](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
});
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
} else {
// fallback case - do a soft-er link by patching the pointer
std::atomic_ref<uint64_t>(Record->HostBranch).store(HostCode, std::memory_order::seq_cst);
// This case is common between calls and jumps as the thunk callsite can be left untouched.
std::atomic_ref<uint64_t>(Record->HostCode).store(HostCode, std::memory_order::seq_cst);
#ifdef _M_ARM_64
// Make memory write visible to other threads reading the same location
asm volatile("dc cvau, %0; dsb ish" : : "r"(Record->HostBranch) :);
asm volatile("dc cvau, %0; dsb ish" : : "r"(Record->HostCode) :);
#endif
}
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
uint32_t LdrInst = 0;
ARMEmitter::Emitter LdrEmit(reinterpret_cast<uint8_t*>(&LdrInst), 4);
LdrEmit.ldr(TMP1, reinterpret_cast<uint64_t>(&Record->HostCode) - JumpThunkStartAddress);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(LdrInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
Thread->LookupCache->AddBlockLink(GuestRip, Record, IndirectBlockDelinker);
}
return HostCode;
}
void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::NodeID Node) {}
void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::Ref Node) {}
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
: CPUBackend(*ctx, Thread)
, Arm64Emitter(ctx)
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE128}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsSVE256}
, HostSupportsAVX256 {ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256}
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
, CTX {ctx} {
, CTX {ctx}
, TempAllocator(ctx->CPUBackendAllocator, 0) {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -539,19 +656,17 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = PMF.GetVTableEntry(CTX->SyscallHandler);
}
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore::ExitFunctionLink);
// Platform Specific
auto& AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
AArch64.LUREM = reinterpret_cast<uint64_t>(LUREM);
AArch64.LREM = reinterpret_cast<uint64_t>(LREM);
}
// Must be done after Dispatcher init
ClearCache();
CurrentCodeBuffer = CodeBuffers.GetLatest();
ThreadState->LookupCache->Shared = CurrentCodeBuffer->LookupCache.get();
// Setup dynamic dispatch.
if (ParanoidTSO()) {
@@ -570,16 +685,24 @@ void Arm64JITCore::EmitDetectionString() {
}
void Arm64JITCore::ClearCache() {
// Get the backing code buffer
// NOTE: Holding on to the reference here is required to ensure validity of the WriteLock mutex
auto PrevCodeBuffer = CurrentCodeBuffer;
std::lock_guard lk(PrevCodeBuffer->LookupCache->WriteLock);
auto CodeBuffer = GetEmptyCodeBuffer();
SetBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache);
}
Arm64JITCore::~Arm64JITCore() {}
bool Arm64JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const {
if (WNode.IsImmediate()) {
return false;
}
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINECONSTANT) {
@@ -594,6 +717,10 @@ bool Arm64JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_
}
bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const {
if (WNode.IsImmediate()) {
return false;
}
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
@@ -612,22 +739,6 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
}
}
FEXCore::IR::RegisterClassType Arm64JITCore::GetRegClass(IR::NodeID Node) const {
return FEXCore::IR::RegisterClassType {GetPhys(Node).Class};
}
bool Arm64JITCore::IsFPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
}
bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void Arm64JITCore::EmitInterruptChecks(bool CheckTF) {
if (CheckTF) {
ARMEmitter::ForwardLabel l_TFUnset;
@@ -688,24 +799,47 @@ void Arm64JITCore::EmitInterruptChecks(bool CheckTF) {
#endif
}
void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF) {
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &HeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
EmitInterruptChecks(CheckTF);
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, const FEXCore::IR::RegisterAllocationData* RAData,
bool CheckTF) {
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
CallReturnTargets.clear();
PendingJumpThunks.clear();
uint32_t SSACount = IR->GetSSACount();
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
this->IR = IR;
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
if ((GetCursorOffset() + BufferRange) > (CurrentCodeBuffer->Size - Utils::FEX_PAGE_SIZE)) {
CTX->ClearCodeCache(ThreadState);
}
uint32_t BufferRange = 0x1000 + SSACount * 24;
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
@@ -715,9 +849,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
JITCodeHeader* CodeHeader = GetCursorAddress<JITCodeHeader*>();
CursorIncrement(sizeof(JITCodeHeader));
#ifdef VIXL_DISASSEMBLER
const auto DisasmBegin = GetCursorAddress<const vixl::aarch64::Instruction*>();
#endif
auto CodeBegin = GetCursorAddress<uint8_t*>();
// AAPCS64
// r30 = LR
@@ -739,34 +871,15 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// X1-X3 = Temp
// X4-r18 = RA
CodeData.BlockEntry = GetCursorAddress<uint8_t*>();
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &JITCodeHeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
EmitInterruptChecks(CheckTF);
SpillSlots = RAData->SpillSlots();
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
SpillSlots = IR->SpillSlots();
PendingTargetLabel = nullptr;
PendingCallReturnTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
@@ -778,6 +891,33 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
b(PendingTargetLabel);
PendingTargetLabel = nullptr;
}
if (BlockIROp->EntryPoint) {
uint64_t BlockStartRIP = Entry + BlockIROp->GuestEntryOffset;
const auto IsReturnTarget = CallReturnTargets.try_emplace(Node).first;
if (PendingTargetLabel) {
// If there is a fallthrough branch to this block, skip over the entrypoint code.
b(&IsTarget->second);
} else if (PendingCallReturnTargetLabel && PendingCallReturnTargetLabel != &IsReturnTarget->second) {
// If we just emitted a call, but the block we're now emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
}
PendingCallReturnTargetLabel = nullptr;
Bind(&IsReturnTarget->second);
CodeData.EntryPoints.emplace(BlockStartRIP, GetCursorAddress<uint8_t*>());
DebugData->GuestOpcodes.push_back({BlockIROp->GuestEntryOffset, GetCursorAddress<uint8_t*>() - CodeData.BlockBegin});
EmitEntryPoint(JITCodeHeaderLabel, CheckTF);
}
if (PendingCallReturnTargetLabel) {
// If there is still a pending call return target, then the block we're emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
PendingCallReturnTargetLabel = nullptr;
}
PendingTargetLabel = nullptr;
@@ -785,25 +925,22 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
switch (IROp->Op) {
#define REGISTER_OP_RT(op, x) \
case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, ID); break
case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, CodeNode); break
#define REGISTER_OP(op, x) \
case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, ID); break
case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, CodeNode); break
#define IROP_DISPATCH_DISPATCH
#include <FEXCore/IR/IRDefines_Dispatch.inc>
#undef REGISTER_OP
default: Op_Unhandled(IROp, ID); break;
default: Op_Unhandled(IROp, CodeNode); break;
}
}
if (DebugData) {
DebugData->Subblocks.push_back({static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t*>() - BlockStartHostCode)});
}
DebugData->Subblocks.push_back({static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockBegin),
static_cast<uint32_t>(GetCursorAddress<uint8_t*>() - BlockStartHostCode)});
}
// Make sure last branch is generated. It certainly can't be eliminated here.
@@ -812,8 +949,32 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
PendingTargetLabel = nullptr;
// CodeSize not including the tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
ARMEmitter::ForwardLabel l_ExitLink;
for (auto& PendingJumpThunk : PendingJumpThunks) {
// Align as 64-bit atomics are used on the HostCode field.
Align(8);
ARMEmitter::ForwardLabel l_DoLink;
uint64_t ThunkAddress = GetCursorAddress<uint64_t>();
Bind(&PendingJumpThunk.Label);
b(&l_DoLink);
br(TMP1);
Bind(&l_DoLink);
ldr(TMP1, &l_ExitLink);
blr(TMP1);
// This is a ExitFunctionLinkData struct
Bind(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
Bind(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
Align(alignof(JITCodeTail));
@@ -878,7 +1039,54 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
JITBlockTail->Size = CodeData.Size;
ClearICache(CodeData.BlockBegin, CodeOnlySize);
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
{
auto CodeBufferLock = std::unique_lock {CodeBuffers.CodeBufferWriteMutex};
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
LOGMAN_THROW_A_FMT(CurrentCodeBuffer->LookupCache.get() == ThreadState->LookupCache->Shared, "INVARIANT VIOLATED: SharedLookupCache "
"doesn't match up!\n");
if (auto Prev = CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(ThreadState->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache);
}
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
SetBuffer(CurrentCodeBuffer->Ptr, CurrentCodeBuffer->Size);
SetCursorOffset(AlignUp(CodeBuffers.LatestOffset, 16));
if ((GetCursorOffset() + TempSize) > (CurrentCodeBuffer->Size - Utils::FEX_PAGE_SIZE)) {
CTX->ClearCodeCache(ThreadState);
}
Align16B();
CodeBuffers.LatestOffset = GetCursorOffset();
}
// Adjust host addresses
const auto Delta = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
CodeData.BlockBegin += Delta;
for (auto& EntryPoint : CodeData.EntryPoints) {
EntryPoint.second += Delta;
}
CodeBegin += Delta;
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
CodeBuffers.LatestOffset = GetCursorOffset();
}
TempAllocator.DelayedDisownBuffer();
ClearICache(CodeBegin, CodeOnlySize);
#ifdef VIXL_DISASSEMBLER
if (Disassemble() & FEXCore::Config::Disassemble::STATS) {
@@ -892,7 +1100,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
if (Disassemble() & FEXCore::Config::Disassemble::BLOCKS) {
const auto DisasmEnd = reinterpret_cast<const vixl::aarch64::Instruction*>(JITBlockTailLocation);
const auto DisasmBegin = reinterpret_cast<const vixl::aarch64::Instruction*>(CodeBegin);
const auto DisasmEnd = reinterpret_cast<const vixl::aarch64::Instruction*>(CodeBegin + CodeOnlySize);
LogMan::Msg::IFmt("Disassemble Begin");
for (auto PCToDecode = DisasmBegin; PCToDecode < DisasmEnd; PCToDecode += 4) {
DisasmDecoder->Decode(PCToDecode);
@@ -903,14 +1112,12 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
#endif
if (DebugData) {
DebugData->HostCodeSize = CodeData.Size;
DebugData->Relocations = &Relocations;
}
DebugData->HostCodeSize = CodeData.Size;
DebugData->Relocations = &Relocations;
this->IR = nullptr;
return CodeData;
return std::move(CodeData);
}
void Arm64JITCore::ResetStack() {
+84 -20
View File
@@ -31,6 +31,10 @@ namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Context {
struct ExitFunctionLinkData;
}
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
@@ -38,9 +42,8 @@ public:
~Arm64JITCore() override;
[[nodiscard]]
CPUBackend::CompiledCode
CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData, bool CheckTF) override;
CPUBackend::CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) override;
void ClearCache() override;
@@ -58,17 +61,28 @@ private:
const bool HostSupportsAFP {};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
FEXCore::Context::ContextImpl* CTX {};
const FEXCore::IR::IRListView* IR {};
uint64_t Entry {};
CPUBackend::CompiledCode CodeData {};
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> CallReturnTargets;
struct PendingJumpThunk {
uint64_t CallerAddress;
uint64_t GuestRIP;
ARMEmitter::ForwardLabel Label;
};
fextl::vector<PendingJumpThunk> PendingJumpThunks;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempAllocator;
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
[[nodiscard]]
ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
ARMEmitter::Register GetReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::GPRFixedClass.Val) {
@@ -81,9 +95,17 @@ private:
}
[[nodiscard]]
ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
ARMEmitter::Register GetReg(IR::Ref Node) const {
return GetReg(IR::PhysicalRegister(Node));
}
[[nodiscard]]
ARMEmitter::Register GetReg(IR::OrderedNodeWrapper Wrap) const {
return GetReg(IR::PhysicalRegister(Wrap));
}
[[nodiscard]]
ARMEmitter::VRegister GetVReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::FPRFixedClass.Val) {
@@ -96,15 +118,18 @@ private:
}
[[nodiscard]]
FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
ARMEmitter::VRegister GetVReg(IR::Ref Node) const {
return GetVReg(IR::PhysicalRegister(Node));
}
[[nodiscard]]
IR::PhysicalRegister GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
ARMEmitter::VRegister GetVReg(IR::OrderedNodeWrapper Wrap) const {
return GetVReg(IR::PhysicalRegister(Wrap));
}
LOGMAN_THROW_A_FMT(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
return PhyReg;
[[nodiscard]]
FEXCore::IR::RegisterClassType GetRegClass(IR::Ref Node) const {
return FEXCore::IR::RegisterClassType {IR::PhysicalRegister(Node).Class};
}
[[nodiscard]]
@@ -114,7 +139,7 @@ private:
LOGMAN_THROW_A_FMT(Const == 0, "Only valid constant");
return ARMEmitter::Reg::zr;
} else {
return GetReg(Src.ID());
return GetReg(Src);
}
}
@@ -224,9 +249,34 @@ private:
}
[[nodiscard]]
bool IsFPR(IR::NodeID Node) const;
bool IsFPR(IR::RegisterClassType Class) const {
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
}
[[nodiscard]]
bool IsGPR(IR::NodeID Node) const;
bool IsGPR(IR::RegisterClassType Class) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
[[nodiscard]]
bool IsGPR(IR::Ref Node) {
return IsGPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsFPR(IR::Ref Node) {
return IsFPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsGPR(IR::OrderedNodeWrapper Wrap) {
return IsGPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
}
[[nodiscard]]
bool IsFPR(IR::OrderedNodeWrapper Wrap) {
return IsFPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
}
[[nodiscard]]
ARMEmitter::ExtendedMemOperand GenerateMemOperand(IR::OpSize AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
@@ -255,10 +305,20 @@ private:
uint32_t End;
};
void EmitLinkedBranch(uint64_t GuestRIP, bool Call) {
PendingJumpThunks.push_back({GetCursorAddress<uint64_t>(), GuestRIP, {}});
auto& Thunk = PendingJumpThunks.back();
Bind(&Thunk.Label);
if (Call) {
bl(&Thunk.Label);
} else {
b(&Thunk.Label);
}
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass {};
const IR::RegisterAllocationData* RAData {};
FEXCore::Core::DebugData* DebugData {};
void ResetStack();
@@ -319,7 +379,7 @@ private:
/** @} */
uint32_t SpillSlots {};
using OpType = void (Arm64JITCore::*)(const IR::IROp_Header* IROp, IR::NodeID Node);
using OpType = void (Arm64JITCore::*)(const IR::IROp_Header* IROp, IR::Ref Node);
using ScalarFMAOpCaller =
std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3)>;
@@ -341,12 +401,14 @@ private:
void EmitInterruptChecks(bool CheckTF);
void EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
OpType RT_StoreMemTSO;
#define DEF_OP(x) void Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header const* IROp, IR::Ref Node)
// Dynamic Dispatcher supporting operations
DEF_OP(ParanoidLoadMemTSO);
@@ -363,6 +425,8 @@ private:
#undef DEF_OP
};
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::Ref Node)
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
+156 -120
View File
@@ -11,11 +11,11 @@ $end_info$
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/MathUtils.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
const auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -53,8 +53,8 @@ DEF_OP(LoadContextPair) {
const auto Op = IROp->C<IR::IROp_LoadContextPair>();
if (Op->Class == FEXCore::IR::GPRClass) {
const auto Dst1 = GetReg(Op->OutValue1.ID());
const auto Dst2 = GetReg(Op->OutValue2.ID());
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
switch (IROp->Size) {
case IR::OpSize::i32Bit: ldp<ARMEmitter::IndexType::OFFSET>(Dst1.W(), Dst2.W(), STATE, Op->Offset); break;
@@ -62,8 +62,8 @@ DEF_OP(LoadContextPair) {
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemPair size: {}", IROp->Size); break;
}
} else {
const auto Dst1 = GetVReg(Op->OutValue1.ID());
const auto Dst2 = GetVReg(Op->OutValue2.ID());
const auto Dst1 = GetVReg(Op->OutValue1);
const auto Dst2 = GetVReg(Op->OutValue2);
switch (IROp->Size) {
case IR::OpSize::i32Bit: ldp<ARMEmitter::IndexType::OFFSET>(Dst1.S(), Dst2.S(), STATE, Op->Offset); break;
@@ -89,7 +89,7 @@ DEF_OP(StoreContext) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize); break;
}
} else {
const auto Src = GetVReg(Op->Value.ID());
const auto Src = GetVReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: strb(Src, STATE, Op->Offset); break;
@@ -120,8 +120,8 @@ DEF_OP(StoreContextPair) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize); break;
}
} else {
const auto Src1 = GetVReg(Op->Value1.ID());
const auto Src2 = GetVReg(Op->Value2.ID());
const auto Src1 = GetVReg(Op->Value1);
const auto Src2 = GetVReg(Op->Value2);
switch (OpSize) {
case IR::OpSize::i32Bit: stp<ARMEmitter::IndexType::OFFSET>(Src1.S(), Src2.S(), STATE, Op->Offset); break;
@@ -137,11 +137,8 @@ DEF_OP(LoadRegister) {
if (Op->Class == IR::GPRClass) {
LOGMAN_THROW_A_FMT(Op->Reg < StaticRegisters.size(), "out of range reg");
const auto reg = StaticRegisters[Op->Reg];
if (GetReg(Node).Idx() != reg.Idx()) {
mov(GetReg(Node).X(), reg.X());
}
mov(GetReg(Node).X(), StaticRegisters[Op->Reg].X());
} else if (Op->Class == IR::FPRClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "out of range reg");
@@ -150,12 +147,10 @@ DEF_OP(LoadRegister) {
const auto guest = StaticFPRegisters[Op->Reg];
const auto host = GetVReg(Node);
if (host.Idx() != guest.Idx()) {
if (HostSupportsAVX256) {
mov(ARMEmitter::SubRegSize::i64Bit, host.Z(), PRED_TMP_32B.Merging(), guest.Z());
} else {
mov(host.Q(), guest.Q());
}
if (HostSupportsAVX256) {
mov(ARMEmitter::SubRegSize::i64Bit, host.Z(), PRED_TMP_32B.Merging(), guest.Z());
} else {
mov(host.Q(), guest.Q());
}
} else {
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", Op->Class);
@@ -180,44 +175,32 @@ DEF_OP(LoadAF) {
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
auto Reg = IR::PhysicalRegister(Node);
if (Op->Class == IR::GPRClass) {
unsigned Reg = Op->Reg == Core::CPUState::PF_AS_GREG ? (StaticRegisters.size() - 2) :
Op->Reg == Core::CPUState::AF_AS_GREG ? (StaticRegisters.size() - 1) :
Op->Reg;
LOGMAN_THROW_A_FMT(Reg < StaticRegisters.size(), "out of range reg");
const auto reg = StaticRegisters[Reg];
const auto Src = GetReg(Op->Value.ID());
if (Src.Idx() != reg.Idx()) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
mov(ARMEmitter::Size::i64Bit, reg, Src);
}
} else if (Op->Class == IR::FPRClass) {
if (Reg.Class == IR::GPRFixedClass) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
mov(ARMEmitter::Size::i64Bit, GetReg(Reg), GetReg(Op->Value));
} else if (Reg.Class == IR::FPRFixedClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "reg out of range");
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
const auto guest = StaticFPRegisters[Op->Reg];
const auto host = GetVReg(Op->Value.ID());
const auto guest = GetVReg(Reg);
const auto host = GetVReg(Op->Value);
if (guest.Idx() != host.Idx()) {
if (HostSupportsAVX256) {
mov(ARMEmitter::SubRegSize::i64Bit, guest.Z(), PRED_TMP_32B.Merging(), host.Z());
} else {
mov(guest.Q(), host.Q());
}
if (HostSupportsAVX256) {
mov(ARMEmitter::SubRegSize::i64Bit, guest.Z(), PRED_TMP_32B.Merging(), host.Z());
} else {
mov(guest.Q(), host.Q());
}
} else {
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", Op->Class);
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", Reg.Class);
}
}
DEF_OP(StorePF) {
const auto Op = IROp->C<IR::IROp_StorePF>();
const auto reg = StaticRegisters[StaticRegisters.size() - 2];
const auto Src = GetReg(Op->Value.ID());
const auto Src = GetReg(Op->Value);
if (Src.Idx() != reg.Idx()) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
@@ -228,7 +211,7 @@ DEF_OP(StorePF) {
DEF_OP(StoreAF) {
const auto Op = IROp->C<IR::IROp_StoreAF>();
const auto reg = StaticRegisters[StaticRegisters.size() - 1];
const auto Src = GetReg(Op->Value.ID());
const auto Src = GetReg(Op->Value);
if (Src.Idx() != reg.Idx()) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
@@ -240,7 +223,7 @@ DEF_OP(LoadContextIndexed) {
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
const auto Index = GetReg(Op->Index.ID());
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
switch (Op->Stride) {
@@ -303,10 +286,10 @@ DEF_OP(StoreContextIndexed) {
const auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
const auto OpSize = IROp->Size;
const auto Index = GetReg(Op->Index.ID());
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
const auto Value = GetReg(Op->Value.ID());
const auto Value = GetReg(Op->Value);
switch (Op->Stride) {
case 1:
@@ -328,7 +311,7 @@ DEF_OP(StoreContextIndexed) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed stride: {}", Op->Stride); break;
}
} else {
const auto Value = GetVReg(Op->Value.ID());
const auto Value = GetVReg(Op->Value);
switch (Op->Stride) {
case 1:
@@ -371,7 +354,7 @@ DEF_OP(SpillRegister) {
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
if (Op->Class == FEXCore::IR::GPRClass) {
const auto Src = GetReg(Op->Value.ID());
const auto Src = GetReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: {
if (SlotOffset > LSByteMaxUnsignedOffset) {
@@ -412,7 +395,7 @@ DEF_OP(SpillRegister) {
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize); break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
const auto Src = GetVReg(Op->Value.ID());
const auto Src = GetVReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i32Bit: {
@@ -552,7 +535,7 @@ DEF_OP(LoadNZCV) {
DEF_OP(StoreNZCV) {
auto Op = IROp->C<IR::IROp_StoreNZCV>();
msr(ARMEmitter::SystemRegister::NZCV, GetReg(Op->Value.ID()));
msr(ARMEmitter::SystemRegister::NZCV, GetReg(Op->Value));
}
DEF_OP(LoadDF) {
@@ -575,7 +558,7 @@ ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(
if (IsInlineConstant(Offset, &Const)) {
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, Const);
} else {
auto RegOffset = GetReg(Offset.ID());
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
@@ -609,7 +592,7 @@ ARMEmitter::Register Arm64JITCore::ApplyMemOperand(IR::OpSize AccessSize, ARMEmi
LoadConstant(ARMEmitter::Size::i64Bit, Tmp, Const);
add(ARMEmitter::Size::i64Bit, Tmp, Base, Tmp, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
} else {
auto RegOffset = GetReg(Offset.ID());
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
@@ -676,7 +659,7 @@ ARMEmitter::SVEMemOperand Arm64JITCore::GenerateSVEMemOperand(IR::OpSize AccessS
// optional extension or shift as part of their behavior.
LOGMAN_THROW_A_FMT(OffsetType.Val == IR::MEM_OFFSET_SXTX.Val, "Currently only the default offset type (SXTX) is supported.");
const auto RegOffset = GetReg(Offset.ID());
const auto RegOffset = GetReg(Offset);
return ARMEmitter::SVEMemOperand(Base.X(), RegOffset.X());
}
@@ -684,7 +667,7 @@ DEF_OP(LoadMem) {
const auto Op = IROp->C<IR::IROp_LoadMem>();
const auto OpSize = IROp->Size;
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -719,11 +702,11 @@ DEF_OP(LoadMem) {
DEF_OP(LoadMemPair) {
const auto Op = IROp->C<IR::IROp_LoadMemPair>();
const auto Addr = GetReg(Op->Addr.ID());
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
const auto Dst1 = GetReg(Op->OutValue1.ID());
const auto Dst2 = GetReg(Op->OutValue2.ID());
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
switch (IROp->Size) {
case IR::OpSize::i32Bit: ldp<ARMEmitter::IndexType::OFFSET>(Dst1.W(), Dst2.W(), Addr, Op->Offset); break;
@@ -731,8 +714,8 @@ DEF_OP(LoadMemPair) {
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemPair size: {}", IROp->Size); break;
}
} else {
const auto Dst1 = GetVReg(Op->OutValue1.ID());
const auto Dst2 = GetVReg(Op->OutValue2.ID());
const auto Dst1 = GetVReg(Op->OutValue1);
const auto Dst2 = GetVReg(Op->OutValue2);
switch (IROp->Size) {
case IR::OpSize::i32Bit: ldp<ARMEmitter::IndexType::OFFSET>(Dst1.S(), Dst2.S(), Addr, Op->Offset); break;
@@ -747,7 +730,7 @@ DEF_OP(LoadMemTSO) {
const auto Op = IROp->C<IR::IROp_LoadMemTSO>();
const auto OpSize = IROp->Size;
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
@@ -844,8 +827,8 @@ DEF_OP(VLoadVectorMasked) {
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto Dst = GetVReg(Node);
const auto MaskReg = GetVReg(Op->Mask.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto MaskReg = GetVReg(Op->Mask);
const auto MemReg = GetReg(Op->Addr);
if (HostSupportsSVE128 || HostSupportsSVE256) {
const auto MemSrc = GenerateSVEMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
@@ -946,9 +929,9 @@ DEF_OP(VStoreVectorMasked) {
const auto CMPPredicate = ARMEmitter::PReg::p0;
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto RegData = GetVReg(Op->Data.ID());
const auto MaskReg = GetVReg(Op->Mask.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto RegData = GetVReg(Op->Data);
const auto MaskReg = GetVReg(Op->Mask);
const auto MemReg = GetReg(Op->Addr);
if (HostSupportsSVE128 || HostSupportsSVE256) {
const auto MemDst = GenerateSVEMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
@@ -1171,13 +1154,13 @@ DEF_OP(VLoadVectorGatherMasked) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto IncomingDst = GetVReg(Op->Incoming.ID());
const auto IncomingDst = GetVReg(Op->Incoming);
const auto MaskReg = GetVReg(Op->Mask.ID());
std::optional<ARMEmitter::Register> BaseAddr = !Op->AddrBase.IsInvalid() ? std::make_optional(GetReg(Op->AddrBase.ID())) : std::nullopt;
const auto VectorIndexLow = GetVReg(Op->VectorIndexLow.ID());
const auto MaskReg = GetVReg(Op->Mask);
std::optional<ARMEmitter::Register> BaseAddr = !Op->AddrBase.IsInvalid() ? std::make_optional(GetReg(Op->AddrBase)) : std::nullopt;
const auto VectorIndexLow = GetVReg(Op->VectorIndexLow);
std::optional<ARMEmitter::VRegister> VectorIndexHigh =
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh.ID())) : std::nullopt;
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh)) : std::nullopt;
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
const bool SupportsSVELoad = (HostSupportsSVE128 || HostSupportsSVE256) &&
@@ -1206,7 +1189,7 @@ DEF_OP(VLoadVectorGatherMasked) {
if (BaseAddr.has_value() || OffsetScale != 1) {
ARMEmitter::Register AddrReg = TMP1;
if (BaseAddr.has_value()) {
AddrReg = GetReg(Op->AddrBase.ID());
AddrReg = GetReg(Op->AddrBase);
} else {
///< OpcodeDispatcher didn't provide a Base address while SVE requires one.
LoadConstant(ARMEmitter::Size::i64Bit, AddrReg, 0);
@@ -1255,13 +1238,13 @@ DEF_OP(VLoadVectorGatherMaskedQPS) {
/// - Matches VGATHERQPS/VPGATHERQD behaviour!
const auto OffsetScale = Op->OffsetScale;
const auto Dst = GetVReg(Node);
const auto IncomingDst = GetVReg(Op->Incoming.ID());
const auto IncomingDst = GetVReg(Op->Incoming);
const auto MaskReg = GetVReg(Op->MaskReg.ID());
std::optional<ARMEmitter::Register> BaseAddr = !Op->AddrBase.IsInvalid() ? std::make_optional(GetReg(Op->AddrBase.ID())) : std::nullopt;
const auto VectorIndexLow = GetVReg(Op->VectorIndexLow.ID());
const auto MaskReg = GetVReg(Op->MaskReg);
std::optional<ARMEmitter::Register> BaseAddr = !Op->AddrBase.IsInvalid() ? std::make_optional(GetReg(Op->AddrBase)) : std::nullopt;
const auto VectorIndexLow = GetVReg(Op->VectorIndexLow);
std::optional<ARMEmitter::VRegister> VectorIndexHigh =
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh.ID())) : std::nullopt;
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh)) : std::nullopt;
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
if (HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4)) {
@@ -1330,8 +1313,8 @@ DEF_OP(VLoadVectorElement) {
const auto ElementSize = IROp->ElementSize;
const auto Dst = GetVReg(Node);
const auto DstSrc = GetVReg(Op->DstSrc.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto DstSrc = GetVReg(Op->DstSrc);
const auto MemReg = GetReg(Op->Addr);
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
@@ -1367,8 +1350,8 @@ DEF_OP(VStoreVectorElement) {
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
const auto ElementSize = IROp->ElementSize;
const auto Value = GetVReg(Op->Value.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto Value = GetVReg(Op->Value);
const auto MemReg = GetReg(Op->Addr);
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
@@ -1403,7 +1386,7 @@ DEF_OP(VBroadcastFromMem) {
const auto ElementSize = IROp->ElementSize;
const auto Dst = GetVReg(Node);
const auto MemReg = GetReg(Op->Address.ID());
const auto MemReg = GetReg(Op->Address);
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
@@ -1444,8 +1427,8 @@ DEF_OP(VBroadcastFromMem) {
DEF_OP(Push) {
const auto Op = IROp->C<IR::IROp_Push>();
const auto ValueSize = IR::OpSizeToSize(Op->ValueSize);
auto Src = GetReg(Op->Value.ID());
const auto AddrSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value);
const auto AddrSrc = GetReg(Op->Addr);
const auto Dst = GetReg(Node);
bool NeedsMoveAfterwards = false;
@@ -1526,11 +1509,34 @@ DEF_OP(Push) {
}
}
DEF_OP(PushTwo) {
const auto Op = IROp->C<IR::IROp_PushTwo>();
const auto ValueSize = IR::OpSizeToSize(Op->ValueSize);
auto Src1 = GetReg(Op->Value1);
auto Src2 = GetReg(Op->Value2);
const auto Dst = GetReg(Op->Addr);
switch (ValueSize) {
case 4: {
stp<ARMEmitter::IndexType::PRE>(Src1.W(), Src2.W(), Dst, -2 * ValueSize);
break;
}
case 8: {
stp<ARMEmitter::IndexType::PRE>(Src1.X(), Src2.X(), Dst, -2 * ValueSize);
break;
}
default: {
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, ValueSize);
break;
}
}
}
DEF_OP(Pop) {
const auto Op = IROp->C<IR::IROp_Pop>();
const auto Size = IR::OpSizeToSize(Op->Size);
const auto Addr = GetReg(Op->InoutAddr.ID());
const auto Dst = GetReg(Op->OutValue.ID());
const auto Addr = GetReg(Op->InoutAddr);
const auto Dst = GetReg(Op->OutValue);
LOGMAN_THROW_A_FMT(Dst != Addr, "Invalid");
@@ -1558,11 +1564,42 @@ DEF_OP(Pop) {
}
}
DEF_OP(PopTwo) {
const auto Op = IROp->C<IR::IROp_PopTwo>();
const auto Size = IR::OpSizeToSize(Op->Size);
const auto Addr = GetReg(Op->InoutAddr);
auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
// ldp x, x is invalid. Explicitly discard the first destination to encode.
if (Dst1 == Dst2) {
Dst1 = ARMEmitter::Reg::zr;
}
LOGMAN_THROW_A_FMT(Dst1 != Addr && Dst2 != Addr, "Invalid");
LOGMAN_THROW_A_FMT(Dst1 != Dst2, "Invalid");
switch (Size) {
case 4: {
ldp<ARMEmitter::IndexType::POST>(Dst1.W(), Dst2.W(), Addr, 2 * Size);
break;
}
case 8: {
ldp<ARMEmitter::IndexType::POST>(Dst1.X(), Dst2.X(), Addr, 2 * Size);
break;
}
default: {
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Op->Size);
break;
}
}
}
DEF_OP(StoreMem) {
const auto Op = IROp->C<IR::IROp_StoreMem>();
const auto OpSize = IROp->Size;
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
@@ -1575,7 +1612,7 @@ DEF_OP(StoreMem) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", OpSize); break;
}
} else {
const auto Src = GetVReg(Op->Value.ID());
const auto Src = GetVReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: {
@@ -1615,8 +1652,8 @@ DEF_OP(StoreMemX87SVEOptPredicate) {
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "StoreMemX87SVEOptPredicate needs SVE support");
const auto RegData = GetVReg(Op->Value.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto RegData = GetVReg(Op->Value);
const auto MemReg = GetReg(Op->Addr);
const auto MemDst = ARMEmitter::SVEMemOperand(MemReg.X(), 0);
switch (IROp->ElementSize) {
@@ -1644,7 +1681,7 @@ DEF_OP(LoadMemX87SVEOptPredicate) {
const auto Op = IROp->C<IR::IROp_LoadMemX87SVEOptPredicate>();
const auto Dst = GetVReg(Node);
const auto Predicate = PRED_X87_SVEOPT;
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "LoadMemX87SVEOptPredicate needs SVE support");
@@ -1674,7 +1711,7 @@ DEF_OP(LoadMemX87SVEOptPredicate) {
DEF_OP(StoreMemPair) {
const auto Op = IROp->C<IR::IROp_StoreMemPair>();
const auto OpSize = IROp->Size;
const auto Addr = GetReg(Op->Addr.ID());
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
const auto Src1 = GetZeroableReg(Op->Value1);
@@ -1685,8 +1722,8 @@ DEF_OP(StoreMemPair) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", OpSize); break;
}
} else {
const auto Src1 = GetVReg(Op->Value1.ID());
const auto Src2 = GetVReg(Op->Value2.ID());
const auto Src1 = GetVReg(Op->Value1);
const auto Src2 = GetVReg(Op->Value2);
switch (OpSize) {
case IR::OpSize::i32Bit: stp<ARMEmitter::IndexType::OFFSET>(Src1.S(), Src2.S(), Addr, Op->Offset); break;
@@ -1701,7 +1738,7 @@ DEF_OP(StoreMemTSO) {
const auto Op = IROp->C<IR::IROp_StoreMemTSO>();
const auto OpSize = IROp->Size;
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
@@ -1751,7 +1788,7 @@ DEF_OP(StoreMemTSO) {
// Half-Barrier.
dmb(ARMEmitter::BarrierScope::ISH);
}
const auto Src = GetVReg(Op->Value.ID());
const auto Src = GetVReg(Op->Value);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
switch (OpSize) {
case IR::OpSize::i8Bit: strb(Src, MemSrc); break;
@@ -1782,16 +1819,16 @@ DEF_OP(MemSet) {
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
const auto Size = IR::OpSizeToSize(Op->Size);
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
const auto Value = GetZeroableReg(Op->Value);
const auto Length = GetReg(Op->Length.ID());
const auto Length = GetReg(Op->Length);
const auto Dst = GetReg(Node);
uint64_t DirectionConstant;
bool DirectionIsInline = IsInlineConstant(Op->Direction, &DirectionConstant);
ARMEmitter::Register DirectionReg = ARMEmitter::Reg::r0;
if (!DirectionIsInline) {
DirectionReg = GetReg(Op->Direction.ID());
DirectionReg = GetReg(Op->Direction);
}
// If Direction > 0 then:
@@ -1808,7 +1845,7 @@ DEF_OP(MemSet) {
if (Op->Prefix.IsInvalid()) {
mov(TMP2, MemReg.X());
} else {
const auto Prefix = GetReg(Op->Prefix.ID());
const auto Prefix = GetReg(Op->Prefix);
add(TMP2, Prefix.X(), MemReg.X());
}
@@ -1971,19 +2008,19 @@ DEF_OP(MemCpy) {
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
const auto Size = IR::OpSizeToSize(Op->Size);
const auto MemRegDest = GetReg(Op->Dest.ID());
const auto MemRegSrc = GetReg(Op->Src.ID());
const auto MemRegDest = GetReg(Op->Dest);
const auto MemRegSrc = GetReg(Op->Src);
const auto Length = GetReg(Op->Length.ID());
const auto Length = GetReg(Op->Length);
uint64_t DirectionConstant;
bool DirectionIsInline = IsInlineConstant(Op->Direction, &DirectionConstant);
ARMEmitter::Register DirectionReg = ARMEmitter::Reg::r0;
if (!DirectionIsInline) {
DirectionReg = GetReg(Op->Direction.ID());
DirectionReg = GetReg(Op->Direction);
}
auto Dst0 = GetReg(Op->OutDstAddress.ID());
auto Dst1 = GetReg(Op->OutSrcAddress.ID());
auto Dst0 = GetReg(Op->OutDstAddress);
auto Dst1 = GetReg(Op->OutSrcAddress);
// If Direction > 0 then:
// MemRegDest is incremented (by size)
// MemRegSrc is incremented (by size)
@@ -2241,7 +2278,7 @@ DEF_OP(ParanoidLoadMemTSO) {
const auto Op = IROp->C<IR::IROp_LoadMemTSO>();
const auto OpSize = IROp->Size;
auto MemReg = GetReg(Op->Addr.ID());
auto MemReg = GetReg(Op->Addr);
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
const auto Dst = GetReg(Node);
@@ -2329,7 +2366,7 @@ DEF_OP(ParanoidStoreMemTSO) {
const auto Op = IROp->C<IR::IROp_StoreMemTSO>();
const auto OpSize = IROp->Size;
auto MemReg = GetReg(Op->Addr.ID());
auto MemReg = GetReg(Op->Addr);
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
const auto Src = GetZeroableReg(Op->Value);
@@ -2362,7 +2399,7 @@ DEF_OP(ParanoidStoreMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", OpSize); break;
}
} else {
const auto Src = GetVReg(Op->Value.ID());
const auto Src = GetVReg(Op->Value);
MemReg = ApplyMemOperand(OpSize, MemReg, TMP4, Op->Offset, Op->OffsetType, Op->OffsetScale);
@@ -2416,7 +2453,7 @@ DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
auto MemReg = GetReg(Op->Addr.ID());
auto MemReg = GetReg(Op->Addr);
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
@@ -2445,7 +2482,7 @@ DEF_OP(CacheLineClean) {
auto Op = IROp->C<IR::IROp_CacheLineClean>();
auto MemReg = GetReg(Op->Addr.ID());
auto MemReg = GetReg(Op->Addr);
// Clean dcache only
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
@@ -2463,7 +2500,7 @@ DEF_OP(CacheLineClean) {
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
auto MemReg = GetReg(Op->Addr.ID());
auto MemReg = GetReg(Op->Addr);
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use this instruction directly
@@ -2483,7 +2520,7 @@ DEF_OP(CacheLineZero) {
DEF_OP(Prefetch) {
auto Op = IROp->C<IR::IROp_Prefetch>();
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
// Access size is only ever handled as 8-byte. Even though it is accesssed as a cacheline.
const auto MemSrc = GenerateMemOperand(IR::OpSize::i64Bit, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
@@ -2527,8 +2564,8 @@ DEF_OP(VStoreNonTemporal) {
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
const auto Value = GetVReg(Op->Value.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto Value = GetVReg(Op->Value);
const auto MemReg = GetReg(Op->Addr);
const auto Offset = Op->Offset;
if (Is256Bit) {
@@ -2552,10 +2589,10 @@ DEF_OP(VStoreNonTemporalPair) {
[[maybe_unused]] const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Is128Bit, "This IR operation only operates at 128-bit wide");
const auto ValueLow = GetVReg(Op->ValueLow.ID());
const auto ValueHigh = GetVReg(Op->ValueHigh.ID());
const auto ValueLow = GetVReg(Op->ValueLow);
const auto ValueHigh = GetVReg(Op->ValueHigh);
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
const auto Offset = Op->Offset;
stnp(ValueLow.Q(), ValueHigh.Q(), MemReg, Offset);
@@ -2570,7 +2607,7 @@ DEF_OP(VLoadNonTemporal) {
const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
const auto Dst = GetVReg(Node);
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemReg = GetReg(Op->Addr);
const auto Offset = Op->Offset;
if (Is256Bit) {
@@ -2587,5 +2624,4 @@ DEF_OP(VLoadNonTemporal) {
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
+27 -13
View File
@@ -16,11 +16,26 @@ $end_info$
#include <FEXCore/Core/SignalDelegator.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(AllocateGPR) {}
DEF_OP(AllocateGPRAfter) {}
DEF_OP(AllocateFPR) {}
DEF_OP(WFET) {
auto Op = IROp->C<IR::IROp_WFET>();
const auto Lower = GetReg(Op->Lower);
const auto Upper = GetReg(Op->Upper);
// Combine registers.
mov(ARMEmitter::Size::i64Bit, TMP1, Lower);
bfi(ARMEmitter::Size::i64Bit, TMP1, Upper, 32, 32);
if (CTX->Config.TSCScale) {
// Scale back to ARM64 TSC scale if necessary
lsr(ARMEmitter::Size::i64Bit, TMP1, TMP1, CTX->Config.TSCScale);
}
// Clear the exclusive monitor so it can't spuriously wake up with that event.
clrex();
// Execute wfet to wait until the TSC.
wfet(TMP1);
}
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
@@ -102,8 +117,8 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg(Op->RoundMode.ID());
auto MXCSR = GetReg(Op->MXCSR.ID());
auto Src = GetReg(Op->RoundMode);
auto MXCSR = GetReg(Op->MXCSR);
// As above, setup the rounding flags in [31:30]
rbit(ARMEmitter::Size::i32Bit, TMP2, Src);
@@ -161,7 +176,7 @@ DEF_OP(PushRoundingMode) {
DEF_OP(PopRoundingMode) {
auto Op = IROp->C<IR::IROp_PopRoundingMode>();
msr(ARMEmitter::SystemRegister::FPCR, GetReg(Op->FPCR.ID()));
msr(ARMEmitter::SystemRegister::FPCR, GetReg(Op->FPCR));
}
DEF_OP(Print) {
@@ -170,17 +185,17 @@ DEF_OP(Print) {
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
if (IsGPR(Op->Value.ID())) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value.ID()));
if (IsGPR(Op->Value)) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue));
} else {
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetVReg(Op->Value.ID()), false);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetVReg(Op->Value.ID()), true);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetVReg(Op->Value), false);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetVReg(Op->Value), true);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue));
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
if (IsGPR(Op->Value.ID())) {
if (IsGPR(Op->Value)) {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r3);
} else {
GenerateIndirectRuntimeCall<void, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
@@ -271,5 +286,4 @@ DEF_OP(Yield) {
yield();
}
#undef DEF_OP
} // namespace FEXCore::CPU
+3 -11
View File
@@ -8,26 +8,19 @@ $end_info$
#include "Interface/Core/JIT/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(Copy) {
auto Op = IROp->C<IR::IROp_Copy>();
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Source.ID()));
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Source));
}
DEF_OP(RMWHandle) {
auto Op = IROp->C<IR::IROp_RMWHandle>();
auto Dest = GetReg(Node);
auto Src = GetReg(Op->Value.ID());
if (Dest != Src) {
mov(ARMEmitter::Size::i64Bit, Dest, Src);
}
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(IROp->Args[0]));
}
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A.ID()), B = GetReg(Op->B.ID());
auto A = GetReg(Op->A), B = GetReg(Op->B);
LOGMAN_THROW_A_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
@@ -39,5 +32,4 @@ DEF_OP(Swap2) {
// Implemented above
}
#undef DEF_OP
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
+23 -8
View File
@@ -14,14 +14,17 @@ $end_info$
#include "Interface/Core/LookupCache.h"
namespace FEXCore {
LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
: BlockLinks_mbr {fextl::pmr::get_default_resource()}
, ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
GuestToHostMap::GuestToHostMap()
: BlockLinks_mbr {fextl::pmr::get_default_resource()} {
BlockLinks_pma = fextl::make_unique<std::pmr::polymorphic_allocator<std::byte>>(&BlockLinks_mbr);
// Setup our PMR map.
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
}
LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
: ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
// Block cache ends up looking like this
// PageMemoryMap[VirtualMemoryRegion >> 12]
@@ -62,18 +65,30 @@ LookupCache::~LookupCache() {
}
void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
auto lk = Shared->AcquireLock();
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE, false);
AllocateOffset = 0;
}
void LookupCache::ClearCache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
void LookupCache::ClearThreadLocalCaches() {
auto lk = Shared->AcquireLock();
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
}
void LookupCache::ClearCache() {
auto lk = Shared->AcquireLock();
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
Shared->ClearCache(lk);
}
void GuestToHostMap::ClearCache(const LockToken&) {
// Allocate a new pointer from the BlockLinks pma again.
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
// All code is gone, clear the block list
+128 -68
View File
@@ -16,6 +16,100 @@
namespace FEXCore {
struct GuestToHostMap {
std::recursive_mutex WriteLock;
struct LockToken {
std::lock_guard<std::recursive_mutex> Lock;
};
[[nodiscard]]
LockToken AcquireLock() {
return LockToken {std::lock_guard {WriteLock}};
}
struct BlockLinkTag {
uint64_t GuestDestination;
FEXCore::Context::ExitFunctionLinkData* HostLink;
bool operator<(const BlockLinkTag& other) const {
if (GuestDestination < other.GuestDestination) {
return true;
} else if (GuestDestination == other.GuestDestination) {
return HostLink < other.HostLink;
} else {
return false;
}
}
};
// Use a monotonic buffer resource to allocate both the std::pmr::map and its members.
// This allows us to quickly clear the block link map by clearing the monotonic allocator.
// If we had allocated the block link map without the MBR, then clearing the map would require slowly
// walking each block member and destructing objects.
//
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, FEXCore::Context::BlockDelinkerFunc>;
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
}
return HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
}
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length, const LockToken&) {
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto& CodePage = CodePages[CurrentPage];
rv |= CodePage.empty();
CodePage.insert(CodePage.end(), Addresses.begin(), Addresses.end());
}
return rv;
}
void ClearCache(const LockToken&);
};
class LookupCache {
public:
struct LookupCacheEntry {
@@ -26,6 +120,13 @@ public:
LookupCache(FEXCore::Context::ContextImpl* CTX);
~LookupCache();
// Swaps out the underlying GuestToHostMap and clears all associated caches.
// This interface requires the previous CodeBuffer to be provided despite not using it. This ensures the shared write lock is still valid.
void ChangeGuestToHostMapping([[maybe_unused]] CPU::CodeBuffer& Prev, GuestToHostMap& NewMap) {
ClearThreadLocalCaches();
Shared = &NewMap;
}
uintptr_t FindBlock(uint64_t Address) {
// Try L1, no lock needed
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
@@ -34,7 +135,7 @@ public:
}
// L2 and L3 need to be locked
std::lock_guard<std::recursive_mutex> lk(WriteLock);
auto lk = Shared->AcquireLock();
// Try L2
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
@@ -56,41 +157,30 @@ public:
}
// Try L3
auto HostCode = BlockList.find(Address);
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value());
return HostCode.value();
}
// Failed to find
return 0;
}
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap* Shared = nullptr;
// Appends Block {Address} to CodePages [Start, Start + Length)
// Appends a list of Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto& CodePage = CodePages[CurrentPage];
rv |= CodePage.size() == 0;
CodePage.push_back(Address);
}
return rv;
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireLock();
return Shared->AddBlockExecutableRange(Addresses, Start, Length, lk);
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
auto lk = Shared->AcquireLock();
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_A_FMT(Inserted, "Duplicate block mapping added");
Shared->AddBlockMapping(Address, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
@@ -99,24 +189,19 @@ public:
L1Entry.HostCode = (uintptr_t)HostCode;
}
void Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
auto lk = Shared->AcquireLock();
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
}
// Remove from BlockList
BlockList.erase(Address);
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -131,23 +216,24 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return;
return ErasedAny;
}
// Page exists, just set the offset to zero
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
return true;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink, const FEXCore::Context::BlockDelinkerFunc& delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
auto lk = Shared->AcquireLock();
Shared->AddBlockLink(GuestDestination, HostLink, delinker, lk);
}
void ClearCache();
void ClearL2Cache();
void ClearThreadLocalCaches();
uintptr_t GetL1Pointer() const {
return L1Pointer;
@@ -169,7 +255,9 @@ public:
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
// This approach has not been fully vetted yet.
// Also note that L1 lookups might be inlined in the JIT Dispatcher and/or block ends.
std::recursive_mutex WriteLock;
auto AcquireLock() {
return Shared->AcquireLock();
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
@@ -226,34 +314,6 @@ private:
uintptr_t PageMemory;
uintptr_t L1Pointer;
struct BlockLinkTag {
uint64_t GuestDestination;
FEXCore::Context::ExitFunctionLinkData* HostLink;
bool operator<(const BlockLinkTag& other) const {
if (GuestDestination < other.GuestDestination) {
return true;
} else if (GuestDestination == other.GuestDestination) {
return HostLink < other.HostLink;
} else {
return false;
}
}
};
// Use a monotonic buffer resource to allocate both the std::pmr::map and its members.
// This allows us to quickly clear the block link map by clearing the monotonic allocator.
// If we had allocated the block link map without the MBR, then clearing the map would require slowly
// walking each block member and destructing objects.
//
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, FEXCore::Context::BlockDelinkerFunc>;
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
size_t TotalCacheSize;
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
File diff suppressed because it is too large. Load diff
+214 -76
View File
@@ -7,6 +7,7 @@
#include "Interface/Context/Context.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -79,6 +80,12 @@ struct LoadSourceOptions {
bool AllowUpperGarbage = false;
};
struct DispatchTableEntry {
uint16_t Op;
uint8_t Count;
X86Tables::OpDispatchPtr Ptr;
};
class OpDispatchBuilder final : public IREmitter {
friend class FEXCore::IR::Pass;
friend class FEXCore::IR::PassManager;
@@ -159,9 +166,13 @@ public:
auto InlineConst = _InlineConstant(Bit);
return _CondJump(Src, InlineConst, InvalidNode, InvalidNode, {Set ? COND_TSTNZ : COND_TSTZ}, OpSize::iInvalid, false);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP) {
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint = BranchHint::None) {
FlushRegisterCache();
return _ExitFunction(NewRIP);
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, InvalidNode, InvalidNode);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
FlushRegisterCache();
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, CallReturnAddress, CallReturnBlock);
}
IRPair<IROp_Break> Break(BreakDefinition Reason) {
FlushRegisterCache();
@@ -187,7 +198,7 @@ public:
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end()) {
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
// If we don't have a jump target to a new block then we have to leave
// Set the RIP to the next instruction and leave
auto RelocatedNextRIP = _EntrypointOffset(GPRSize, NextRIP - Entry);
@@ -231,7 +242,7 @@ public:
template<typename F>
void ForeachDirection(F&& Routine) {
// Otherwise, prepare to branch.
auto Zero = _Constant(0);
auto Zero = Constant(0);
// If the shift is zero, do not touch the flags.
auto ForwardBlock = CreateNewCodeBlockAfter(GetCurrentBlock());
@@ -289,7 +300,7 @@ public:
return ShouldDump;
}
void BeginFunction(uint64_t RIP, const fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks>* Blocks, uint32_t NumInstructions);
void BeginFunction(uint64_t RIP, const fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks>* Blocks, uint32_t NumInstructions, bool Is64BitMode);
void Finalize();
// Dispatch builder functions
@@ -701,7 +712,7 @@ public:
Ref ReconstructX87StateFromFSW_Helper(Ref FSW);
void FLD(OpcodeArgs, IR::OpSize Width);
void FLDFromStack(OpcodeArgs);
void FLD_Const(OpcodeArgs, NamedVectorConstant Constant);
void FLD_Const(OpcodeArgs, NamedVectorConstant K);
void FBLD(OpcodeArgs);
void FBSTP(OpcodeArgs);
@@ -823,11 +834,13 @@ public:
void PHADDS(OpcodeArgs);
void PHSUBS(OpcodeArgs);
void CLWB(OpcodeArgs);
void CLWBOrTPause(OpcodeArgs);
void CLFLUSHOPT(OpcodeArgs);
void LoadFenceOrXRSTOR(OpcodeArgs);
void MemFenceOrXSAVEOPT(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void UMonitorOrCLRSSBSY(OpcodeArgs);
void UMWaitOp(OpcodeArgs);
void CLZeroOp(OpcodeArgs);
void RDTSCPOp(OpcodeArgs);
void RDPIDOp(OpcodeArgs);
@@ -1157,6 +1170,7 @@ public:
// End of AVX 256-bit implementation
void InvalidOp(OpcodeArgs);
void NoExecOp(OpcodeArgs);
void SetPackedRFLAG(bool Lower8, Ref Src);
Ref GetPackedRFLAG(uint32_t FlagsMask = ~0U);
@@ -1183,7 +1197,7 @@ public:
CalculateDeferredFlags();
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
const auto VectorSize = GetGuestVectorLength();
// Write backwards. This is a heuristic to improve coalescing, since we
@@ -1208,13 +1222,15 @@ public:
Ref Value = RegCache.Value[Index];
if (Index >= GPR0Index && Index <= GPR15Index) {
_StoreRegister(Value, Index - GPR0Index, GPRClass, GPRSize);
Ref R = _StoreRegister(Value, GPRSize);
R->Reg = PhysicalRegister(GPRFixedClass, Index - GPR0Index).Raw;
} else if (Index == PFIndex) {
_StorePF(Value, GPRSize);
} else if (Index == AFIndex) {
_StoreAF(Value, GPRSize);
} else if (Index >= FPR0Index && Index <= FPR15Index) {
_StoreRegister(Value, Index - FPR0Index, FPRClass, VectorSize);
Ref R = _StoreRegister(Value, VectorSize);
R->Reg = PhysicalRegister(FPRFixedClass, Index - FPR0Index).Raw;
} else if (Index == DFIndex) {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(Core::CPUState, flags[X86State::RFLAG_DF_RAW_LOC]));
} else {
@@ -1241,7 +1257,7 @@ public:
_StoreContext(Size, Class, Value, Offset);
// If Partial and MMX register, then we need to store all 1s in bits 64-80
if (Partial && Index >= MM0Index && Index <= MM7Index) {
_StoreContext(OpSize::i16Bit, IR::GPRClass, _Constant(0xFFFF), Offset + 8);
_StoreContext(OpSize::i16Bit, IR::GPRClass, Constant(0xFFFF), Offset + 8);
}
}
}
@@ -1254,6 +1270,10 @@ public:
RegCache.Partial &= ~Mask;
}
IR::OpSize GetGPROpSize() const {
return Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
protected:
void RecordX87Use() override {
CurrentHeader->HasX87 = true;
@@ -1307,6 +1327,7 @@ private:
struct JumpTargetInfo {
Ref BlockEntry;
bool HaveEmitted;
bool IsEntryPoint;
};
FEXCore::Context::ContextImpl* CTX {};
@@ -1492,7 +1513,7 @@ private:
#undef OpcodeArgs
Ref AppendSegmentOffset(Ref Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
Ref GetSegment(uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
Ref GetSegment(uint32_t Flags, uint32_t DefaultPrefix = FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX, bool Override = false);
void UpdatePrefixFromSegment(Ref Segment, uint32_t SegmentReg);
@@ -1502,12 +1523,14 @@ private:
Ref GetRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0);
bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
[[nodiscard]]
static bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
// Literals are immediates as sources but memory addresses as destinations.
return !(Load && Operand.IsLiteral()) && !Operand.IsGPR();
}
bool IsNonTSOReg(MemoryAccessType Access, uint8_t Reg) {
[[nodiscard]]
static bool IsNonTSOReg(MemoryAccessType Access, uint8_t Reg) {
return Access == MemoryAccessType::DEFAULT && Reg == X86State::REG_RSP;
}
@@ -1614,7 +1637,7 @@ private:
}
void ZeroNZCV() {
CachedNZCV = _Constant(0);
CachedNZCV = Constant(0);
NZCVDirty = true;
}
@@ -1627,9 +1650,9 @@ private:
// This is currently worse for 8/16-bit, but that should be optimized. TODO
if (SrcSize >= OpSize::i32Bit) {
if (SetPF) {
CalculatePF(_SubWithFlags(SrcSize, Res, _Constant(0)));
CalculatePF(_SubWithFlags(SrcSize, Res, Constant(0)));
} else {
_SubNZCV(SrcSize, Res, _Constant(0));
_SubNZCV(SrcSize, Res, Constant(0));
}
CFInverted = true;
@@ -1700,7 +1723,7 @@ private:
} else {
// Invert as a GPR
unsigned Bit = IndexNZCV(FEXCore::X86State::RFLAG_CF_RAW_LOC);
SetNZCV(_Xor(OpSize::i32Bit, GetNZCV(), _Constant(1u << Bit)));
SetNZCV(_Xor(OpSize::i32Bit, GetNZCV(), Constant(1u << Bit)));
CalculateDeferredFlags();
}
@@ -1738,7 +1761,7 @@ private:
}
HandleNZCVWrite();
_SubNZCV(OpSize::i32Bit, _Constant(0), Value);
_SubNZCV(OpSize::i32Bit, Constant(0), Value);
CFInverted = true;
}
@@ -1763,25 +1786,25 @@ private:
StoreRegister(Core::CPUState::AF_AS_GREG, false, Value);
} else if (BitOffset == FEXCore::X86State::RFLAG_DF_RAW_LOC) {
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, _Constant(1), Value, ShiftType::LSL, 1));
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF = _Select(FEXCore::IR::COND_EQ, Value, _Constant(0), _And(OpSize::i32Bit, PackedTF, _Constant(~1)), _Constant(1));
auto NewPackedTF = _Select(FEXCore::IR::COND_EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContext(OpSize::i8Bit, GPRClass, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
} else {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
}
}
void SetAF(unsigned Constant) {
void SetAF(unsigned K) {
// AF is stored in bit 4 of the AF flag byte, with garbage in the other
// bits. This allows us to defer the extract in the usual case. When it is
// read, bit 4 is extracted. In order to write a constant value of AF, that
// means we need to left-shift here to compensate.
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Constant(Constant << 4));
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(K << 4));
}
void ZeroPF_AF();
@@ -1797,7 +1820,8 @@ private:
InvalidateReg(Core::CPUState::AF_AS_GREG);
}
CondClassType CondForNZCVBit(unsigned BitOffset, bool Invert) {
[[nodiscard]]
static CondClassType CondForNZCVBit(unsigned BitOffset, bool Invert) {
switch (BitOffset) {
case X86State::RFLAG_SF_RAW_LOC: return {Invert ? COND_PL : COND_MI};
case X86State::RFLAG_ZF_RAW_LOC: return {Invert ? COND_NEQ : COND_EQ};
@@ -1823,7 +1847,8 @@ private:
static const int AVXHigh0Index = 48;
static const int AVXHigh15Index = 63;
uint32_t CacheIndexToContextOffset(int Index) {
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
@@ -1832,7 +1857,8 @@ private:
}
}
RegisterClassType CacheIndexClass(int Index) {
[[nodiscard]]
static RegisterClassType CacheIndexClass(int Index) {
if ((Index >= MM0Index && Index <= MM7Index) || Index >= FPR0Index) {
return FPRClass;
} else {
@@ -1840,7 +1866,8 @@ private:
}
}
IR::OpSize CacheIndexToOpSize(int Index) {
[[nodiscard]]
static IR::OpSize CacheIndexToOpSize(int Index) {
// MMX registers are rounded up to 128-bit since they are shared with 80-bit
// x87 registers, even though MMX is logically only 64-bit.
if (Index >= AVXHigh0Index || ((Index >= MM0Index && Index <= MM7Index))) {
@@ -1948,7 +1975,7 @@ private:
}
Ref LoadGPR(uint8_t Reg) {
return LoadRegCache(Reg, GPR0Index + Reg, GPRClass, CTX->GetGPROpSize());
return LoadRegCache(Reg, GPR0Index + Reg, GPRClass, GetGPROpSize());
}
Ref LoadContext(IR::OpSize Size, uint8_t Index) {
@@ -2002,14 +2029,14 @@ private:
auto Value = _Bfe(OpSize::i32Bit, 1, IndexNZCV(BitOffset), GetNZCV());
if (Invert) {
return _Xor(OpSize::i32Bit, Value, _Constant(1));
return _Xor(OpSize::i32Bit, Value, Constant(1));
} else {
return Value;
}
} else {
// Because we explicitly inverted for CF above, we use the unsafe
// _NZCVSelect rather than the safe CF-aware version.
return _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(BitOffset, Invert), _Constant(1), _Constant(0));
return _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(BitOffset, Invert), Constant(1), Constant(0));
}
} else if (BitOffset == FEXCore::X86State::RFLAG_PF_RAW_LOC) {
return LoadGPR(Core::CPUState::PF_AS_GREG);
@@ -2017,7 +2044,7 @@ private:
return LoadGPR(Core::CPUState::AF_AS_GREG);
} else if (BitOffset == FEXCore::X86State::RFLAG_DF_RAW_LOC) {
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), _Constant(63));
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(Core::CPUState, flags[BitOffset]));
}
@@ -2025,14 +2052,7 @@ private:
// Returns (DF ? -Size : Size)
Ref LoadDir(const unsigned Size) {
auto Dir = LoadDF();
auto Shift = FEXCore::ilog2(Size);
if (Shift) {
return _Lshl(CTX->GetGPROpSize(), Dir, _Constant(Shift));
} else {
return Dir;
}
return ARef(LoadDF()).Lshl(FEXCore::ilog2(Size)).Ref();
}
// Returns DF ? (X - Size) : (X + Size)
@@ -2096,7 +2116,7 @@ private:
// Zero AF. Note that the comparison sets the raw PF to 0/1 above, so
// PF[4] is 0 so the XOR with PF will have no effect, so setting the AF
// byte to zero will indeed zero AF as intended.
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(0));
}
// Convert NZCV from the Arm representation to an eXternal representation
@@ -2113,7 +2133,7 @@ private:
void ConvertNZCVToX87() {
LOGMAN_THROW_A_FMT(NZCVDirty && CachedNZCV, "NZCV must be saved");
Ref V = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_OF_RAW_LOC, false), _Constant(1), _Constant(0));
Ref V = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_OF_RAW_LOC, false), Constant(1), Constant(0));
if (CTX->HostFeatures.SupportsFlagM2) {
// Convert to x86 flags, saves us from or'ing after.
@@ -2121,8 +2141,8 @@ private:
}
// CF is inverted after FCMP
Ref C = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_CF_RAW_LOC, true), _Constant(1), _Constant(0));
Ref Z = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_ZF_RAW_LOC, false), _Constant(1), _Constant(0));
Ref C = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_CF_RAW_LOC, true), Constant(1), Constant(0));
Ref Z = _NZCVSelect(OpSize::i32Bit, CondForNZCVBit(FEXCore::X86State::RFLAG_ZF_RAW_LOC, false), Constant(1), Constant(0));
if (!CTX->HostFeatures.SupportsFlagM2) {
C = _Or(OpSize::i32Bit, C, V);
@@ -2130,7 +2150,7 @@ private:
}
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(C);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(V);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(Z);
}
@@ -2180,9 +2200,9 @@ private:
return CachedNamedVectorConstants[NamedConstant][log2_size_bytes];
}
auto Constant = _LoadNamedVectorConstant(Size, NamedConstant);
CachedNamedVectorConstants[NamedConstant][log2_size_bytes] = Constant;
return Constant;
auto K = _LoadNamedVectorConstant(Size, NamedConstant);
CachedNamedVectorConstants[NamedConstant][log2_size_bytes] = K;
return K;
}
Ref LoadAndCacheIndexedNamedVectorConstant(IR::OpSize Size, FEXCore::IR::IndexNamedVectorConstant NamedIndexedConstant, uint32_t Index) {
IndexNamedVectorMapKey Key {
@@ -2196,9 +2216,9 @@ private:
return it->second;
}
auto Constant = _LoadNamedVectorIndexedConstant(Size, NamedIndexedConstant, Index);
CachedIndexedNamedVectorConstants.insert_or_assign(Key, Constant);
return Constant;
auto K = _LoadNamedVectorIndexedConstant(Size, NamedIndexedConstant, Index);
CachedIndexedNamedVectorConstants.insert_or_assign(Key, K);
return K;
}
Ref LoadUncachedZeroVector(IR::OpSize Size) {
@@ -2216,7 +2236,7 @@ private:
CachedIndexedNamedVectorConstants.clear();
}
std::pair<bool, CondClassType> DecodeNZCVCondition(uint8_t OP);
std::optional<CondClassType> DecodeNZCVCondition(uint8_t OP);
Ref SelectBit(Ref Cmp, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
Ref SelectCC(uint8_t OP, IR::OpSize ResultSize, Ref TrueValue, Ref FalseValue);
@@ -2253,7 +2273,7 @@ private:
}
// Otherwise, prepare to branch.
auto Zero = _Constant(0);
auto Zero = Constant(0);
// If the shift is zero, do not touch the flags.
auto SetBlock = CreateNewCodeBlockAfter(GetCurrentBlock());
@@ -2310,7 +2330,7 @@ private:
Ref CalculateFlags_ADD(IR::OpSize SrcSize, Ref Src1, Ref Src2, bool UpdateCF = true);
void CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High);
void CalculateFlags_UMUL(Ref High);
void CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2);
void CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res);
void CalculateFlags_ShiftLeft(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2);
void CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRight(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2);
@@ -2321,20 +2341,7 @@ private:
void CalculateFlags_ZCNT(IR::OpSize SrcSize, Ref Result);
/** @} */
Ref AndConst(FEXCore::IR::OpSize Size, Ref Node, uint64_t Const) {
uint64_t NodeConst;
if (IsValueConstant(WrapNode(Node), &NodeConst)) {
return _Constant(NodeConst & Const);
} else {
return _And(Size, Node, _Constant(Const));
}
}
/** @} */
Ref GetX87Top();
Ref GetX87Tag(Ref Value, Ref AbridgedFTW);
void SetX87FTW(Ref FTW);
Ref GetX87FTW_Helper();
void SetX87Top(Ref Value);
@@ -2342,8 +2349,8 @@ private:
void ChgStateX87_MMX() override {
LOGMAN_THROW_A_FMT(MMXState == MMXState_X87, "Expected state to be x87");
_StackForceSlow();
SetX87Top(_Constant(0)); // top reset to zero
StoreContext(AbridgedFTWIndex, _Constant(0xFFFFUL)); // all valid
SetX87Top(Constant(0)); // top reset to zero
StoreContext(AbridgedFTWIndex, Constant(0xFFFFUL)); // all valid
MMXState = MMXState_MMX;
}
@@ -2368,10 +2375,12 @@ private:
bool BlockSetRIP {false};
bool Multiblock {};
bool Is64BitMode {};
uint64_t Entry {};
IROp_IRHeader* CurrentHeader {};
bool IsTSOEnabled(FEXCore::IR::RegisterClassType Class) {
[[nodiscard]]
bool IsTSOEnabled(FEXCore::IR::RegisterClassType Class) const {
if (ForceTSO == ForceTSOMode::ForceEnabled) {
return true;
} else if (ForceTSO == ForceTSOMode::ForceDisabled) {
@@ -2401,7 +2410,7 @@ private:
Ref _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
if (AtomicTSO) {
return _LoadMemTSO(Class, Size, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
@@ -2421,7 +2430,7 @@ private:
A.Offset = 0;
}
Out.Base = LoadEffectiveAddress(this, A, CTX->GetGPROpSize(), true, false);
Out.Base = LoadEffectiveAddress(this, A, GetGPROpSize(), true, false);
return Out;
}
@@ -2452,7 +2461,7 @@ private:
Ref _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, IR::OpSize Size, AddressMode A, Ref Value, IR::OpSize Align = IR::OpSize::i8Bit) {
bool AtomicTSO = IsTSOEnabled(Class) && !A.NonTSO;
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, AtomicTSO, Class != GPRClass, Size);
if (AtomicTSO) {
return _StoreMemTSO(Class, Size, Value, A.Base, A.Index, Align, A.IndexType, A.IndexScale);
@@ -2494,10 +2503,139 @@ private:
return Value;
}
Ref VZeroExtendOperand(OpSize Size, X86Tables::DecodedOperand Op, Ref Value) {
bool IsMMX = Op.IsGPR() && Op.Data.GPR.GPR >= X86State::REG_MM_0;
bool AlreadyExtended = Op.IsGPRDirect() || Op.IsGPRIndirect() || IsMMX;
return AlreadyExtended ? Value : _VMov(Size, Value);
}
void Push(IR::OpSize Size, Ref Value) {
auto OldSP = LoadGPRRegister(X86State::REG_RSP);
auto NewSP = _Push(CTX->GetGPROpSize(), Size, Value, OldSP);
auto NewSP = _Push(GetGPROpSize(), Size, Value, OldSP);
StoreGPRRegister(X86State::REG_RSP, NewSP);
FlushRegisterCache();
}
struct ArithRef {
IREmitter* E {};
bool IsConstant {};
union {
Ref R {};
uint64_t C;
};
ArithRef() {}
ArithRef(IREmitter* IREmit, Ref Reference)
: E(IREmit)
, IsConstant(false)
, R(Reference) {}
ArithRef(IREmitter* IREmit, uint64_t K)
: E(IREmit)
, IsConstant(true)
, C(K) {}
ArithRef Neg() {
return IsConstant ? ArithRef(E, -C) : ArithRef(E, E->_Neg(OpSize::i64Bit, R));
}
ArithRef And(uint64_t K) {
return IsConstant ? ArithRef(E, C & K) : ArithRef(E, E->_And(OpSize::i64Bit, R, E->Constant(K)));
}
ArithRef Presub(uint64_t K) {
return IsConstant ? ArithRef(E, K - C) : ArithRef(E, E->_Sub(OpSize::i64Bit, E->Constant(K), R));
}
ArithRef Lshl(uint64_t Shift) {
if (Shift == 0) {
return *this;
} else if (IsConstant) {
return ArithRef(E, C << Shift);
} else {
return ArithRef(E, E->_Lshl(OpSize::i64Bit, R, E->Constant(Shift)));
}
}
ArithRef Bfe(unsigned Start, unsigned Size) {
if (IsConstant) {
return ArithRef(E, (C >> Start) & ((1ull << Size) - 1));
} else {
return ArithRef(E, E->_Bfe(OpSize::i64Bit, Size, Start, R));
}
}
ArithRef Sbfe(unsigned Start, unsigned Size) {
if (IsConstant) {
uint64_t SourceMask = Size == 64 ? ~0ULL : ((1ULL << Size) - 1);
SourceMask <<= Start;
int64_t NewConstant = (C & SourceMask) >> Start;
NewConstant <<= 64 - Size;
NewConstant >>= 64 - Size;
return ArithRef(E, NewConstant);
} else {
return ArithRef(E, E->_Sbfe(OpSize::i64Bit, Size, Start, R));
}
}
Ref BfiInto(Ref Bitfield, unsigned Start, unsigned Size) {
if (IsConstant && (Size > 0 && Size < 64)) {
uint64_t SourceMask = (1ULL << Size) - 1;
uint64_t SourceMaskShifted = SourceMask << Start;
if (C == 0) {
return E->_And(OpSize::i64Bit, Bitfield, E->_InlineConstant(~SourceMaskShifted));
} else if (C == SourceMask) {
return E->_Or(OpSize::i64Bit, Bitfield, E->_InlineConstant(SourceMaskShifted));
}
}
if (IsConstant) {
return E->_Bfi(OpSize::i64Bit, Size, Start, Bitfield, E->Constant(C));
} else {
return E->_Bfi(OpSize::i64Bit, Size, Start, Bitfield, R);
}
}
ArithRef MaskBit(OpSize Size) {
if (IsConstant) {
uint64_t ShiftMask = Size == OpSize::i64Bit ? 63 : 31;
uint64_t Result = 1ull << (C & ShiftMask);
if (ShiftMask == 31) {
Result &= ((1ull << 32) - 1);
}
return ArithRef(E, Result);
} else {
return ArithRef(E, E->_Lshl(Size, E->Constant(1), R));
}
}
Ref Ref() {
return IsConstant ? E->Constant(C) : R;
}
bool IsDefinitelyZero() const {
return IsConstant && C == 0;
}
};
ArithRef ARef(Ref R) {
uint64_t C;
if (IsValueConstant(WrapNode(R), &C)) {
return ARef(C);
} else {
return ArithRef(this, R);
}
}
ArithRef ARef(uint64_t K) {
return ArithRef(this, K);
}
void InstallHostSpecificOpcodeHandlers();
@@ -2510,9 +2648,9 @@ private:
constexpr inline void InstallToTable(auto& FinalTable, const auto& LocalTable) {
for (const auto& Op : LocalTable) {
auto OpNum = std::get<0>(Op);
auto Dispatcher = std::get<2>(Op);
for (uint8_t i = 0; i < std::get<1>(Op); ++i) {
auto OpNum = Op.Op;
auto Dispatcher = Op.Ptr;
for (uint8_t i = 0; i < Op.Count; ++i) {
auto& TableOp = FinalTable[OpNum + i];
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
if (TableOp.OpcodeDispatcher) {
@@ -23,7 +23,7 @@ class OrderedNode;
void OpDispatchBuilder::InstallAVX128Handlers() {
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
static constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> AVX128Table[] = {
static constexpr DispatchTableEntry AVX128Table[] = {
{OPD(1, 0b00, 0x10), 1, &OpDispatchBuilder::AVX128_VMOVAPS},
{OPD(1, 0b01, 0x10), 1, &OpDispatchBuilder::AVX128_VMOVAPS},
{OPD(1, 0b10, 0x10), 1, &OpDispatchBuilder::AVX128_VMOVSS},
@@ -426,7 +426,7 @@ void OpDispatchBuilder::InstallAVX128Handlers() {
#undef OPD
#define OPD(group, pp, opcode) (((group - X86Tables::TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
static constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> VEX128TableGroupOps[] {
static constexpr DispatchTableEntry VEX128TableGroupOps[] {
// VPSRLI
{OPD(X86Tables::TYPE_VEX_GROUP_12, 1, 0b010), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::AVX128_VectorShiftImmImpl, OpSize::i16Bit, IROps::OP_VUSHRI>},
@@ -465,7 +465,7 @@ void OpDispatchBuilder::InstallAVX128Handlers() {
#undef OPD
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> VEX128_PCLMUL[] = {
constexpr DispatchTableEntry VEX128_PCLMUL[] = {
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::AVX128_VPCLMULQDQ},
};
#undef OPD
@@ -778,7 +778,7 @@ void OpDispatchBuilder::AVX128_VectorXOR(OpcodeArgs) {
void OpDispatchBuilder::AVX128_VZERO(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto IsVZEROALL = DstSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
if (IsVZEROALL) {
// NOTE: Despite the name being VZEROALL, this will still only ever
@@ -845,7 +845,7 @@ void OpDispatchBuilder::AVX128_MOVQ(OpcodeArgs) {
// This instruction is a bit special that if the destination is a register then it'll ZEXT the 64bit source to 256bit
if (Op->Dest.IsGPR()) {
// Zero bits [127:64] as well.
Src.Low = _VMov(OpSize::i64Bit, Src.Low);
Src.Low = VZeroExtendOperand(OpSize::i64Bit, Op->Src[0], Src.Low);
Ref ZeroVector = LoadZeroVector(OpSize::i128Bit);
Src.High = ZeroVector;
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Src);
@@ -968,7 +968,7 @@ void OpDispatchBuilder::AVX128_VBROADCAST(OpcodeArgs) {
}
} else {
// Get the address to broadcast from into a GPR.
Ref Address = MakeSegmentAddress(Op, Op->Src[0], CTX->GetGPROpSize());
Ref Address = MakeSegmentAddress(Op, Op->Src[0], GetGPROpSize());
Src.Low = _VBroadcastFromMem(OpSize::i128Bit, ElementSize, Address);
}
@@ -1022,7 +1022,7 @@ void OpDispatchBuilder::AVX128_InsertCVTGPR_To_FPR(OpcodeArgs) {
if (Op->Src[1].IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Op->Src[1], CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Op->Src[1], GetGPROpSize(), Op->Flags);
Result.Low = _VSToFGPRInsert(OpSize::i128Bit, DstElementSize, SrcSize, Src1.Low, Src2, false);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
@@ -1054,7 +1054,7 @@ void OpDispatchBuilder::AVX128_CVTFPR_To_GPR(OpcodeArgs) {
if (Op->Src[0].IsGPR()) {
Src = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false);
} else {
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSizeFromSrc(Op), Op->Flags);
Src.Low = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcElementSize, Op->Flags);
}
Ref Result = CVTFPR_To_GPRImpl(Op, Src.Low, SrcElementSize, HostRoundingMode);
@@ -1094,7 +1094,7 @@ void OpDispatchBuilder::AVX128_VPSIGN(OpcodeArgs) {
template<IR::OpSize ElementSize>
void OpDispatchBuilder::AVX128_UCOMISx(OpcodeArgs) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false);
@@ -1179,7 +1179,7 @@ void OpDispatchBuilder::AVX128_MOVBetweenGPR_FPR(OpcodeArgs) {
RefPair Result {};
if (Op->Src[0].IsGPR()) {
// Loading from GPR and moving to Vector.
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], CTX->GetGPROpSize(), Op->Flags);
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], GetGPROpSize(), Op->Flags);
// zext to 128bit
Result.Low = _VCastFromGPR(OpSize::i128Bit, OpSizeFromSrc(Op), Src);
} else {
@@ -1227,7 +1227,7 @@ void OpDispatchBuilder::AVX128_PExtr(OpcodeArgs) {
Index &= NumElements - 1;
if (Op->Dest.IsGPR()) {
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
// Extract already zero extends the result.
Ref Result = _VExtractToGPR(OpSize::i128Bit, OverridenElementSize, Src.Low, Index);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, GPRSize, OpSize::iInvalid);
@@ -1309,7 +1309,7 @@ void OpDispatchBuilder::AVX128_MOVMSK(OpcodeArgs) {
// Inserting the full lower 32-bits offset 31 so the sign bit ends up at offset 63.
GPR = _Bfi(OpSize::i64Bit, 32, 31, GPR, GPR);
// Shift right to only get the two sign bits we care about.
return _Lshr(OpSize::i64Bit, GPR, _Constant(62));
return _Lshr(OpSize::i64Bit, GPR, Constant(62));
};
auto Mask4Byte = [this](Ref Src) {
@@ -1341,7 +1341,7 @@ void OpDispatchBuilder::AVX128_MOVMSK(OpcodeArgs) {
auto GPRHigh = Mask8Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 2);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, CTX->GetGPROpSize(), OpSize::iInvalid);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
}
void OpDispatchBuilder::AVX128_MOVMSKB(OpcodeArgs) {
@@ -1383,7 +1383,7 @@ void OpDispatchBuilder::AVX128_PINSRImpl(OpcodeArgs, IR::OpSize ElementSize, con
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
Result.Low = _VInsGPR(OpSize::i128Bit, ElementSize, Index, Src1.Low, Src2);
} else {
// If loading from memory then we only load the element size
@@ -2055,7 +2055,7 @@ void OpDispatchBuilder::AVX128_VMASKMOVImpl(OpcodeArgs, IR::OpSize ElementSize,
auto Mask = AVX128_LoadSource_WithOpSize(Op, MaskOp, Op->Flags, !Is128Bit);
const auto MakeAddress = [this, Op](const X86Tables::DecodedOperand& Data) {
return MakeSegmentAddress(Op, Data, CTX->GetGPROpSize());
return MakeSegmentAddress(Op, Data, GetGPROpSize());
};
if (IsStore) {
@@ -2148,7 +2148,7 @@ void OpDispatchBuilder::AVX128_VectorVariableBlend(OpcodeArgs) {
}
void OpDispatchBuilder::AVX128_SaveAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
RefPair Pair = LoadContextPair(OpSize::i128Bit, AVXHigh0Index + i);
@@ -2157,7 +2157,7 @@ void OpDispatchBuilder::AVX128_SaveAVXState(Ref MemBase) {
}
void OpDispatchBuilder::AVX128_RestoreAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
auto YMMHRegs = LoadMemPair(FPRClass, OpSize::i128Bit, MemBase, i * 16 + 576);
@@ -2168,7 +2168,7 @@ void OpDispatchBuilder::AVX128_RestoreAVXState(Ref MemBase) {
}
void OpDispatchBuilder::AVX128_DefaultAVXState() {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
auto ZeroRegister = LoadZeroVector(OpSize::i128Bit);
for (uint32_t i = 0; i < NumRegs; i++) {
@@ -2212,8 +2212,8 @@ void OpDispatchBuilder::AVX128_VTESTP(OpcodeArgs) {
// For 256-bit, we need to split up the operation. This is nontrivial.
// Let's go the simple route here.
Ref ZF, CFInv;
Ref ZeroConst = _Constant(0);
Ref OneConst = _Constant(1);
Ref ZeroConst = Constant(0);
Ref OneConst = Constant(1);
const auto ElementSizeInBits = IR::OpSizeAsBits(ElementSize);
@@ -2294,8 +2294,8 @@ void OpDispatchBuilder::AVX128_PTest(OpcodeArgs) {
Test1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i16Bit, Test1, 0);
Test2 = _VExtractToGPR(OpSize::i128Bit, OpSize::i16Bit, Test2, 0);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
auto ZeroConst = Constant(0);
auto OneConst = Constant(1);
Test2 = _Select(FEXCore::IR::COND_NEQ, Test2, ZeroConst, OneConst, ZeroConst);
@@ -2328,7 +2328,7 @@ void OpDispatchBuilder::AVX128_VPERMD(OpcodeArgs) {
RefPair Result {};
Ref IndexMask = _VectorImm(OpSize::i128Bit, OpSize::i32Bit, 0b111);
Ref AddConst = _Constant(0x03020100);
Ref AddConst = Constant(0x03020100);
Ref Repeating3210 = _VDupFromGPR(OpSize::i128Bit, OpSize::i32Bit, AddConst);
Result.Low = DoPerm(Src, Indices.Low, IndexMask, Repeating3210);
@@ -2454,20 +2454,21 @@ void OpDispatchBuilder::AVX128_VFMAImpl(OpcodeArgs, IROps IROp, uint8_t Src1Idx,
}
void OpDispatchBuilder::AVX128_VFMAScalarImpl(OpcodeArgs, IROps IROp, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx) {
const auto Size = GetDstSize(Op);
const auto Is128Bit = Size == Core::CPUState::XMM_SSE_REG_SIZE;
LOGMAN_THROW_A_FMT(Is128Bit, "This can't be 256-bit");
const auto SrcSize = OpSizeFromSrc(Op);
const OpSize ElementSize = Op->Flags & X86Tables::DecodeFlags::FLAG_OPTION_AVX_W ? OpSize::i64Bit : OpSize::i32Bit;
auto Dest = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, !Is128Bit).Low;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, !Is128Bit).Low;
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit).Low;
auto Dest = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false).Low;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Src[0], Op->Flags, false).Low;
Ref Src2 {};
if (Op->Src[1].IsGPR()) {
Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, false).Low;
} else {
Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], SrcSize, Op->Flags);
}
Ref Sources[3] = {Dest, Src1, Src2};
DeriveOp(Result_Low, IROp,
_VFMLAScalarInsert(OpSize::i128Bit, ElementSize, Dest, Sources[Src1Idx - 1], Sources[Src2Idx - 1], Sources[AddendIdx - 1]));
_VFMLAScalarInsert(OpSize::i128Bit, SrcSize, Dest, Sources[Src1Idx - 1], Sources[Src2Idx - 1], Sources[AddendIdx - 1]));
AVX128_StoreResult_WithOpSize(Op, Op->Dest, AVX128_Zext(Result_Low));
}
@@ -2517,9 +2518,9 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherImpl(OpSize Size, O
///< BaseAddr doesn't need to exist, calculate that here.
Ref BaseAddr = VSIB.BaseAddr;
if (BaseAddr && VSIB.Displacement) {
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, _Constant(VSIB.Displacement));
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, Constant(VSIB.Displacement));
} else if (VSIB.Displacement) {
BaseAddr = _Constant(VSIB.Displacement);
BaseAddr = Constant(VSIB.Displacement);
} else if (!BaseAddr) {
BaseAddr = Invalid();
}
@@ -2612,9 +2613,9 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_VPGatherQPSImpl(Ref Dest, R
///< BaseAddr doesn't need to exist, calculate that here.
Ref BaseAddr = VSIB.BaseAddr;
if (BaseAddr && VSIB.Displacement) {
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, _Constant(VSIB.Displacement));
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, Constant(VSIB.Displacement));
} else if (VSIB.Displacement) {
BaseAddr = _Constant(VSIB.Displacement);
BaseAddr = Constant(VSIB.Displacement);
} else if (!BaseAddr) {
BaseAddr = Invalid();
}
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable[] = {
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable[] = {
// Instructions
{0x00, 6, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ALUOp, FEXCore::IR::IROps::OP_ADD, FEXCore::IR::IROps::OP_ATOMICFETCHADD, 0>},
@@ -76,12 +76,12 @@ constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispat
{0xFC, 2, &OpDispatchBuilder::FLAGControlOp},
};
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable_64[] = {
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable_64[] = {
{0x63, 1, &OpDispatchBuilder::MOVSXDOp},
{0xA0, 4, &OpDispatchBuilder::MOVOffsetOp},
};
constexpr inline std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_BaseOpTable_32[] = {
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable_32[] = {
{0x06, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x07, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x0E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX>},
@@ -11,10 +11,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
#include <array>
#include <cstdint>
#include <tuple>
#include <utility>
namespace FEXCore::IR {
class OrderedNode;
@@ -25,21 +22,13 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref RotatedNode {};
if (CTX->HostFeatures.SupportsSHA) {
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
} else {
// SHA1 extension missing, manually rotate.
// Emulate rotate.
auto ShiftLeft = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Dest, 30);
RotatedNode = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeft, Dest, 2);
}
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
@@ -62,153 +51,49 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
auto Src2 = SHADataShuffle(Src);
// The result is swizzled differently than expected
Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
} else {
// Shift the incoming source left by a 32-bit element, inserting Zeros.
// This could be slightly improved to use a VInsGPR with the zero register.
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
auto Src2Shift = _VExtr(OpSize::i128Bit, OpSize::i8Bit, Src, ZeroRegister, 12);
auto Xor1 = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, Src2Shift);
// Emulate rotate.
auto ShiftLeftXor1 = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Xor1, 1);
auto RotatedXor1 = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXor1, Xor1, 31);
// Element0 didn't get XOR'd with anything, so do it now.
auto ExtractUpper = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, RotatedXor1, 3);
auto XorLower = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, ExtractUpper);
// Emulate rotate.
auto ShiftLeftXorLower = _VShlI(OpSize::i128Bit, OpSize::i32Bit, XorLower, 1);
auto RotatedXorLower = _VUShraI(OpSize::i128Bit, OpSize::i32Bit, ShiftLeftXorLower, XorLower, 31);
Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 0, 0, RotatedXor1, RotatedXorLower);
}
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
auto Src2 = SHADataShuffle(Src);
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
using FnType = Ref (*)(OpDispatchBuilder&, Ref, Ref, Ref);
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1c?
return Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._Andn(OpSize::i32Bit, D, B));
};
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p with different key
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1m
return Self.BitwiseAtLeastTwo(B, C, D);
};
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref { // sha1p
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
constexpr std::array<uint32_t, 4> k_array {
0x5A827999U,
0x6ED9EBA1U,
0x8F1BBCDCU,
0xCA62C1D6U,
};
constexpr std::array<FnType, 4> fn_array {
f0,
f1,
f2,
f3,
};
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Ref ConstantVector {};
switch (Imm8) {
case 0:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0);
break;
case 1:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1);
break;
case 2:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2);
break;
case 3:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3);
break;
}
Ref ConstantVector {};
switch (Imm8) {
case 0:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K0);
break;
case 1:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K1);
break;
case 2:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K2);
break;
case 3:
ConstantVector = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_SHA1RNDS_K3);
break;
}
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
Src2 = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src2, ConstantVector);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
Src2 = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src2, ConstantVector);
switch (Imm8) {
case 0: Result = SHADataShuffle(_VSha1C(Src1, ZeroRegister, Src2)); break;
case 2: Result = SHADataShuffle(_VSha1M(Src1, ZeroRegister, Src2)); break;
case 1:
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
} else {
const FnType Fn = fn_array[Imm8];
auto K = _Constant(OpSize::i32Bit, k_array[Imm8]);
auto W0E = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto B = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto C = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto D = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
auto A1 =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto D1 = C;
auto E1 = D;
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
auto ANext =
_Add(OpSize::i32Bit,
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(OpSize::i32Bit, 2));
auto DNext = C;
auto ENext = D;
return {ANext, BNext, CNext, DNext, ENext};
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Dest3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Dest2, std::get<2>(Final));
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Dest1, std::get<3>(Final));
switch (Imm8) {
case 0: Result = SHADataShuffle(_VSha1C(Src1, ZeroRegister, Src2)); break;
case 2: Result = SHADataShuffle(_VSha1M(Src1, ZeroRegister, Src2)); break;
case 1:
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
@@ -218,69 +103,20 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Result = _VSha256U0(Dest, Src);
} else {
const auto Sigma0 = [this](Ref W) -> Ref {
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 7)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 18))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 3)));
};
auto W4 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto W3 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto W2 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto W1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
auto W0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
auto Sig3 = _Add(OpSize::i32Bit, W3, Sigma0(W4));
auto Sig2 = _Add(OpSize::i32Bit, W2, Sigma0(W3));
auto Sig1 = _Add(OpSize::i32Bit, W1, Sigma0(W2));
auto Sig0 = _Add(OpSize::i32Bit, W0, Sigma0(W1));
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, Sig3);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, Sig2);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, Sig1);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, Sig0);
}
auto Result = _VSha256U0(Dest, Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](Ref W) -> Ref {
return _Xor(
OpSize::i32Bit,
_Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 17)), _Ror(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 19))),
_Lshr(OpSize::i32Bit, W, _Constant(OpSize::i32Bit, 10)));
};
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Src2 = _VZip2(OpSize::i128Bit, OpSize::i64Bit, DupDst, Src);
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Src2 = _VZip2(OpSize::i128Bit, OpSize::i64Bit, DupDst, Src);
Result = _VSha256U1(Src1, Src2);
} else {
auto W14 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto W15 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto W16 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0), Sigma1(W14));
auto W17 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1), Sigma1(W15));
auto W18 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2), Sigma1(W16));
auto W19 = _Add(OpSize::i32Bit, _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3), Sigma1(W17));
auto D3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, W19);
auto D2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, D3, W18);
auto D1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, D2, W17);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, D1, W16);
}
auto Result = _VSha256U1(Src1, Src2);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
@@ -301,81 +137,27 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
Ref Result;
if (CTX->HostFeatures.SupportsSHA) {
auto shuffle_abcd = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `abcd` configuration from x86 format.
auto Tmp = _VZip2(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_abcd = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `abcd` configuration from x86 format.
auto Tmp = _VZip2(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_efgh = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `efgh` configuration from x86 format.
auto Tmp = _VZip(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto shuffle_efgh = [this](Ref Src1, Ref Src2) -> Ref {
// Generates a suitable SHA256 `efgh` configuration from x86 format.
auto Tmp = _VZip(OpSize::i128Bit, OpSize::i64Bit, Src2, Src1);
return _VRev64(OpSize::i128Bit, OpSize::i32Bit, Tmp);
};
auto ABCD = shuffle_abcd(Dest, Src);
auto EFGH = shuffle_efgh(Dest, Src);
auto ABCD = shuffle_abcd(Dest, Src);
auto EFGH = shuffle_efgh(Dest, Src);
// x86 uses only the bottom 64-bits of the key, so duplicate to match ARM64 semantics.
auto Key = _VDupElement(OpSize::i128Bit, OpSize::i64Bit, XMM0, 0);
// x86 uses only the bottom 64-bits of the key, so duplicate to match ARM64 semantics.
auto Key = _VDupElement(OpSize::i128Bit, OpSize::i64Bit, XMM0, 0);
auto A = _VSha256H(ABCD, EFGH, Key);
auto B = _VSha256H2(EFGH, ABCD, Key);
Result = shuffle_abcd(A, B);
} else {
const auto Ch = [this](Ref E, Ref F, Ref G) -> Ref {
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
};
const auto Sigma0 = [this](Ref A) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(OpSize::i32Bit, 2)), A, ShiftType::ROR, 13),
A, ShiftType::ROR, 22);
};
const auto Sigma1 = [this](Ref E) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(OpSize::i32Bit, 6)), E, ShiftType::ROR, 11),
E, ShiftType::ROR, 25);
};
auto E0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 1);
auto F0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 0);
auto G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Ref Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
auto WK0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 0);
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
auto H0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 0);
Q0 = _Add(OpSize::i32Bit, Q0, H0);
auto A0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 3);
auto B0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Src, 2);
auto C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto A1 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q0, BitwiseAtLeastTwo(A0, B0, C0)), Sigma0(A0));
auto D0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 2);
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
Ref Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
auto WK1 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, XMM0, 1);
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
// Rematerialize G0. Costs a move but saves spilling, coming out ahead.
G0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 1);
Q1 = _Add(OpSize::i32Bit, Q1, G0);
auto A2 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q1, BitwiseAtLeastTwo(A1, A0, B0)), Sigma0(A1));
// Rematerialize C0. As with G0.
C0 = _VExtractToGPR(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto E2 = _Add(OpSize::i32Bit, Q1, C0);
auto Res3 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 3, Dest, A2);
auto Res2 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 2, Res3, A1);
auto Res1 = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 1, Res2, E2);
Result = _VInsGPR(OpSize::i128Bit, OpSize::i32Bit, 0, Res1, E1);
}
auto A = _VSha256H(ABCD, EFGH, Key);
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_DDDTable[] = {
constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
@@ -28,7 +28,7 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
void OpDispatchBuilder::ZeroPF_AF() {
// PF is stored inverted, so invert it when we zero.
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(_Constant(1));
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(Constant(1));
SetAF(0);
}
@@ -228,21 +228,26 @@ void OpDispatchBuilder::CalculateAF(Ref Src1, Ref Src2) {
// We only care about bit 4 in the subsequent XOR. If we'll XOR with 0,
// there's no sense XOR'ing at all. If we'll XOR with 1, that's just
// inverting.
uint64_t Const;
if (IsValueConstant(WrapNode(Src2), &Const)) {
if (Const & (1u << 4)) {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Not(OpSize::i32Bit, Src1));
} else {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Src1);
}
for (unsigned i = 0; i < 2; ++i) {
Ref SrcA = i ? Src1 : Src2;
Ref SrcB = i ? Src2 : Src1;
return;
uint64_t Const;
if (IsValueConstant(WrapNode(SrcA), &Const)) {
if (Const & (1u << 4)) {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Not(OpSize::i32Bit, SrcB));
} else {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(SrcB);
}
return;
}
}
// We store the XOR of the arguments. At read time, we XOR with the
// appropriate bit of the result (available as the PF flag) and extract the
// appropriate bit. Again 64-bit to avoid masking.
Ref XorRes = _Xor(OpSize::i64Bit, Src1, Src2);
Ref XorRes = Src1 == Src2 ? Constant(0) : _Xor(OpSize::i64Bit, Src1, Src2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -276,7 +281,7 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
CFInverted = false;
} else {
// Need to zero-extend for correct comparisons below
Src2 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src2);
Src2 = ARef(Src2).Bfe(0, IR::OpSizeAsBits(SrcSize)).Ref();
// Note that we do not extend Src2PlusCF, since we depend on proper
// 32-bit arithmetic to correctly handle the Src2 = 0xffff case.
@@ -316,7 +321,7 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2
} else {
// Zero extend for correct comparison behaviour with Src1 = 0xffff.
Src1 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src1);
Src2 = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Src2);
Src2 = ARef(Src2).Bfe(0, IR::OpSizeAsBits(SrcSize)).Ref();
auto Src2PlusCF = IncrementByCarry(OpSize, Src2);
@@ -426,13 +431,9 @@ void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
CFInverted = true;
}
void OpDispatchBuilder::CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res, Ref Src1, Ref Src2) {
void OpDispatchBuilder::CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res) {
InvalidateAF();
CalculatePF(Res);
// SF/ZF/CF/OF
SetNZ_ZeroCV(SrcSize, Res);
SetNZP_ZeroCV(SrcSize, Res);
}
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift) {
@@ -8,7 +8,7 @@ constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F38Table[] = {
constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_66, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_NONE, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADDP, OpSize::i16Bit>},
@@ -8,7 +8,7 @@ namespace FEXCore::IR {
#define PF_3A_66 1
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> Table[] = {
constexpr DispatchTableEntry Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
@@ -42,8 +42,8 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
auto REX0 = OpDispatchTableGenH0F3AREX.template operator()<0>();
auto REX1 = OpDispatchTableGenH0F3AREX.template operator()<1>();
auto concat = []<typename T, size_t N1, size_t N2>(std::array<T, N1> const& lhs,
std::array<T, N2> const& rhs) consteval -> std::array<T, N1 + N2> {
auto concat = []<typename T, size_t N1, size_t N2>(const std::array<T, N1>& lhs,
const std::array<T, N2>& rhs) consteval -> std::array<T, N1 + N2> {
std::array<T, N1 + N2> Table {};
for (size_t i = 0; i < N1; ++i) {
Table[i] = lhs[i];
@@ -60,12 +60,12 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATableNeedsREX0[] = {
constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_H0F3ATable_64[] = {
constexpr DispatchTableEntry OpDispatch_H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i64Bit>},
{OPD(1, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i64Bit>},
};
@@ -5,7 +5,7 @@
namespace FEXCore::IR {
using X86Tables::OpToIndex;
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_PrimaryGroupTables[] = {
constexpr DispatchTableEntry OpDispatch_PrimaryGroupTables[] = {
// GROUP 1
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -8,7 +8,7 @@ constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryGroupTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 6
{OPD(FEXCore::X86Tables::TYPE_GROUP_6, PF_NONE, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_6, PF_F3, 3), 1, &OpDispatchBuilder::PermissionRestrictedOp},
@@ -113,11 +113,13 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_NONE, 7), 1, &OpDispatchBuilder::StoreFenceOrCLFlush}, // SFENCE (or CLFLUSH)
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 5), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 6), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 6), 1, &OpDispatchBuilder::UMonitorOrCLRSSBSY},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 6), 1, &OpDispatchBuilder::CLWB},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 6), 1, &OpDispatchBuilder::CLWBOrTPause},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_66, 7), 1, &OpDispatchBuilder::CLFLUSHOPT},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F2, 6), 1, &OpDispatchBuilder::UMWaitOp},
// GROUP 16
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_NONE, 0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, true, 1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_NONE, 1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 1>},
@@ -154,7 +156,7 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_F2, 0), 8, &OpDispatchBuilder::NOPOp},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryGroupTables_64[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables_64[] = {
// GROUP 15
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 0), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::ReadSegmentReg, OpDispatchBuilder::Segment::FS>},
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryModRMTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryModRMTables[] = {
// REG /1
{((0 << 3) | 0), 1, &OpDispatchBuilder::UnimplementedOp},
{((0 << 3) | 1), 1, &OpDispatchBuilder::UnimplementedOp},
@@ -3,7 +3,7 @@
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable[] = {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
@@ -150,7 +150,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
#endif
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryRepModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
@@ -181,7 +181,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryRepNEModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
@@ -207,7 +207,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_SecondaryOpSizeModTables[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
@@ -314,7 +314,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable_64[] = {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable_64[] = {
{0x05, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SyscallOp, true>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
@@ -322,7 +322,7 @@ constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDisp
{0xA9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
};
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_TwoByteOpTable_32[] = {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable_32[] = {
{0x05, 1, &OpDispatchBuilder::NOPOp},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
@@ -4,7 +4,7 @@
namespace FEXCore::IR {
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDispatch_VEXTable[] = {
constexpr DispatchTableEntry OpDispatch_VEXTable[] = {
{OPD(2, 0b00, 0xF2), 1, &OpDispatchBuilder::ANDNBMIOp}, {OPD(2, 0b00, 0xF5), 1, &OpDispatchBuilder::BZHI},
{OPD(2, 0b10, 0xF5), 1, &OpDispatchBuilder::PEXT}, {OPD(2, 0b11, 0xF5), 1, &OpDispatchBuilder::PDEP},
{OPD(2, 0b11, 0xF6), 1, &OpDispatchBuilder::MULX}, {OPD(2, 0b00, 0xF7), 1, &OpDispatchBuilder::BEXTRBMIOp},
@@ -16,7 +16,7 @@ constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpDis
#undef OPD
#define OPD(group, pp, opcode) (((group - X86Tables::InstType::TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> OpDispatch_VEXGroupTable[] = {
constexpr DispatchTableEntry OpDispatch_VEXGroupTable[] = {
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b001), 1, &OpDispatchBuilder::BLSRBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b010), 1, &OpDispatchBuilder::BLSMSKBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b011), 1, &OpDispatchBuilder::BLSIBMIOp},
@@ -78,7 +78,7 @@ void OpDispatchBuilder::VMOVAPS_VMOVAPDOp(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (Is128Bit && Op->Dest.IsGPR()) {
Src = _VMov(OpSize::i128Bit, Src);
Src = VZeroExtendOperand(OpSize::i128Bit, Op->Src[0], Src);
}
StoreResult(FPRClass, Op, Src, OpSize::iInvalid);
}
@@ -90,7 +90,7 @@ void OpDispatchBuilder::VMOVUPS_VMOVUPDOp(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
if (Is128Bit && Op->Dest.IsGPR()) {
Src = _VMov(OpSize::i128Bit, Src);
Src = VZeroExtendOperand(OpSize::i128Bit, Op->Src[0], Src);
}
StoreResult(FPRClass, Op, Src, OpSize::i8Bit);
}
@@ -434,7 +434,7 @@ Ref OpDispatchBuilder::InsertCVTGPR_To_FPRImpl(OpcodeArgs, IR::OpSize DstSize, I
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
return _VSToFGPRInsert(DstSize, DstElementSize, SrcSize, Src1, Src2, ZeroUpperBits);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
@@ -706,7 +706,7 @@ void OpDispatchBuilder::MOVQOp(OpcodeArgs, VectorOpType VectorType) {
const auto gpr = Op->Dest.Data.GPR.GPR;
const auto gprIndex = gpr - X86State::REG_XMM_0;
auto Reg = _VMov(OpSize::i64Bit, Src);
auto Reg = VZeroExtendOperand(OpSize::i64Bit, Op->Src[0], Src);
StoreXMMRegister_WithAVXInsert(VectorType, gprIndex, Reg);
} else {
// This is simple, just store the result
@@ -740,8 +740,8 @@ void OpDispatchBuilder::MOVMSKOp(OpcodeArgs, IR::OpSize ElementSize) {
// Inserting the full lower 32-bits offset 31 so the sign bit ends up at offset 63.
GPR = _Bfi(OpSize::i64Bit, 32, 31, GPR, GPR);
// Shift right to only get the two sign bits we care about.
GPR = _Lshr(OpSize::i64Bit, GPR, _Constant(62));
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, CTX->GetGPROpSize(), OpSize::iInvalid);
GPR = _Lshr(OpSize::i64Bit, GPR, Constant(62));
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
} else if (Size == OpSize::i128Bit && ElementSize == OpSize::i32Bit) {
// Shift all the sign bits to the bottom of their respective elements.
Src = _VUShrI(Size, OpSize::i32Bit, Src, 31);
@@ -753,20 +753,21 @@ void OpDispatchBuilder::MOVMSKOp(OpcodeArgs, IR::OpSize ElementSize) {
Src = _VAddV(Size, OpSize::i32Bit, Src);
// Extract to a GPR.
Ref GPR = _VExtractToGPR(Size, OpSize::i32Bit, Src, 0);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, CTX->GetGPROpSize(), OpSize::iInvalid);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GPR, GetGPROpSize(), OpSize::iInvalid);
} else {
Ref CurrentVal = _Constant(0);
Ref CurrentVal = Constant(0);
for (unsigned i = 0; i < NumElements; ++i) {
// Extract the top bit of the element
Ref Tmp = _VExtractToGPR(Size, ElementSize, Src, i);
Tmp = _Bfe(ElementSize, 1, IR::OpSizeAsBits(ElementSize) - 1, Tmp);
// Shift it to the correct location
Tmp = _Lshl(ElementSize, Tmp, _Constant(i));
// Or it with the current value
CurrentVal = _Or(OpSize::i64Bit, CurrentVal, Tmp);
// Shift it to the correct location and or it with the current value
if (i != 0) {
CurrentVal = _Orlshl(OpSize::i64Bit, CurrentVal, Tmp, i);
} else {
CurrentVal = Tmp;
}
}
StoreResult(GPRClass, Op, CurrentVal, OpSize::iInvalid);
}
@@ -1503,7 +1504,7 @@ void OpDispatchBuilder::VBROADCASTOp(OpcodeArgs, IR::OpSize ElementSize) {
Result = _VDupElement(DstSize, ElementSize, Src, 0);
} else {
// Get the address to broadcast from into a GPR.
Ref Address = MakeSegmentAddress(Op, Op->Src[0], CTX->GetGPROpSize());
Ref Address = MakeSegmentAddress(Op, Op->Src[0], GetGPROpSize());
Result = _VBroadcastFromMem(DstSize, ElementSize, Address);
}
@@ -1522,7 +1523,7 @@ Ref OpDispatchBuilder::PINSROpImpl(OpcodeArgs, IR::OpSize ElementSize, const X86
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
return _VInsGPR(Size, ElementSize, Index, Src1, Src2);
}
@@ -1643,7 +1644,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs, IR::OpSize ElementSize) {
Index &= NumElements - 1;
if (Op->Dest.IsGPR()) {
const auto GPRSize = CTX->GetGPROpSize();
const auto GPRSize = GetGPROpSize();
// Extract already zero extends the result.
Ref Result = _VExtractToGPR(OpSize::i128Bit, OverridenElementSize, Src, Index);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, GPRSize, OpSize::iInvalid);
@@ -2064,7 +2065,7 @@ Ref OpDispatchBuilder::CVTGPR_To_FPRImpl(OpcodeArgs, IR::OpSize DstElementSize,
Ref Converted {};
if (Src2Op.IsGPR()) {
// If the source is a GPR then convert directly from the GPR.
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, CTX->GetGPROpSize(), Op->Flags);
auto Src2 = LoadSource_WithOpSize(GPRClass, Op, Src2Op, GetGPROpSize(), Op->Flags);
Converted = _Float_FromGPR_S(DstElementSize, SrcSize, Src2);
} else if (SrcSize != DstElementSize) {
// If the source is from memory but the Source size and destination size aren't the same,
@@ -2120,7 +2121,7 @@ Ref OpDispatchBuilder::CVTFPR_To_GPRImpl(OpcodeArgs, Ref Src, IR::OpSize SrcElem
Ref Converted = _Float_ToGPR_ZS(GPRSize, SrcElementSize, Src);
bool Dst32 = GPRSize == OpSize::i32Bit;
Ref MaxI = Dst32 ? _Constant(0x80000000) : _Constant(0x8000000000000000);
Ref MaxI = Dst32 ? Constant(0x80000000) : Constant(0x8000000000000000);
Ref MaxF = LoadAndCacheNamedVectorConstant(SrcElementSize, (SrcElementSize == OpSize::i32Bit) ?
(Dst32 ? NAMED_VECTOR_CVTMAX_F32_I32 : NAMED_VECTOR_CVTMAX_F32_I64) :
(Dst32 ? NAMED_VECTOR_CVTMAX_F64_I32 : NAMED_VECTOR_CVTMAX_F64_I64));
@@ -2133,7 +2134,7 @@ void OpDispatchBuilder::CVTFPR_To_GPR(OpcodeArgs) {
// If loading a vector, use the full size, so we don't
// unnecessarily zero extend the vector. Otherwise, if
// memory, then we want to load the element size exactly.
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : SrcElementSize;
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
Ref Result = CVTFPR_To_GPRImpl(Op, Src, SrcElementSize, HostRoundingMode);
StoreResult(GPRClass, Op, Result, OpSize::iInvalid);
@@ -2354,7 +2355,7 @@ void OpDispatchBuilder::VMASKMOVOpImpl(OpcodeArgs, IR::OpSize ElementSize, IR::O
const X86Tables::DecodedOperand& MaskOp, const X86Tables::DecodedOperand& DataOp) {
const auto MakeAddress = [this, Op](const X86Tables::DecodedOperand& Data) {
return MakeSegmentAddress(Op, Data, CTX->GetGPROpSize());
return MakeSegmentAddress(Op, Data, GetGPROpSize());
};
Ref Mask = LoadSource_WithOpSize(FPRClass, Op, MaskOp, DataSize, Op->Flags);
@@ -2397,7 +2398,7 @@ void OpDispatchBuilder::MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType) {
Ref Result {};
if (Op->Src[0].IsGPR()) {
// Loading from GPR and moving to Vector.
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], CTX->GetGPROpSize(), Op->Flags);
Ref Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], GetGPROpSize(), Op->Flags);
// zext to 128bit
Result = _VCastFromGPR(OpSize::i128Bit, OpSizeFromSrc(Op), Src);
} else {
@@ -2503,7 +2504,7 @@ Ref OpDispatchBuilder::XSaveBase(X86Tables::DecodedOp Op) {
void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// NOTE: Mask should be EAX and EDX concatenated, but we only need to test
// for features that are in the lower 32 bits, so EAX only is sufficient.
const auto OpSize = CTX->GetGPROpSize();
const auto OpSize = GetGPROpSize();
const auto StoreIfFlagSet = [this, OpSize](uint32_t BitIndex, auto fn, uint32_t FieldSize = 1) {
Ref Mask = LoadGPRRegister(X86State::REG_RAX);
@@ -2538,8 +2539,7 @@ void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// We need to save MXCSR and MXCSR_MASK if either SSE or AVX are requested to be saved
{
StoreIfFlagSet(
1, [this, Op] { SaveMXCSRState(XSaveBase(Op)); }, 2);
StoreIfFlagSet(1, [this, Op] { SaveMXCSRState(XSaveBase(Op)); }, 2);
}
// Update XSTATE_BV region of the XSAVE header
@@ -2552,7 +2552,7 @@ void OpDispatchBuilder::XSaveOpImpl(OpcodeArgs) {
// XSTATE_BV section of the header is 8 bytes in size, but we only really
// care about setting at most 3 bits in the first byte. We zero out the rest.
_StoreMem(GPRClass, OpSize::i64Bit, RequestedFeatures, Base, _Constant(512), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, OpSize::i64Bit, RequestedFeatures, Base, Constant(512), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
}
}
@@ -2576,11 +2576,11 @@ void OpDispatchBuilder::SaveX87State(OpcodeArgs, Ref MemBase) {
_StoreMem(GPRClass, OpSize::i16Bit, MemBase, FCW, OpSize::i16Bit);
}
{ _StoreMem(GPRClass, OpSize::i16Bit, ReconstructFSW_Helper(), MemBase, _Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, OpSize::i16Bit, ReconstructFSW_Helper(), MemBase, Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1); }
{
// Abridged FTW
_StoreMem(GPRClass, OpSize::i8Bit, LoadContext(AbridgedFTWIndex), MemBase, _Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, OpSize::i8Bit, LoadContext(AbridgedFTWIndex), MemBase, Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
}
// BYTE | 0 1 | 2 3 | 4 | 5 | 6 7 | 8 9 | a b | c d | e f |
@@ -2634,7 +2634,7 @@ void OpDispatchBuilder::SaveX87State(OpcodeArgs, Ref MemBase) {
}
void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
_StoreMemPair(FPRClass, OpSize::i128Bit, LoadXMMRegister(i), LoadXMMRegister(i + 1), MemBase, i * 16 + 160);
@@ -2643,11 +2643,11 @@ void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
void OpDispatchBuilder::SaveMXCSRState(Ref MemBase) {
// Store MXCSR and the mask for all bits.
_StoreMemPair(GPRClass, OpSize::i32Bit, GetMXCSR(), _Constant(0xFFFF), MemBase, 24);
_StoreMemPair(GPRClass, OpSize::i32Bit, GetMXCSR(), Constant(0xFFFF), MemBase, 24);
}
void OpDispatchBuilder::SaveAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
Ref Upper0 = _VDupElement(OpSize::i256Bit, OpSize::i128Bit, LoadXMMRegister(i + 0), 1);
@@ -2661,7 +2661,7 @@ Ref OpDispatchBuilder::GetMXCSR() {
Ref MXCSR = _LoadContext(OpSize::i32Bit, GPRClass, offsetof(FEXCore::Core::CPUState, mxcsr));
// Mask out unsupported bits
// Keeps FZ, RC, exception masks, and DAZ
MXCSR = _And(OpSize::i32Bit, MXCSR, _Constant(0xFFC0));
MXCSR = _And(OpSize::i32Bit, MXCSR, Constant(0xFFC0));
return MXCSR;
}
@@ -2671,12 +2671,12 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
RestoreX87State(Mem);
RestoreSSEState(Mem);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Mem, _Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Mem, Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
}
void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
const auto OpSize = CTX->GetGPROpSize();
const auto OpSize = GetGPROpSize();
// If a bit in our XSTATE_BV is set, then we restore from that region of the XSAVE area,
// otherwise, if not set, then we need to set the relevant data the bit corresponds to
@@ -2688,7 +2688,7 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
// Note: we rematerialize Base/Mask in each block to avoid crossblock
// liveness.
Ref Base = XSaveBase(Op);
Ref Mask = _LoadMem(GPRClass, OpSize::i64Bit, Base, _Constant(512), OpSize::i64Bit, MEM_OFFSET_SXTX, 1);
Ref Mask = _LoadMem(GPRClass, OpSize::i64Bit, Base, Constant(512), OpSize::i64Bit, MEM_OFFSET_SXTX, 1);
Ref BitFlag = _Bfe(OpSize, FieldSize, BitIndex, Mask);
auto CondJump_ = CondJump(BitFlag, {COND_NEQ});
@@ -2714,13 +2714,11 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
// x87
{
RestoreIfFlagSetOrDefault(
0, [this, Op] { RestoreX87State(XSaveBase(Op)); }, [this, Op] { DefaultX87State(Op); });
RestoreIfFlagSetOrDefault(0, [this, Op] { RestoreX87State(XSaveBase(Op)); }, [this, Op] { DefaultX87State(Op); });
}
// SSE
{
RestoreIfFlagSetOrDefault(
1, [this, Op] { RestoreSSEState(XSaveBase(Op)); }, [this] { DefaultSSEState(); });
RestoreIfFlagSetOrDefault(1, [this, Op] { RestoreSSEState(XSaveBase(Op)); }, [this] { DefaultSSEState(); });
}
// AVX
if (CTX->HostFeatures.SupportsAVX) {
@@ -2733,9 +2731,9 @@ void OpDispatchBuilder::XRstorOpImpl(OpcodeArgs) {
RestoreIfFlagSetOrDefault(
1,
[this, Op] {
Ref Base = XSaveBase(Op);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Base, _Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
Ref Base = XSaveBase(Op);
Ref MXCSR = _LoadMem(GPRClass, OpSize::i32Bit, Base, Constant(24), OpSize::i32Bit, MEM_OFFSET_SXTX, 1);
RestoreMXCSRState(MXCSR);
},
[] { /* Intentionally do nothing*/ }, 2);
}
@@ -2746,13 +2744,13 @@ void OpDispatchBuilder::RestoreX87State(Ref MemBase) {
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
{
auto NewFSW = _LoadMem(GPRClass, OpSize::i16Bit, MemBase, _Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, OpSize::i16Bit, MemBase, Constant(2), OpSize::i16Bit, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
}
{
// Abridged FTW
StoreContext(AbridgedFTWIndex, _LoadMem(GPRClass, OpSize::i8Bit, MemBase, _Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1));
StoreContext(AbridgedFTWIndex, _LoadMem(GPRClass, OpSize::i8Bit, MemBase, Constant(4), OpSize::i8Bit, MEM_OFFSET_SXTX, 1));
}
for (uint32_t i = 0; i < Core::CPUState::NUM_MMS; i += 2) {
@@ -2764,7 +2762,7 @@ void OpDispatchBuilder::RestoreX87State(Ref MemBase) {
}
void OpDispatchBuilder::RestoreSSEState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
auto XMMRegs = LoadMemPair(FPRClass, OpSize::i128Bit, MemBase, i * 16 + 160);
@@ -2776,7 +2774,7 @@ void OpDispatchBuilder::RestoreSSEState(Ref MemBase) {
void OpDispatchBuilder::RestoreMXCSRState(Ref MXCSR) {
// Mask out unsupported bits
MXCSR = _And(OpSize::i32Bit, MXCSR, _Constant(0xFFC0));
MXCSR = _And(OpSize::i32Bit, MXCSR, Constant(0xFFC0));
_StoreContext(OpSize::i32Bit, GPRClass, MXCSR, offsetof(FEXCore::Core::CPUState, mxcsr));
// We only support the rounding mode and FTZ bit being set
@@ -2785,7 +2783,7 @@ void OpDispatchBuilder::RestoreMXCSRState(Ref MXCSR) {
}
void OpDispatchBuilder::RestoreAVXState(Ref MemBase) {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i += 2) {
Ref XMMReg0 = LoadXMMRegister(i + 0);
@@ -2810,7 +2808,7 @@ void OpDispatchBuilder::DefaultX87State(OpcodeArgs) {
}
void OpDispatchBuilder::DefaultSSEState() {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
Ref ZeroVector = LoadZeroVector(OpSize::i128Bit);
for (uint32_t i = 0; i < NumRegs; ++i) {
@@ -2819,7 +2817,7 @@ void OpDispatchBuilder::DefaultSSEState() {
}
void OpDispatchBuilder::DefaultAVXState() {
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
for (uint32_t i = 0; i < NumRegs; i++) {
Ref Reg = LoadXMMRegister(i);
@@ -2876,7 +2874,7 @@ void OpDispatchBuilder::VPALIGNROp(OpcodeArgs) {
template<IR::OpSize ElementSize>
void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
Ref Src1 = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetGuestVectorLength(), Op->Flags);
Ref Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
@@ -2887,12 +2885,12 @@ template void OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>(OpcodeArgs);
template void OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>(OpcodeArgs);
void OpDispatchBuilder::LDMXCSR(OpcodeArgs) {
Ref Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags);
Ref Dest = LoadSource_WithOpSize(GPRClass, Op, Op->Dest, OpSize::i32Bit, Op->Flags);
RestoreMXCSRState(Dest);
}
void OpDispatchBuilder::STMXCSR(OpcodeArgs) {
StoreResult(GPRClass, Op, GetMXCSR(), OpSize::iInvalid);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, GetMXCSR(), OpSize::i32Bit, OpSize::iInvalid);
}
template<IR::OpSize ElementSize>
@@ -3006,7 +3004,7 @@ void OpDispatchBuilder::MOVQ2DQ(OpcodeArgs) {
if constexpr (ToXMM) {
const auto Index = Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0;
Src = _VMov(OpSize::i128Bit, Src);
Src = VZeroExtendOperand(OpSize::i128Bit, Op->Src[0], Src);
StoreXMMRegister(Index, Src);
} else {
// This is simple, just store the result
@@ -3952,8 +3950,8 @@ void OpDispatchBuilder::PTestOpImpl(OpSize Size, Ref Dest, Ref Src) {
Test1 = _VExtractToGPR(Size, OpSize::i16Bit, Test1, 0);
Test2 = _VExtractToGPR(Size, OpSize::i16Bit, Test2, 0);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
auto ZeroConst = Constant(0);
auto OneConst = Constant(1);
Test2 = _Select(FEXCore::IR::COND_NEQ, Test2, ZeroConst, OneConst, ZeroConst);
@@ -3978,7 +3976,7 @@ void OpDispatchBuilder::VTESTOpImpl(OpSize SrcSize, IR::OpSize ElementSize, Ref
const auto ElementSizeInBits = IR::OpSizeAsBits(ElementSize);
const auto MaskConstant = uint64_t {1} << (ElementSizeInBits - 1);
Ref Mask = _VDupFromGPR(SrcSize, ElementSize, _Constant(MaskConstant));
Ref Mask = _VDupFromGPR(SrcSize, ElementSize, Constant(MaskConstant));
Ref AndTest = _VAnd(SrcSize, OpSize::i8Bit, Src2, Src1);
Ref AndNotTest = _VAndn(SrcSize, OpSize::i8Bit, Src2, Src1);
@@ -3992,8 +3990,8 @@ void OpDispatchBuilder::VTESTOpImpl(OpSize SrcSize, IR::OpSize ElementSize, Ref
Ref AndGPR = _VExtractToGPR(SrcSize, OpSize::i16Bit, MaxAnd, 0);
Ref AndNotGPR = _VExtractToGPR(SrcSize, OpSize::i16Bit, MaxAndNot, 0);
Ref ZeroConst = _Constant(0);
Ref OneConst = _Constant(1);
Ref ZeroConst = Constant(0);
Ref OneConst = Constant(1);
Ref CFInv = _Select(IR::COND_NEQ, AndNotGPR, ZeroConst, OneConst, ZeroConst);
@@ -4582,7 +4580,7 @@ void OpDispatchBuilder::VPERMDOp(OpcodeArgs) {
// Get rid of any junk unrelated to the relevant selector index bits (bits [2:0])
Ref IndexMask = _VectorImm(DstSize, OpSize::i32Bit, 0b111);
Ref AddConst = _Constant(0x03020100);
Ref AddConst = Constant(0x03020100);
Ref Repeating3210 = _VDupFromGPR(DstSize, OpSize::i32Bit, AddConst);
Ref FinalIndices = VPERMDIndices(OpSizeFromDst(Op), Indices, IndexMask, Repeating3210);
@@ -4731,7 +4729,7 @@ void OpDispatchBuilder::VPBLENDWOp(OpcodeArgs) {
void OpDispatchBuilder::VZEROOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto IsVZEROALL = DstSize == OpSize::i256Bit;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto NumRegs = Is64BitMode ? 16U : 8U;
if (IsVZEROALL) {
// NOTE: Despite the name being VZEROALL, this will still only ever
@@ -4817,7 +4815,7 @@ Ref OpDispatchBuilder::VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize,
Ref ShiftedIndices = _VShlI(DstSize, OpSize::i8Bit, IndexTrn3, IndexShift);
uint64_t VConstant = IsPD ? 0x0706050403020100 : 0x03020100;
Ref VectorConst = _VDupFromGPR(DstSize, ElementSize, _Constant(VConstant));
Ref VectorConst = _VDupFromGPR(DstSize, ElementSize, Constant(VConstant));
Ref FinalIndices {};
if (Is256Bit) {
@@ -4876,7 +4874,7 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask
IntermediateResult = _VPCMPISTRX(Src1, Src2, Control);
}
Ref ZeroConst = _Constant(0);
Ref ZeroConst = Constant(0);
if (IsMask) {
// For the masked variant of the instructions, if control[6] is set, then we
@@ -4913,7 +4911,7 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask
Ref ResultNoFlags = _Bfe(OpSize::i32Bit, 16, 0, IntermediateResult);
Ref IfZero = _Constant(16 >> (Control & 1));
Ref IfZero = Constant(16 >> (Control & 1));
Ref IfNotZero = UseMSBIndex ? _FindMSB(IR::OpSize::i32Bit, ResultNoFlags) : _FindLSB(IR::OpSize::i32Bit, ResultNoFlags);
Ref Result = _Select(IR::COND_EQ, ResultNoFlags, ZeroConst, IfZero, IfNotZero);
@@ -4946,9 +4944,14 @@ void OpDispatchBuilder::VFMAImpl(OpcodeArgs, IROps IROp, bool Scalar, uint8_t Sr
const OpSize ElementSize = Op->Flags & X86Tables::DecodeFlags::FLAG_OPTION_AVX_W ? OpSize::i64Bit : OpSize::i32Bit;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, Size, Op->Flags);
Ref Src1 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Size, Op->Flags);
Ref Src2 {};
if (Op->Src[1].IsGPR()) {
Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[1], Size, Op->Flags);
} else {
Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
}
Ref Sources[3] = {
Dest,
@@ -5081,9 +5084,9 @@ void OpDispatchBuilder::VPGATHER(OpcodeArgs) {
///< BaseAddr doesn't need to exist, calculate that here.
Ref BaseAddr = VSIB.BaseAddr;
if (BaseAddr && VSIB.Displacement) {
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, _Constant(VSIB.Displacement));
BaseAddr = _Add(OpSize::i64Bit, BaseAddr, Constant(VSIB.Displacement));
} else if (VSIB.Displacement) {
BaseAddr = _Constant(VSIB.Displacement);
BaseAddr = Constant(VSIB.Displacement);
} else if (!BaseAddr) {
BaseAddr = Invalid();
}
@@ -31,21 +31,13 @@ Ref OpDispatchBuilder::GetX87Top() {
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
Ref OpDispatchBuilder::GetX87Tag(Ref Value, Ref AbridgedFTW) {
Ref RegValid = _And(OpSize::i32Bit, _Lshr(OpSize::i32Bit, AbridgedFTW, Value), _Constant(1));
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref X87Valid = _Constant(static_cast<uint8_t>(FPState::X87Tag::Valid));
return _Select(FEXCore::IR::COND_EQ, RegValid, _Constant(0), X87Empty, X87Valid);
}
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref X87Empty = Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref NewAbridgedFTW {};
for (int i = 0; i < 8; i++) {
Ref RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
Ref RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, _Constant(1), _Constant(0));
Ref RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, Constant(1), Constant(0));
if (i) {
NewAbridgedFTW = _Orlshl(OpSize::i32Bit, NewAbridgedFTW, RegValid, i);
@@ -92,9 +84,9 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
_PopStackDestroy();
}
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant Constant) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, Constant);
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit, true);
}
@@ -112,16 +104,16 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
SaveNZCV();
// Extract sign and make integer absolute
auto zero = _Constant(0);
auto zero = Constant(0);
_SubNZCV(OpSize::i64Bit, Data, zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, _Constant(0x8000), zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClassType {COND_MI});
// left justify the absolute integer
auto shift = _Sub(OpSize::i64Bit, _Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shift = _Sub(OpSize::i64Bit, Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shifted = _Lshl(OpSize::i64Bit, absolute, shift);
auto adjusted_exponent = _Sub(OpSize::i64Bit, _Constant(0x3fff + 63), shift);
auto adjusted_exponent = _Sub(OpSize::i64Bit, Constant(0x3fff + 63), shift);
auto zeroed_exponent = _Select(COND_EQ, absolute, zero, zero, adjusted_exponent);
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
@@ -133,7 +125,7 @@ void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, CTX->GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
@@ -156,6 +148,30 @@ void OpDispatchBuilder::FSTToStack(OpcodeArgs) {
void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref Data = _ReadStackValue(0);
// For 16-bit integers, we need to manually check for overflow
// since _F80CVTInt doesn't handle 16-bit overflow detection properly
if (Size == OpSize::i16Bit) {
// Extract the 80-bit float value to check for special cases
// Get the upper 64 bits which contain sign and exponent and then the exponent from upper.
Ref Upper = _VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, Data, 1);
Ref Exponent = _And(OpSize::i64Bit, Upper, Constant(0x7fff));
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
Ref IsSpecial = _NZCVSelect(OpSize::i64Bit, {COND_EQ}, Constant(1), Constant(0));
// For overflow detection, check if exponent indicates a value >= 2^15
// Biased exponent for 2^15 is 0x3fff + 15 = 0x400e
_SubWithFlags(OpSize::i64Bit, Exponent, Constant(0x400e));
Ref IsOverflow = _NZCVSelect(OpSize::i64Bit, {COND_UGE}, Constant(1), Constant(0));
// Set Invalid Operation flag if overflow or special value
Ref InvalidFlag = _Or(OpSize::i64Bit, IsSpecial, IsOverflow);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(InvalidFlag);
}
Data = _F80CVTInt(Size, Data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, OpSize::i8Bit);
@@ -312,14 +328,25 @@ void OpDispatchBuilder::FSUB(OpcodeArgs, IR::OpSize Width, bool Integer, bool Re
}
Ref OpDispatchBuilder::GetX87FTW_Helper() {
Ref FTW = _Constant(0);
// AbridgedFTWIndex has 1-bit per slot (8 slots). Duplicate each bit to get
// 2-bits per slot (16-bit result). Duplicating bits is equivalent to
// Morton interleaving a number with itself. To interleave efficiently two
// bytes, we use the well-known bit twiddling algorithm:
//
// https://graphics.stanford.edu/~seander/bithacks.html#InterleaveBMN
Ref X = LoadContext(AbridgedFTWIndex);
X = _Orlshl(OpSize::i32Bit, X, X, 4);
X = _And(OpSize::i32Bit, X, Constant(0x0f0f0f0f));
X = _Orlshl(OpSize::i32Bit, X, X, 2);
X = _And(OpSize::i32Bit, X, Constant(0x33333333));
X = _Orlshl(OpSize::i32Bit, X, X, 1);
X = _And(OpSize::i32Bit, X, Constant(0x55555555));
X = _Orlshl(OpSize::i32Bit, X, X, 1);
for (int i = 0; i < 8; i++) {
Ref RegTag = GetX87Tag(_Constant(i), LoadContext(AbridgedFTWIndex));
FTW = _Orlshl(OpSize::i32Bit, FTW, RegTag, i * 2);
}
return FTW;
// The above sequence sets valid to 11 and empty to 00, so invert to finalize.
static_assert(static_cast<uint8_t>(FPState::X87Tag::Valid) == 0b00);
static_assert(static_cast<uint8_t>(FPState::X87Tag::Empty) == 0b11);
return _Xor(OpSize::i32Bit, X, Constant(0xffff));
}
void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
@@ -356,33 +383,33 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
}
@@ -412,13 +439,13 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 1));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, Constant(IR::OpSizeToSize(Size) * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
Ref MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(IR::OpSizeToSize(Size) * 2));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, Constant(IR::OpSizeToSize(Size) * 2));
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
}
}
@@ -452,44 +479,44 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto OneConst = Constant(1);
auto SevenConst = Constant(7);
const auto LoadSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
@@ -501,9 +528,9 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(OpSize::i128Bit, OpSize::i16Bit, data, 4);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// reset to default
FNINIT(Op);
@@ -520,29 +547,29 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = NewFCW;
auto roundShift = _Constant(10);
auto roundMask = _Constant(3);
auto roundShift = Constant(10);
auto roundMask = Constant(3);
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
roundingMode = _And(OpSize::i32Bit, roundingMode, roundMask);
_SetRoundingMode(roundingMode, false, roundingMode);
}
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto OneConst = Constant(1);
auto SevenConst = Constant(7);
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
auto low = Constant(~0ULL);
auto high = Constant(0xFFFF);
Ref Mask = _VLoadTwoGPRs(low, high);
const auto StoreSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
@@ -558,9 +585,9 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref RegHigh =
_LoadMem(FPRClass, OpSize::i16Bit, Mem, _Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_LoadMem(FPRClass, OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(OpSize::i128Bit, OpSize::i16Bit, 4, 0, Reg, RegHigh);
if (ReducedPrecisionMode) {
Reg = _F80CVT(OpSize::i64Bit, Reg); // Convert to double precision
@@ -589,13 +616,13 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
if (Offset != 0) {
_F80StackXchange(Offset);
}
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
}
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0x3FF0000000000000)) :
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
}
@@ -636,7 +663,7 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
if (WhichFlags == FCOMIFlags::FLAGS_X87) {
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
} else {
@@ -646,7 +673,7 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
// PF is stored inverted, so invert from the host flag.
// TODO: This could perhaps be optimized?
auto PF = _Xor(OpSize::i32Bit, HostFlag_Unordered, _Constant(1));
auto PF = _Xor(OpSize::i32Bit, HostFlag_Unordered, Constant(1));
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(PF);
}
@@ -668,7 +695,7 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
HostFlag_ZF = _Or(OpSize::i32Bit, HostFlag_ZF, HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
}
@@ -676,7 +703,7 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
void OpDispatchBuilder::X87OpHelper(OpcodeArgs, FEXCore::IR::IROps IROp, bool ZeroC2) {
DeriveOp(Result, IROp, _F80SCALEStack());
if (ZeroC2) {
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(Constant(0));
}
}
@@ -702,7 +729,7 @@ void OpDispatchBuilder::X87ModifySTP(OpcodeArgs, bool Inc) {
Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
// Start with the top value
auto Top = T ? T : GetX87Top();
Ref FSW = _Lshl(OpSize::i64Bit, Top, _Constant(11));
Ref FSW = _Lshl(OpSize::i64Bit, Top, Constant(11));
// We must construct the FSW from our various bits
auto C0 = GetRFLAG(FEXCore::X86State::X87FLAG_C0_LOC);
@@ -717,6 +744,9 @@ Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
auto C3 = GetRFLAG(FEXCore::X86State::X87FLAG_C3_LOC);
FSW = _Orlshl(OpSize::i64Bit, FSW, C3, 14);
auto IE = GetRFLAG(FEXCore::X86State::X87FLAG_IE_LOC);
FSW = _Or(OpSize::i64Bit, FSW, IE);
return FSW;
}
@@ -730,14 +760,14 @@ void OpDispatchBuilder::X87FNSTSW(OpcodeArgs) {
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
auto Zero = _Constant(0);
auto Zero = Constant(0);
if (ReducedPrecisionMode) {
_SetRoundingMode(Zero, false, Zero);
}
// Init FCW to 0x037F
auto NewFCW = _Constant(OpSize::i16Bit, 0x037F);
auto NewFCW = Constant(0x037F);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Set top to zero
@@ -797,8 +827,8 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
default: LOGMAN_MSG_A_FMT("Unhandled FCMOV op: 0x{:x}", Opcode); break;
}
auto ZeroConst = _Constant(0);
auto AllOneConst = _Constant(0xffff'ffff'ffff'ffffull);
auto ZeroConst = Constant(0);
auto AllOneConst = Constant(0xffff'ffff'ffff'ffffull);
Ref SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SrcCond);
@@ -817,8 +847,8 @@ void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
// Claim this is a normal number
// We don't support anything else
auto TopValid = _StackValidTag(0);
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
auto ZeroConst = Constant(0);
auto OneConst = Constant(1);
// In the case of top being invalid then C3:C2:C0 is 0b101
auto C3 = _Select(FEXCore::IR::COND_NEQ, TopValid, OneConst, OneConst, ZeroConst);
@@ -36,12 +36,12 @@ void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
}
}
@@ -87,7 +87,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
}
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(Num));
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit, true);
}
@@ -377,21 +377,21 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Gpr = _VExtractToGPR(OpSize::i64Bit, OpSize::i64Bit, Node, 0);
// zero case
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, _Constant(0xfff0'0000'0000'0000UL));
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = _Sub(OpSize::i64Bit, ExpNZ, _Constant(1023));
ExpNZ = _Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
Ref ExpNZV = _Float_FromGPR_S(OpSize::i64Bit, OpSize::i64Bit, ExpNZ);
Ref SigNZ = _And(OpSize::i64Bit, Gpr, _Constant(0x800f'ffff'ffff'ffffLL));
SigNZ = _Or(OpSize::i64Bit, SigNZ, _Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZ = _And(OpSize::i64Bit, Gpr, Constant(0x800f'ffff'ffff'ffffLL));
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
_TestNZ(OpSize::i64Bit, Gpr, _Constant(0x7fff'ffff'ffff'ffffUL));
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, ExpZV, ExpNZV);
@@ -187,7 +187,7 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END | FLAGS_CALL , 4, nullptr}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1, nullptr}},
@@ -128,7 +128,7 @@ std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps = []() co
// GROUP 5
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END | FLAGS_CALL , 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END, 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0, nullptr}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 5), 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END, 0, nullptr}},
@@ -223,11 +223,11 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// GROUP 12
{OPD(TYPE_GROUP_12, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 2), 1, X86InstInfo{"PSRLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 4), 1, X86InstInfo{"PSRAW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 6), 1, X86InstInfo{"PSLLW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_12, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_12, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -260,11 +260,11 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// GROUP 13
{OPD(TYPE_GROUP_13, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 2), 1, X86InstInfo{"PSRLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 4), 1, X86InstInfo{"PSRAD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 6), 1, X86InstInfo{"PSLLD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_13, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_13, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -297,11 +297,11 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
// GROUP 14
{OPD(TYPE_GROUP_14, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 2), 1, X86InstInfo{"PSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 6), 1, X86InstInfo{"PSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(TYPE_GROUP_14, PF_NONE, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_14, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -347,7 +347,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 5), 1, X86InstInfo{"INCSSPQ", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"CLRSSBSY", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"UMONITOR/CLRSSBSY", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -365,7 +365,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_15, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 6), 1, X86InstInfo{"UMWAIT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// GROUP 16
@@ -384,8 +384,8 @@ std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps = []() consteval {
{0x24, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SD", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2B, 1, X86InstInfo{"MOVNTSD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSD2SI", TYPE_INST, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2D, 1, X86InstInfo{"CVTSD2SI", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSD2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2D, 1, X86InstInfo{"CVTSD2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2E, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x30, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
@@ -20,21 +20,21 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS",TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b00, 0x14), 1, X86InstInfo{"VUNPCKLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x14), 1, X86InstInfo{"VUNPCKLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -42,26 +42,26 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b00, 0x15), 1, X86InstInfo{"VUNPCKHPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x15), 1, X86InstInfo{"VUNPCKHPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOV(L)HPS",TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b10, 0x16), 1, X86InstInfo{"VMOVSHDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b00, 0x50), 1, X86InstInfo{"VMOVMSKPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x50), 1, X86InstInfo{"VMOVMSKPD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b00, 0x51), 1, X86InstInfo{"VSQRTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x51), 1, X86InstInfo{"VSQRTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x51), 1, X86InstInfo{"VSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x51), 1, X86InstInfo{"VSQRTSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x51), 1, X86InstInfo{"VSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x51), 1, X86InstInfo{"VSQRTSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x52), 1, X86InstInfo{"VRSQRTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x52), 1, X86InstInfo{"VRSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x52), 1, X86InstInfo{"VRSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x53), 1, X86InstInfo{"VRCPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -100,11 +100,11 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b00, 0xC2), 1, X86InstInfo{"VCMPccPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC2), 1, X86InstInfo{"VCMPccPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b10, 0xC2), 1, X86InstInfo{"VCMPccSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b11, 0xC2), 1, X86InstInfo{"VCMPccSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b10, 0xC2), 1, X86InstInfo{"VCMPccSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_L_IGNORE | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b11, 0xC2), 1, X86InstInfo{"VCMPccSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_L_IGNORE | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC4), 1, X86InstInfo{"VPINSRW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_SF_SRC_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC5), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC4), 1, X86InstInfo{"VPINSRW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_SF_SRC_GPR | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(1, 0b01, 0xC5), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(1, 0b00, 0xC6), 1, X86InstInfo{"VSHUFPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC6), 1, X86InstInfo{"VSHUFPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -119,38 +119,38 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b00, 0x29), 1, X86InstInfo{"VMOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x29), 1, X86InstInfo{"VMOVAPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2A), 1, X86InstInfo{"VCVTSI2SS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b11, 0x2A), 1, X86InstInfo{"VCVTSI2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b10, 0x2A), 1, X86InstInfo{"VCVTSI2SS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x2A), 1, X86InstInfo{"VCVTSI2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_INST, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b10, 0x2D), 1, X86InstInfo{"VCVTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b11, 0x2D), 1, X86InstInfo{"VCVTSD2SI", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b10, 0x2D), 1, X86InstInfo{"VCVTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x2D), 1, X86InstInfo{"VCVTSD2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x2E), 1, X86InstInfo{"VUCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2E), 1, X86InstInfo{"VUCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x2E), 1, X86InstInfo{"VUCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b01, 0x2E), 1, X86InstInfo{"VUCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x59), 1, X86InstInfo{"VMULPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x59), 1, X86InstInfo{"VMULPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x59), 1, X86InstInfo{"VMULSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x59), 1, X86InstInfo{"VMULSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x59), 1, X86InstInfo{"VMULSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x59), 1, X86InstInfo{"VMULSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x5A), 1, X86InstInfo{"VCVTPS2PD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5A), 1, X86InstInfo{"VCVTPD2PS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5A), 1, X86InstInfo{"VCVTSS2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5A), 1, X86InstInfo{"VCVTSD2SS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5A), 1, X86InstInfo{"VCVTSS2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_L_IGNORE | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5A), 1, X86InstInfo{"VCVTSD2SS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_L_IGNORE |FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5B), 1, X86InstInfo{"VCVTDQ2PS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5B), 1, X86InstInfo{"VCVTPS2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -158,23 +158,23 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b00, 0x5C), 1, X86InstInfo{"VSUBPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5C), 1, X86InstInfo{"VSUBPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5C), 1, X86InstInfo{"VSUBSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5C), 1, X86InstInfo{"VSUBSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5C), 1, X86InstInfo{"VSUBSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x5C), 1, X86InstInfo{"VSUBSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x5D), 1, X86InstInfo{"VMINPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5D), 1, X86InstInfo{"VMINPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5D), 1, X86InstInfo{"VMINSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5D), 1, X86InstInfo{"VMINSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5D), 1, X86InstInfo{"VMINSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x5D), 1, X86InstInfo{"VMINSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x5E), 1, X86InstInfo{"VDIVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5E), 1, X86InstInfo{"VDIVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5E), 1, X86InstInfo{"VDIVSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5E), 1, X86InstInfo{"VDIVSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5E), 1, X86InstInfo{"VDIVSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x5E), 1, X86InstInfo{"VDIVSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b00, 0x5F), 1, X86InstInfo{"VMAXPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5F), 1, X86InstInfo{"VMAXPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5F), 1, X86InstInfo{"VMAXSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5F), 1, X86InstInfo{"VMAXSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5F), 1, X86InstInfo{"VMAXSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b11, 0x5F), 1, X86InstInfo{"VMAXSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(1, 0b01, 0x68), 1, X86InstInfo{"VPUNPCKHBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x69), 1, X86InstInfo{"VPUNPCKHWD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -182,7 +182,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b01, 0x6B), 1, X86InstInfo{"VPACKSSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6C), 1, X86InstInfo{"VPUNPCKLQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6D), 1, X86InstInfo{"VPUNPCKHQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6E), 1, X86InstInfo{"VMOV*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x6E), 1, X86InstInfo{"VMOV*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0 | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -193,8 +193,8 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b01, 0x7D), 1, X86InstInfo{"VHSUBPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x7D), 1, X86InstInfo{"VHSUBPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_VEX_L_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -212,7 +212,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b01, 0xD3), 1, X86InstInfo{"VPSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD5), 1, X86InstInfo{"VPMULLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_L_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD7), 1, X86InstInfo{"VPMOVMSKB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(1, 0b01, 0xD8), 1, X86InstInfo{"VPSUBUSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -246,7 +246,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF1), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF2), 1, X86InstInfo{"VPSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -254,7 +254,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(1, 0b01, 0xF4), 1, X86InstInfo{"VPMULUDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF5), 1, X86InstInfo{"VPMADDWD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF6), 1, X86InstInfo{"VPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF7), 1, X86InstInfo{"VMASKMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF7), 1, X86InstInfo{"VMASKMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(1, 0b01, 0xF8), 1, X86InstInfo{"VPSUBB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF9), 1, X86InstInfo{"VPSUBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -278,18 +278,18 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b01, 0x09), 1, X86InstInfo{"VPSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0A), 1, X86InstInfo{"VPSIGND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0B), 1, X86InstInfo{"VPMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0C), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0D), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0E), 1, X86InstInfo{"VTESTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0F), 1, X86InstInfo{"VTESTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0C), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0D), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0E), 1, X86InstInfo{"VTESTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0F), 1, X86InstInfo{"VTESTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x13), 1, X86InstInfo{"VCVTPH2PS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x16), 1, X86InstInfo{"VPERMPS", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x13), 1, X86InstInfo{"VCVTPH2PS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x16), 1, X86InstInfo{"VPERMPS", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x17), 1, X86InstInfo{"VPTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x18), 1, X86InstInfo{"VBROADCASTSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x19), 1, X86InstInfo{"VBROADCASTSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1A), 1, X86InstInfo{"VBROADCASTF128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x18), 1, X86InstInfo{"VBROADCASTSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x19), 1, X86InstInfo{"VBROADCASTSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1A), 1, X86InstInfo{"VBROADCASTF128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_SF_MOD_MEM_ONLY | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1C), 1, X86InstInfo{"VPABSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1D), 1, X86InstInfo{"VPABSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1E), 1, X86InstInfo{"VPABSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -305,10 +305,10 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b01, 0x29), 1, X86InstInfo{"VPCMPEQQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2B), 1, X86InstInfo{"VPACKUSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2C), 1, X86InstInfo{"VMASKMOVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2D), 1, X86InstInfo{"VMASKMOVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2E), 1, X86InstInfo{"VMASKMOVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2F), 1, X86InstInfo{"VMASKMOVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2C), 1, X86InstInfo{"VMASKMOVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2D), 1, X86InstInfo{"VMASKMOVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2E), 1, X86InstInfo{"VMASKMOVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2F), 1, X86InstInfo{"VMASKMOVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x30), 1, X86InstInfo{"VPMOVZXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x31), 1, X86InstInfo{"VPMOVZXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -316,7 +316,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b01, 0x33), 1, X86InstInfo{"VPMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x34), 1, X86InstInfo{"VPMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x35), 1, X86InstInfo{"VPMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x36), 1, X86InstInfo{"VPERMD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x36), 1, X86InstInfo{"VPERMD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x37), 1, X86InstInfo{"VPCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x38), 1, X86InstInfo{"VPMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -329,17 +329,17 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b01, 0x3F), 1, X86InstInfo{"VPMAXUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x40), 1, X86InstInfo{"VPMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x41), 1, X86InstInfo{"VPHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x41), 1, X86InstInfo{"VPHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 0, nullptr}},
{OPD(2, 0b01, 0x45), 1, X86InstInfo{"VPSRLV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x46), 1, X86InstInfo{"VPSRAVD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x46), 1, X86InstInfo{"VPSRAVD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x47), 1, X86InstInfo{"VPSLLV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x58), 1, X86InstInfo{"VPBROADCASTD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x59), 1, X86InstInfo{"VPBROADCASTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x5A), 1, X86InstInfo{"VBROADCASTI128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x58), 1, X86InstInfo{"VPBROADCASTD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x59), 1, X86InstInfo{"VPBROADCASTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x5A), 1, X86InstInfo{"VBROADCASTI128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_SF_MOD_MEM_ONLY | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x78), 1, X86InstInfo{"VPBROADCASTB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x79), 1, X86InstInfo{"VPBROADCASTW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x78), 1, X86InstInfo{"VPBROADCASTB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x79), 1, X86InstInfo{"VPBROADCASTW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x8C), 1, X86InstInfo{"VPMASKMOV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x8E), 1, X86InstInfo{"VPMASKMOV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -353,31 +353,31 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b01, 0x97), 1, X86InstInfo{"VFMSUBADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x98), 1, X86InstInfo{"VFMADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x99), 1, X86InstInfo{"VFMADD132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x99), 1, X86InstInfo{"VFMADD132_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0x9A), 1, X86InstInfo{"VFMSUB132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9B), 1, X86InstInfo{"VFMSUB132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9B), 1, X86InstInfo{"VFMSUB132_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0x9C), 1, X86InstInfo{"VFNMADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9D), 1, X86InstInfo{"VFNMADD132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9D), 1, X86InstInfo{"VFNMADD132_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0x9E), 1, X86InstInfo{"VFNMSUB132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9F), 1, X86InstInfo{"VFNMSUB132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x9F), 1, X86InstInfo{"VFNMSUB132_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xA8), 1, X86InstInfo{"VFMADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xA9), 1, X86InstInfo{"VFMADD213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xA9), 1, X86InstInfo{"VFMADD213_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xAA), 1, X86InstInfo{"VFMSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAB), 1, X86InstInfo{"VFMSUB213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAB), 1, X86InstInfo{"VFMSUB213_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xAC), 1, X86InstInfo{"VFNMADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAD), 1, X86InstInfo{"VFNMADD213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAD), 1, X86InstInfo{"VFNMADD213_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xAE), 1, X86InstInfo{"VFNMSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAF), 1, X86InstInfo{"VFNMSUB213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xAF), 1, X86InstInfo{"VFNMSUB213_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xB8), 1, X86InstInfo{"VFMADD231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xB9), 1, X86InstInfo{"VFMADD231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xB9), 1, X86InstInfo{"VFMADD231_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xBA), 1, X86InstInfo{"VFMSUB231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBB), 1, X86InstInfo{"VFMSUB231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBB), 1, X86InstInfo{"VFMSUB231_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xBC), 1, X86InstInfo{"VFNMADD231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBD), 1, X86InstInfo{"VFNMADD231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBD), 1, X86InstInfo{"VFNMADD231_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xBE), 1, X86InstInfo{"VFNMSUB231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBF), 1, X86InstInfo{"VFNMSUB231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xBF), 1, X86InstInfo{"VFNMSUB231_S", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_IGNORE, 0, nullptr}},
{OPD(2, 0b01, 0xA6), 1, X86InstInfo{"VFMADDSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xA7), 1, X86InstInfo{"VFMSUBADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -398,25 +398,25 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(2, 0b10, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b11, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b00, 0xF5), 1, X86InstInfo{"BZHI", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b00, 0xF5), 1, X86InstInfo{"BZHI", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_2ND_SRC, 0, nullptr}},
// AMD reference manual is incorrect. PEXT actually maps to 0b10, not 0b01.
{OPD(2, 0b10, 0xF5), 1, X86InstInfo{"PEXT", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF5), 1, X86InstInfo{"PDEP", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b10, 0xF5), 1, X86InstInfo{"PEXT", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF5), 1, X86InstInfo{"PDEP", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b00, 0xF7), 1, X86InstInfo{"BEXTR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF7), 1, X86InstInfo{"SHLX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b10, 0xF7), 1, X86InstInfo{"SARX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b00, 0xF7), 1, X86InstInfo{"BEXTR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF7), 1, X86InstInfo{"SHLX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b10, 0xF7), 1, X86InstInfo{"SARX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_2ND_SRC, 0, nullptr}},
// VEX Map 3
{OPD(3, 0b01, 0x00), 1, X86InstInfo{"VPERMQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x01), 1, X86InstInfo{"VPERMPD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x02), 1, X86InstInfo{"VPBLENDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x04), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x05), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x06), 1, X86InstInfo{"VPERM2F128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x00), 1, X86InstInfo{"VPERMQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_1 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x01), 1, X86InstInfo{"VPERMPD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_1 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x02), 1, X86InstInfo{"VPBLENDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x04), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x05), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x06), 1, X86InstInfo{"VPERM2F128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_REX_W_0 | FLAGS_XMM_FLAGS | FLAGS_VEX_L_1, 1, nullptr}},
{OPD(3, 0b01, 0x08), 1, X86InstInfo{"VROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x09), 1, X86InstInfo{"VROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -427,41 +427,41 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(3, 0b01, 0x0E), 1, X86InstInfo{"VPBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0F), 1, X86InstInfo{"VPALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x14), 1, X86InstInfo{"VPEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x15), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x16), 1, X86InstInfo{"VPEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x17), 1, X86InstInfo{"VEXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x14), 1, X86InstInfo{"VPEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x15), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x16), 1, X86InstInfo{"VPEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x17), 1, X86InstInfo{"VEXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x18), 1, X86InstInfo{"VINSERTF128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x19), 1, X86InstInfo{"VEXTRACTF128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x1D), 1, X86InstInfo{"VCVTPS2PH", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x18), 1, X86InstInfo{"VINSERTF128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x19), 1, X86InstInfo{"VEXTRACTF128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x1D), 1, X86InstInfo{"VCVTPS2PH", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x20), 1, X86InstInfo{"VPINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(3, 0b01, 0x21), 1, X86InstInfo{"VINSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x22), 1, X86InstInfo{"VPINSR{D,Q}", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(3, 0b01, 0x20), 1, X86InstInfo{"VPINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(3, 0b01, 0x21), 1, X86InstInfo{"VINSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x22), 1, X86InstInfo{"VPINSR{D,Q}", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
{OPD(3, 0b01, 0x38), 1, X86InstInfo{"VINSERTI128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x39), 1, X86InstInfo{"VEXTRACTI128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x38), 1, X86InstInfo{"VINSERTI128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x39), 1, X86InstInfo{"VEXTRACTI128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x40), 1, X86InstInfo{"VDPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(3, 0b01, 0x42), 1, X86InstInfo{"VMPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_L_1 | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4A), 1, X86InstInfo{"VBLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4B), 1, X86InstInfo{"VBLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4C), 1, X86InstInfo{"VPBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4A), 1, X86InstInfo{"VBLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4B), 1, X86InstInfo{"VBLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4C), 1, X86InstInfo{"VPBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_REX_W_0 | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x5C), 1, X86InstInfo{"VFMADDSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x5D), 1, X86InstInfo{"VFMADDSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x5E), 1, X86InstInfo{"VFMSUBADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x5F), 1, X86InstInfo{"VFMSUBADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x60), 1, X86InstInfo{"VPCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x61), 1, X86InstInfo{"VPCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x62), 1, X86InstInfo{"VPCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x63), 1, X86InstInfo{"VPCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x60), 1, X86InstInfo{"VPCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(3, 0b01, 0x61), 1, X86InstInfo{"VPCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(3, 0b01, 0x62), 1, X86InstInfo{"VPCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(3, 0b01, 0x63), 1, X86InstInfo{"VPCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_L_0, 1, nullptr}},
{OPD(3, 0b01, 0x68), 1, X86InstInfo{"VFMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x69), 1, X86InstInfo{"VFMADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
@@ -481,9 +481,9 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
{OPD(3, 0b01, 0x7E), 1, X86InstInfo{"VFNMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0x7F), 1, X86InstInfo{"VFNMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_L_0 | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b11, 0xF0), 1, X86InstInfo{"RORX", TYPE_INST, FLAGS_MODRM, 1, nullptr}},
{OPD(3, 0b11, 0xF0), 1, X86InstInfo{"RORX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_L_0, 1, nullptr}},
// VEX Map 4 - 31 (Reserved)
};
@@ -500,21 +500,21 @@ std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps = []() conste
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
constexpr U8U8InfoStruct VEXGroupTable[] = {
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b010), 1, X86InstInfo{"VPSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b100), 1, X86InstInfo{"VPSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b110), 1, X86InstInfo{"VPSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b010), 1, X86InstInfo{"VPSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b100), 1, X86InstInfo{"VPSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b110), 1, X86InstInfo{"VPSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b010), 1, X86InstInfo{"VPSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b011), 1, X86InstInfo{"VPSRLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b110), 1, X86InstInfo{"VPSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b111), 1, X86InstInfo{"VPSLLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b010), 1, X86InstInfo{"VPSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b011), 1, X86InstInfo{"VPSRLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b110), 1, X86InstInfo{"VPSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b111), 1, X86InstInfo{"VPSLLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b010), 1, X86InstInfo{"VLDMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b011), 1, X86InstInfo{"VSTMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b010), 1, X86InstInfo{"VLDMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_L_0 | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b011), 1, X86InstInfo{"VSTMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_L_0 | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b001), 1, X86InstInfo{"BLSR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b010), 1, X86InstInfo{"BLSMSK", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
@@ -32,13 +32,15 @@ constexpr uint32_t FLAG_REX_WIDENING = (1 << 7);
constexpr uint32_t FLAG_REX_XGPR_B = (1 << 8);
constexpr uint32_t FLAG_REX_XGPR_X = (1 << 9);
constexpr uint32_t FLAG_REX_XGPR_R = (1 << 10);
constexpr uint32_t FLAG_ES_PREFIX = (1 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (1 << 12);
constexpr uint32_t FLAG_SS_PREFIX = (1 << 13);
constexpr uint32_t FLAG_DS_PREFIX = (1 << 14);
constexpr uint32_t FLAG_FS_PREFIX = (1 << 15);
constexpr uint32_t FLAG_GS_PREFIX = (1 << 16);
constexpr uint32_t FLAG_SEGMENTS = (0b11'1111 << 11);
constexpr uint32_t FLAG_NO_PREFIX = (0b000 << 11);
constexpr uint32_t FLAG_ES_PREFIX = (0b001 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (0b010 << 11);
constexpr uint32_t FLAG_SS_PREFIX = (0b011 << 11);
constexpr uint32_t FLAG_DS_PREFIX = (0b100 << 11);
constexpr uint32_t FLAG_FS_PREFIX = (0b101 << 11);
constexpr uint32_t FLAG_GS_PREFIX = (0b110 << 11);
constexpr uint32_t FLAG_SEGMENTS = (0b111 << 11);
// Bits 14, 15, 16 - Unused
constexpr uint32_t FLAG_REP_PREFIX = (1 << 17);
constexpr uint32_t FLAG_REPNE_PREFIX = (1 << 18);
@@ -353,6 +355,14 @@ constexpr InstFlagType FLAGS_VEX_1ST_SRC = (0b10ULL << 22);
constexpr InstFlagType FLAGS_VEX_2ND_SRC = (0b11ULL << 22);
// Whether or not the instruction has a VSIB byte
constexpr InstFlagType FLAGS_VEX_VSIB = (1ULL << 24);
constexpr InstFlagType FLAGS_VEX_L_IGNORE = (1ULL << 25);
constexpr InstFlagType FLAGS_VEX_L_0 = (1ULL << 26);
constexpr InstFlagType FLAGS_VEX_L_1 = (1ULL << 27);
constexpr InstFlagType FLAGS_REX_W_0 = (1ULL << 28);
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
Loaded 100 of 363 files, more files were not shown because too many files have changed in this diff. Show more